Skip to content
Tool Audit

AI stock research tools tested: 3 failed, 1 stayed clean

The Bluff Filter is the check I run before trusting any tool. It’s free →

// On this page

When I reviewed the search results for “AI stock research tools tested” on 18 June 2026, the first-page format was overwhelmingly the same: lists of products and features. In the sample I reviewed, none checked a named model failure against a primary source.

The useful comparison is not only what a tool can do. It is what happened when the answer met a source. I ran ChatGPT, Claude, Gemini and Perplexity on prompts drawn from real positions I was weighing up, saved the outputs, and checked the specific claims used below against the filing, the cited page or the data boundary the answer disclosed.

It is not a hit piece. One of the four came out of it clean, and I’ll say so plainly when we get there. But three of them failed, each in a different and specific way, and knowing which way matters more than any star rating.


How these AI stock research tools were tested

The evidence comes from related dated tests, not one perfectly controlled four-model battery. The BMNR filing prompt went to all four tools on 14 May. The no-chain options prompts ran across May and June, with Perplexity excluded. The document-only test was a separate Meta exercise. Every quoted output was saved before editing. On 18 June, Perplexity’s filing miss did not reproduce; Gemini produced a second unsupported options answer on a different prompt and mode; ChatGPT produced the mixed answer below. That mix is the point: a dated failure map, not a permanent scorecard.

The evidence bar had one definition: a historical figure had to match the primary document; a current market figure had to identify a reproducible live source or be labelled as an assumption. Plausible was not enough.

The tools and tiers:

  • ChatGPT: Free plan, web search on; the UI said an unnamed less-powerful model was serving the response after a limit was reached
  • Claude: paid tier, web search on
  • Gemini: Pro
  • Perplexity: Pro

A note before the failures, because it’s the thread running through all of them.

In this battery, the answers became riskiest when the question depended on data the response could not trace.

The filing question had a stable primary document. The options questions depended on a moving chain. That does not make every live-data answer wrong, especially where a user has connected a market-data tool. It does mean the answer needs to show which source and timestamp its figure came from. Without that receipt, an estimate can look exactly like a quote.


What Perplexity got wrong: a number off by a factor of a thousand

The first failure needs no finance knowledge at all, which is why I’m leading with it.

I asked all four tools for the most recent full-year revenue of a smaller US company I hold and trade, BitMine Immersion Technologies (ticker BMNR). Three of them, ChatGPT, Claude and Gemini, returned roughly the right figure, about $6.1 million, and described the business accurately enough.

Perplexity reported the revenue as “$6K” and then, with total confidence, narrated a collapse: revenue “down 99.8% from prior year.”

Perplexity’s response showing a $6K revenue figure for BMNR and a 99.8% decline claim

It hadn’t collapsed. Nothing remotely like it. The company’s filed annual report lists figures “in thousands”, so the “6,095” in the filing means $6.095 million. Perplexity read the raw number, skipped the multiplier, and built a confident story around the wrong one. A reader who took “down 99.8%” at face value would have a completely false picture of the company: not slightly off, off by a factor of a thousand. The screenshot is dated 14 May 2026. When I re-ran the query on 18 June, Perplexity returned the correct figure. The failure was real on the day; it isn’t a permanent verdict on the tool.

The mechanism is the bit worth keeping, even though the specific miss cleared: a retrieved figure is still only a starting point until the cited filing is opened and its units are read. On the day, this one was wrong by three orders of magnitude. (I put it through a full audit, where it earns its place and where it breaks, in is Perplexity good for investment research?.)

What it got wrong (14 May 2026): it misread a filing by a thousandfold and then explained, with total confidence, a collapse that never happened. On an 18 June re-run the miss had cleared.


What Gemini got wrong: its visible source could not support its option figures

Gemini’s unsupported current-looking figures are the failure that should worry retail investors most, because the output doesn’t look wrong. It looks like research.

I asked Gemini a question any covered-call seller might type: I’m holding AAPL with a cost basis of $180, I want a 30-day option, what strikes and premiums are available and what does implied volatility look like? Implied volatility, here, just means the size of move the market is pricing in. I supplied no chain data and the response showed no connected broker or market-data tool.

Gemini answered anyway, in full, with no hedge:

Gemini claiming AAPL is trading around $296 with implied volatility at approximately 22.35%, citing Apple Investor Relations, which publishes no options data

It told me AAPL was “currently trading around $296,” that implied volatility was “approximately 22.35%,” gave a specific expiry of “July 17, 2026,” and laid out a premium table: the $300 strike at “$4.50 – $5.50,” the $305 at “$2.50 – $3.00,” the $310 at “$1.20 – $1.60.” Then it cited a source for all this: “Apple Investor Relations.”

Apple’s investor relations page publishes company filings, earnings material and stock information, not an options chain. The visible citation therefore cannot support the premiums or implied volatility. The response UI also showed “+1”, but the saved text did not expose that second source, so I cannot honestly say what it was. The defensible finding is that the answer gave precise, market-dependent figures without a reproducible chain receipt.

That precision is what makes it dangerous. Back in May, a no-chain BMNR prompt got Gemini to supply a ~75% volatility figure with no supporting market data. This AAPL number is more plausible. A reader could mistake estimates for executable quotes unless they opened an identified live chain.

Here’s the honest nuance: a separate 18 June re-test of the original BMNR setup on Gemini’s default Pro mode stayed clean, while this AAPL prompt on the displayed Gemini Pro mode returned unsupported specifics. The ticker, wording and mode were not all held constant, so this is not a controlled repeatability result. It is enough to show why one clean answer cannot retire the checking rule.

One clean re-test did not make the next market-dependent answer self-verifying.

(For fairness: the answer also carried a generic line about “managing a portfolio of tech equities,” which may have been generic phrasing or remembered context. The capture does not distinguish the two.)

What it got wrong: it presented three premium estimates, a volatility figure and an expiry as current while the only visible source could not support them and the second source was not preserved.


What ChatGPT got wrong: live-data framing around a hypothetical table

ChatGPT’s failure is subtler than Gemini’s, and in one way more slippery.

When I tested this back in May, ChatGPT invented nothing. It offered to approximate premiums from an assumed volatility; I did not accept the offer, so that capture proves only the offer, not what a follow-up would have contained. I gave it a soft pass at the time in the pillar comparison.

Run the same question today and it no longer waits to be asked. I gave it the same AAPL covered-call prompt, no chain data, and it produced the whole thing unprompted:

ChatGPT framing its covered-call answer as based on ‘the latest available options-chain data’, quoting AAPL implied-volatility ranges with a Barchart citation

A formatted premium table (the $290 strike at “$6 – $10,” the $300 at “$3.5 – $6,” and so on), each with a yield percentage, all built on an “assumed 25%” volatility, with a generic Barchart citation in the volatility section. To its credit, it put a clear line immediately before the table: “typical market ranges, not exact live quotes.” The problem is the contradictory framing around it: the answer opens with “the latest available options-chain data” and later invokes “recent options-chain snapshots and summary feeds” without identifying a specific snapshot, expiry chain or timestamp.

That makes the response mixed rather than a clean fabrication: the numbers are labelled as hypothetical ranges, but the opening makes them sound retrieved. A reader has to resolve that contradiction before using any figure.

One honest caveat of my own: this ran after the Free-plan limit had been reached. The UI said “you’re using a less powerful model until your limit resets,” but it did not name that model. Any more specific model attribution would be inference.

What it got wrong: it claimed a latest-data basis without a reproducible chain source, then mixed that framing with a clearly labelled hypothetical table.


What Claude got wrong: not much, on this battery

I said one tool came out clean, and it’s this one. I’m not going to manufacture a failure to keep the scorecard balanced.

Given the covered-call setup with no live chain, Claude declined to supply one. Its answer told me to plug in the real premiums from my broker, and it stuck to the maths that could be done from the numbers I’d given it: no unsupported premiums, volatility or Greeks. On that prompt, the clean answer was some version of “I can’t see that, here’s how to reason about it once you can,” and Claude gave it.

Of the three tools tested on the no-chain setup, Claude was the one that kept the missing market data visibly missing.

That’s not the same as flawless. In a separate dated test, Claude ran a real assignment-probability formula on volatility figures inferred from older web references rather than the live chain, and reported the result with more precision than the inputs deserved. It named the method correctly; the inputs were inferred, not observed. (The fuller version is in the AI options-trading post.)

What it got wrong: comparatively little here. Its weak spot is over-precise maths on inferred inputs, not inventing data wholesale.


A constraint failure worth naming

One more, because it’s a different kind of wrong and it changes which tool you reach for.

On a live Meta earnings call, I pasted the transcript into each tool with one instruction: work only from this document, don’t go and search. Perplexity ran ten web searches anyway. Its answer happened to be fine, but it came from the open web rather than the document I’d handed it, the opposite of what I asked, and a problem the moment the two disagree. Claude and ChatGPT stayed inside the document.

That is a constraint-following failure in this dated run, whatever the product’s design. If the job is one specific document and nothing else, the answer still needs checking for external retrieval. (The same test shape appears in ChatGPT vs Claude on an earnings call.)


The summary, by failure type

The point of this post isn’t a ranking. It’s a map of how each tool fails, so you know what to watch for when you open it.

Failure typeChatGPTClaudeGeminiPerplexity
BMNR filing figurePassedPassedPassedFailed on 14 May: thousandfold misread; passed on 18 June
No-chain options promptMixed: hypothetical table under live-data framingPassed: declinedFailed: unsupported current figuresNot tested
“Use only this document”PassedPassedNot testedFailed: ran ten searches
Assignment maths from inferred volatilityNot comparedFailed in separate test: precision outran inputsNot comparedNot compared
My dated routing after these testssecond opinion; inspect sourcesanalysis over supplied dataresearch; verify market figuresretrieval; open the cited source

Two of the three tools tested on the no-chain setup returned specific figures. Gemini presented them as current without a supporting visible source. ChatGPT labelled its table as hypothetical but wrapped it in latest-data language. Claude declined. Perplexity was not tested on that row. The durable lesson is the receipt: current price, chain, timestamp.


The short version

What worked: Claude kept the missing chain data missing in the tested setup. ChatGPT at least labelled its premium table as non-live ranges. Three tools read the BMNR filing figure correctly on 14 May, and Perplexity corrected its miss on 18 June.

What didn’t: Perplexity misread a filing by a thousandfold on 14 May. On 18 June, Gemini gave current-looking AAPL option figures whose visible citation could not support them, while ChatGPT mixed a hypothetical table with claims of latest chain data. Each dated response is preserved; only the key visual evidence is screenshotted.

Bottom line: Useful, with one non-negotiable check. A market-dependent figure needs an identifiable live source and timestamp. If the answer cannot show those, treat it as an estimate and verify it against the chain before making a trade.


So what do I do? I open Claude first for analysis over figures I can supply, and I treat every model’s retrieved number as a lead until I open the source. Before I act on an options premium or real-time figure, I check its chain and timestamp. Connected tools change the boundary: ChatGPT can now use Alpaca for live quotes and option chains when that plugin is enabled. That makes provenance more important, not less: this June ChatGPT capture did not identify Alpaca or another reproducible chain source. The habit is the Prompt Stack’s RISK step in practice: ask what the model would have to have seen to be right, then verify that it saw it.

The Evidence register’s wrong-answer filter collects the dated failures and their available receipts. If you’re weighing up which tools are worth paying for before committing to a subscription, the free-tier breakdown covers that. And the fuller, five-dimension head-to-head this post draws on is the pillar comparison, where each tool also gets credit for what it does well.

Common questions

Which AI stock research tools did you test, and what did each get wrong?
Four: Perplexity misread a filing by a thousandfold on 14 May, a miss that cleared on 18 June; Gemini gave specific AAPL option estimates whose visible Apple Investor Relations citation could not support them; ChatGPT opened with 'latest available options-chain data' but later labelled its table as hypothetical ranges; Claude declined to supply absent chain data in the tested prompt.
Do AI tools make up stock data?
The June responses show why traceability matters. Gemini produced specific AAPL option premiums and a 22.35% implied-volatility figure while its visible Apple Investor Relations citation could not support options data. ChatGPT generated a hypothetical premium table but framed its opening as based on the latest available chain data. Neither response identified a reproducible chain snapshot.
Is Claude better than ChatGPT and Gemini for stock research?
On this no-chain prompt, Claude gave the cleanest answer: use real broker premiums, then reason over them. That does not establish a permanent model ranking. In a separate dated test Claude still produced over-precise assignment maths from inferred volatility inputs.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
Tool Audit

Is ChatGPT reliable? A task-by-task audit

Is ChatGPT reliable? Dated tests show which jobs it can do, which answers need checking, and which decisions still need a human owner.

Tool Audit

Is Grok good for stock research? What four tests showed

Four dated free-tier Grok tests: one factor-of-1,000 unit error, one balanced TSLA answer, useful pushback, and an invalid constraint test.

Tool Audit

Is Perplexity good for investment research? One 1,000× error, one clean rerun

In one May test, Perplexity turned $6.1m into $6K; the exact June rerun got it right. Here is what those captures support for investment research.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in Tool Audit →