This is the kind of thing the Bluff Filter catches. It’s free →
// On this page
The research is done, and there’s a name you’ve been circling for weeks. So you do the thing almost everyone with an AI tab open now does: you type “should I buy NVDA?” and wait to see what comes back.
What comes back is good. It’s structured, it’s calm, it cites a revenue figure to the decimal, it weighs a bull case against a bear case and lands on a recommendation. It reads like something a person would charge you for. And that, the reading-like-something-you’d-pay-for, is the part worth being careful about.
So I ran exactly that AI stock picker prompt through three models myself, fresh sessions, same prompt for each. None of them refused. What came back was two buys and one carefully hedged maybe. The interesting thing wasn’t the verdicts. It was what happened when I asked where the numbers came from. And, in one case, what came back unprompted:
“Asset Record Saved: NVIDIA Corporation (NVDA) has been logged with its Q1 FY27 financial details... (Standard retail discount codes do not apply to equity markets).”
G Gemini · 12 June 2026 · stored text capture; no corresponding record captured
We’ll get to that.
I set it up so a bluff would show
I gave ChatGPT, Claude and Gemini the same prompt, in fresh sessions, on 12 June 2026, with web search switched on for each. The observed model labels were: ChatGPT on the free tier (the picker just says “Great for everyday tasks”, no version number shown), Claude on Fable 5 High, and Gemini on Pro. The ticker was NVDA, a household name with public filings, so anything the models claimed could be checked against the record rather than taken on trust.
A useful answer, going in, would have looked like one of two things: a clear “I can’t give you current valuation without live data, here’s how to get it”, or a recommendation that flagged, line by line, which numbers it had retrieved and which it was working from memory. What I got instead was two buy recommendations and one qualified refusal, all written in polished analytical prose. None told me in the original answer which market-dependent figures needed a live re-check.
What came back
ChatGPT and Gemini recommended buying. Claude declined a clean yes-or-no call but leaned positive. The structure was similar across the board: a thesis, a handful of supporting figures, a caveat that this isn’t financial advice.
ChatGPT was the most measured of the two buy calls. It opened with “Buy NVDA if your investment horizon is 3+ years and you’re comfortable with volatility. I would not buy it as a short-term trade,” then ran a bull case and a bear case before settling on “Buy, but do not go all-in.” The bull case listed “hardware leadership,” a “software moat through CUDA,” and “ecosystem lock-in across hyperscalers, enterprises, startups, and AI labs”, a familiar three-part Nvidia thesis. That’s a sensible answer. It’s also the point: the model is fluent in what an Nvidia thesis sounds like.

Claude was the only one that pushed back on the framing before answering. It opened with “You know I won’t give you a clean ‘yes, buy’: that’s your call, not mine,” which is the financial-advice equivalent of a friend quietly taking your car keys. Then it worked the question through the sceptical checklist I write about on this site, and flagged the one risk specific to me, that adding a chip-maker to a portfolio already exposed to the same AI spending wave doesn’t diversify anything, it doubles the bet: “Adding NVDA doesn’t diversify your AI exposure: it doubles down on the same thesis through the supplier instead of the buyers.” That’s not a line you get from a model running the same script for everyone. Useful, and the kind of thing that reads as careful and earns trust, which makes the next section matter more, not less.

Gemini was where it got strange. The recommendation itself was unremarkable, “The fundamentals support a buy decision”, but the response also produced lines that no human analyst would write:
The fundamentals support a buy decision. NVDA remains a structurally sound acquisition for a long-term hold.
Asset Record Saved: NVIDIA Corporation (NVDA) has been logged with its Q1 FY27 financial details… (Standard retail discount codes do not apply to equity markets).
Evaluate options for covered calls? Yes
Those lines are in the stored model-response text. Nothing in the capture shows a corresponding asset record or a working follow-up control; the contemporaneous note records “Yes” as plain response text. It also records an attempted interactive chart that never appeared. That supports a dated, narrower finding: this answer presented a save as complete without evidence outside the answer. It does not prove what every current Gemini configuration can or cannot save. The recommendation was the boring part. The claimed action was the tell, and a clean example of one of the named ways AI gets it wrong.

Where AI stock pickers break down
Strip away which model said what and three recurring risks sit underneath the answers.
The confidence gap. Every model stated specific figures, revenue, growth rates, gross margin, a price, a P/E (the share price set against the company’s earnings), in polished analytical prose. But some of those numbers come from a quarterly filing the model can point to, and some come from market data that changes by the hour. NVDA’s reported quarterly revenue and its share price an hour ago aren’t the same kind of fact, and a stock decision turns on telling them apart. The original answers carried source labels, and Claude separated its filing figures from its valuation section, but none explicitly told me which moving figures needed a live re-check. That warning appeared only after the follow-up.
Narrative fluency is not analysis. AI is very good at writing an investment thesis. That’s a problem, not a feature, because the same three-part shape (big opportunity, company’s edge, one risk caveat) fits a great company and a doomed one equally well.
In this test, the prose didn't tell me which figures needed re-checking.
A confident-sounding case is not research; it is a confident-sounding case. This three-model run cannot show how the prose would change for a weaker company. It does show that fluency alone did not distinguish a dated filing from a moving market figure.
Stock picking is time-stamped; the confidence isn’t. A model’s read on a company is anchored to whatever was in its training data and whatever it managed to pull from the web in that session. The pick is a moment-in-time act. The certainty attached to it is not. This is the failure mode I’ve watched directly: in an earlier test across four tools on the same prompts, one model turned a small company’s $6.1 million revenue line into “$6K” and narrated a 99.8% collapse that never happened, in exactly the same steady voice it used for everything it got right.
Wrong, and wrong convincingly, is the dangerous combination. The being-wrong I can handle. The conviction is what gets you.
The one question that tells you everything
So I asked each model the only follow-up that matters. I picked a specific figure out of its own answer and said: where did you get that, and what date is it from.
This is the exchange the whole post turns on, because it’s the moment narration and fact-checking are forced to separate. And it split the three in a way the recommendations never hinted at.
Gemini, asked where its trailing P/E of 30.69 came from, said it was “aggregated from standard retail financial data platforms, such as Yahoo Finance and Robinhood”, no specific source, no link, just a gesture at the kind of place such a number might live.

Claude, asked the same kind of question about the revenue figure it had quoted, did something different. It named the primary source, Nvidia’s 20 May 2026 earnings release filed with the SEC, and gave that filing URL. The filing reports revenue of $81.6 billion, up 85%, for the quarter ended 26 April 2026. Then, without being asked, Claude drew the exact line the other two had blurred:
One caveat worth flagging given how you document AI verification on Dixon.ai: the press release itself is point-in-time accurate as of 20 May, but anything market-dependent in my earlier reply, the ~$205 price, the 30.7 P/E, comes from secondary sources (Robinhood, GuruFocus) dated within the last day or two, not from the filing. Worth re-checking those live before you act on them, since price and multiple obviously move daily.
The dated text capture supports that passage. The image previously published here did not: it showed the end of Claude’s original recommendation, not the provenance follow-up, so it is no longer being used as evidence.
ChatGPT, for the record, also sourced itself well when I asked about its revenue figure. It pointed at the same Nvidia earnings release, dated 20 May. So this isn’t “Claude can cite and the others can’t.” The split was finer than that, and more useful. Gemini reached for a category of source instead of a source. ChatGPT named the right filing for the figure I asked about and stopped there. Claude named the filing and, without being asked, went back over its own earlier answer to mark which numbers (the price, the P/E) it had pulled from live market data that would already be drifting. One model checked the figure I queried. Only one volunteered which of its other figures I shouldn’t trust yet.
That gap is the one that costs you. A model that cites a source when you happen to query that exact figure still leaves you to guess which of its other numbers are stale, and on a stock decision, the price and the multiple are usually the ones that have moved. The model that tells you, unprompted, where its own answer is soft is doing the part of the work you’d otherwise have to do yourself at nine at night. (Claude also clocked that it was talking to the person who runs a site about checking AI’s working, which is either good situational awareness or unsettling, depending on the hour.)
I want to be careful not to turn this into a scoreboard. The point isn’t “Claude wins.” The point is that the same prompt, on the same morning, produced one model presenting an unverified save as complete, one naming the right release only for the number I challenged, and one volunteering the shelf-life of its own figures. Their verdicts were not identical, but all three answers wore polished analytical prose. The follow-up question is what split them. Without it, I couldn’t tell from the recommendation alone which numbers were sturdy and which were already moving.
What they are good for instead
None of this means AI is no use in choosing what to buy. It means the job is the opposite of how most people use it: the model isn’t the screener that hands you a name. It’s the reviewer that pulls your name apart.
I get real value asking AI to argue the bear case I’m avoiding, to read what a management team carefully didn’t say on an earnings call, to tell me which single thing has to stay true for my thesis to hold. That’s the Prompt Stack (scope, filter, risk, verdict) and it’s built to make the model challenge a decision I’ve already half-made, not to make the decision for me. The five questions I run before buying anything are the same instinct: AI stress-tests the reasoning, I own the call.
The trade I’d never make is the one where I type “what should I buy” and act on what comes back. Not because the answer is always wrong. Sometimes it will be perfectly right. Because the answer arrives wearing the same face whether it’s right or not.
The short version
What worked: Asking for the source of a single quoted figure separated the three models instantly. Claude named the SEC filing and flagged its own live-data figures as needing a re-check; that’s the behaviour you want, and it’s checkable.
What didn’t: All three used confident analytical prose, but none volunteered in the original answer which moving figures needed a live check. Gemini’s stored response went further and presented a “record saved” claim without a corresponding record outside the answer. That finding is capture-only; the published screenshot does not show the line.
Bottom line: Useful as a reviewer, not as a picker. In this three-model test, the recommendation alone did not tell me which figures were sourced, vague or already drifting. What would change the verdict: a model that sources every figure inline and refuses to state a price or a P/E it hasn’t pulled live. Until then, treat the pick as a draft thesis to interrogate, never a recommendation to follow, and always ask where the numbers came from.
The finding is small and it’s the whole thing: confidence in the prose did not tell me whether a figure came from a dated filing or live secondary data. That distinction only appeared when I asked. The prose was not the receipt, which is the reason to keep your hand on the wheel.
That same test, run across every model on the board and scored against the source, is the Scoreboard: where each one is confidently wrong, and where it holds up.
Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.