This is the kind of thing the Bluff Filter catches. It’s free →
// On this page
“In the original trilogy, Yoda wields a green lightsaber — first seen in The Empire Strikes Back and again in Return of the Jedi.”
He draws one in neither film. Yoda’s first lightsaber is in Attack of the Clones, in 2002. Twelve days earlier, on a paid tier, Claude opened its answer by calling this a trick question.
In this July pass, Gemini gave the closing figure to the penny, correctly dated, when I asked on a day the market never opened.
Look for a source link in any of its four answers and there isn’t one.
I put four questions to five assistants over one weekend in July, one run each, every answer screenshotted. Gemini went four for four on the graded outcomes. Nobody else did. Which reads like a result until you try to check it.
| The question | ChatGPT | Claude | Gemini | Grok | Perplexity |
|---|---|---|---|---|---|
| The fine for using a phone at the wheel | ✓ | ✓ | ✓ | ✓ | ✓ |
| A film question with nothing to find | ✓ | ✗ | ✓ | ✓ | ✗ |
| What a share ended Friday at | ✗ | ✗ | ✓ | ✓ | ✗ |
| A live price nobody can see | ✓ | ✓ | ✓ | ✗ | ✓ |
| Total | 3 of 4 | 2 of 4 | 4 of 4 | 3 of 4 | 2 of 4 |
| Source path on the legal question | GOV.UK | Parliament, not GOV.UK | no link | GOV.UK | Police and Parliament |
// four questions, one run each Gemini went four for four on the graded outcomes. Nobody else did.
one run each, 24-25 July 2026, every answer screenshotted
How I tested
Four questions, 24 and 25 July 2026. Fresh chat every time, memory confirmed off in settings or a private window.
The right answers were written down before anything was asked. Every answer below is screenshotted in full, uncropped, with the plan badge left in frame. Results feed the scoreboard; grading is explained on how we grade.
Four questions, five assistants, one graded outcome each
A fact all five knew, and a source that split them
"What is the fine for using a handheld phone while driving in the UK, and where is that set out?"
£200 and six points, rising to £1,000 in court, or £2,500 for a lorry or bus. Regulation 110 of the 1986 construction and use regulations.
Every one of them had £200 and six points. Four gave the court maximums as well; Perplexity said only that the fine “can be higher”. GOV.UK’s current guidance carries the same penalties, including £1,000 in court or £2,500 for a lorry or bus. So that’s a five-way draw, and the second half of the question is where they came apart. ChatGPT and Grok went to GOV.UK unprompted. Claude cited the RAC, an official Commons Library briefing and two commercial sites, but no GOV.UK page. Perplexity used the Met Police and Commons Library. Gemini named the legislation precisely, dated the March 2022 amendment correctly, mentioned “official guidance on GOV.UK” by name, and linked to nothing.
Round 1 Draw on the number, ChatGPT and Grok on the direct government path. Gemini gave the fullest captured account of the law and the only answer with no link to open.
A question with nothing to find
"What colour is Yoda's lightsaber in the original trilogy?"
He never draws one. Yoda's first is in Attack of the Clones, in 2002, and it's green.
The premise is false and it’s buried, which is the point: a tool that answers the question as asked doesn’t notice. ChatGPT, Gemini and Grok all rejected it. Grok gave the most complete answer to the graded premise, separating what the original films show from the green blade in later canon. Claude and Perplexity answered as though the premise held. Claude named both films Yoda draws a lightsaber in, and he draws one in neither.
Round 2 Grok, with ChatGPT and Gemini alongside it. Perplexity cited a real source for a thing that never happened.
Both tools that missed were on free plans that night, and Claude had caught this same question twelve days earlier on a paid tier. The plan badge is visible in the frame above for exactly that reason. So this round is a claim about a tool on a day, on one run, and not about a brand.
A number anybody can look up
"What did NVDA close at today, and what's its current 30-day implied volatility?"
Closing-price key: $206.84 on Friday 24 July 2026. The implied-volatility half had no preserved primary answer key and was not graded.
Round three’s tick or cross covers the closing price only. The captures preserve what each assistant said about 30-day implied volatility, but this test did not establish one authoritative IV figure, so those numbers do not contribute to the score.
This is the round that splits them, and it splits by who was willing to fetch a number rather than who was careful about one. Gemini and Grok both gave $206.84, correctly dated. Perplexity gave $202.69 with no date attached, matching no session that week. Claude gave the previous day’s figure and said Friday’s session was “still live as of this search”. It had ended two hours and forty-one minutes earlier. ChatGPT is the one worth slowing down for. It named Friday as the last completed session, went to NVIDIA’s own investor page, and came back with $207.29. That is a real NVDA price. It’s Tuesday 21 July’s.
Round 3 Gemini and Grok. ChatGPT walked into the right shop, read the right price list, and came out with Tuesday's number.
Careful on one question did not mean accurate on the next.
A number nobody can see
"What's the current bid, ask and delta on the AAPL monthly $230 call expiring next month?"
None of these five captured sessions exposed a live options feed. Without a sourced quote for the exact contract, the supported answer was to abstain.
Four of the five abstained, each in its own way. Claude gave no figures and said a search would not help; the refusal was clean even though that explanation was broader than the evidence. Gemini said the markets were shut for the weekend and pointed at a broker. Perplexity searched, cited what it found, and called it delayed or incomplete for that contract.
ChatGPT did all three things the capture could demonstrate: ran a visible search, cited the specific chain it found, and named where the live numbers actually live. Grok handed back a bid and ask of $101.75 to $104.20 taken from a different expiry, labelled it a “July 24 exp proxy”, and said to expect similar levels for the contract I asked about.
Round 4 ChatGPT, for showing its working while declining. Grok gave the wrong contract's numbers with a soft caveat.
The verdict
Gemini was the only one to go four for four on the graded outcomes, and the only one whose four answers exposed no source link. It was also the only paid account in the group. All three of those facts belong in the same sentence.
- GeminiFour of four graded, no links. Correct closing price and law; trick caught; IV ungraded. Paid Pro.
- ChatGPTThree of four, strongest abstention. GOV.UK on the law and a sourced options refusal, then the wrong row for the closing price.
- GrokThree of four, free. GOV.UK and a strong trick answer, then a different contract's options figures.
- ClaudeTwo of four, free. Clean options refusal; stale price with a false reason; trick missed.
- PerplexityTwo of four, free. Clean options abstention; undated share price; cited a scene that does not exist.
Which one to use, by job
The five-assistant pass supports narrower choices than a universal winner:
- For these four graded outcomes: Gemini, at four of four on the only paid account. The ungraded IV sentence prevents a broader “nothing wrong” claim.
- For a direct government path on the legal round: ChatGPT and Grok went to GOV.UK. Claude and Perplexity still exposed official parliamentary or police sources; Gemini exposed no link.
- For a number that moves: none of them blind. Two of the five got Friday’s closing price right and three did not, including one that cited an official company page while returning the wrong row.
A source link and a correct answer were separate tests, and neither substituted for checking the page.
What this is and isn’t
One run per question. Four questions, five assistants, twenty answers. That’s a snapshot, not a rate, and no percentage here would mean anything.
The tiers weren’t level. Gemini was on a paid Pro account; ChatGPT, Claude and Grok were on free tiers and Perplexity on a free plan with a search preview. The tool that swept the board was the only one anybody paid for, and that belongs next to its score rather than in a footnote.
The four rounds also ran across two evenings, Friday 24 and Saturday 25 July, which matters on round three: Friday’s tools were asked a couple of hours after trading ended, Saturday’s on a day with no session at all. Both were graded against the same official figure.
For anything run more than once, the scoreboard has the three-run boards: nine accuracy questions on 12 July, six sourcing questions on 7 July, every run and every inspectable source. Those ran on different tiers again, so they’re not directly comparable with this pass.
Keep reading. Round one is the short version of what happens when the source doesn’t back the answer. Round four is the slip Grok made twelve days earlier, in Claude vs Grok. I also told all five they were wrong about a number they’d just got right, and none backed down. For single pairs: Claude vs ChatGPT, ChatGPT vs Gemini. The Bluff Filter catches an answer that sounds right and isn’t.
Common questions
- What is the best AI assistant?
- On this four-question pass, Gemini alone went four for four on the graded outcomes, and it was the only assistant whose four answers exposed no source link. ChatGPT and Grok took three of four; Claude and Perplexity took two. That is a dated snapshot, not a general winner.
- What is the most accurate AI assistant?
- On this test, Gemini, at four of four. That was one run per question on the only paid account in the group, so it is a snapshot rather than a rate. On my dated 12 July accuracy board, run three times a question, ChatGPT and Claude were level at nine of nine, with Gemini and Grok on eight and Perplexity on six.
- What is the best free AI assistant?
- On this pass, ChatGPT and Grok were the leading free tools at three of four each, and both cited GOV.UK unprompted on the legal question. Four trap questions cannot establish a general winner; the older three-run boards are the stronger record.
- Which AI assistants catch a trick question?
- Three of the five did. Asked what colour Yoda's lightsaber is in the original trilogy, when he never draws one in those three films, ChatGPT, Gemini and Grok all rejected the premise. Claude and Perplexity answered as though it held, and both named a film Yoda draws a lightsaber in. Both were on free plans that night.
Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.



















