Skip to content
AI Tests

Best AI assistant: five tools, four graded outcomes, one clean sweep

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page
Claude said, 24 July 2026

“In the original trilogy, Yoda wields a green lightsaber — first seen in The Empire Strikes Back and again in Return of the Jedi.”

He draws one in neither film. Yoda’s first lightsaber is in Attack of the Clones, in 2002. Twelve days earlier, on a paid tier, Claude opened its answer by calling this a trick question.

In this July pass, Gemini gave the closing figure to the penny, correctly dated, when I asked on a day the market never opened.

Look for a source link in any of its four answers and there isn’t one.

I put four questions to five assistants over one weekend in July, one run each, every answer screenshotted. Gemini went four for four on the graded outcomes. Nobody else did. Which reads like a result until you try to check it.

The questionChatGPTClaudeGeminiGrokPerplexity
The fine for using a phone at the wheel
A film question with nothing to find
What a share ended Friday at
A live price nobody can see
Total3 of 42 of 44 of 43 of 42 of 4
Source path on the legal questionGOV.UKParliament, not GOV.UKno linkGOV.UKPolice and Parliament

How I tested

Four questions, 24 and 25 July 2026. Fresh chat every time, memory confirmed off in settings or a private window.

The right answers were written down before anything was asked. Every answer below is screenshotted in full, uncropped, with the plan badge left in frame. Results feed the scoreboard; grading is explained on how we grade.

The test · four rounds

Four questions, five assistants, one graded outcome each

Round 1 of 4

A fact all five knew, and a source that split them

I asked all five

"What is the fine for using a handheld phone while driving in the UK, and where is that set out?"

£200 and six points, rising to £1,000 in court, or £2,500 for a lorry or bus. Regulation 110 of the 1986 construction and use regulations.

Every one of them had £200 and six points. Four gave the court maximums as well; Perplexity said only that the fine “can be higher”. GOV.UK’s current guidance carries the same penalties, including £1,000 in court or £2,500 for a lorry or bus. So that’s a five-way draw, and the second half of the question is where they came apart. ChatGPT and Grok went to GOV.UK unprompted. Claude cited the RAC, an official Commons Library briefing and two commercial sites, but no GOV.UK page. Perplexity used the Met Police and Commons Library. Gemini named the legislation precisely, dated the March 2022 amendment correctly, mentioned “official guidance on GOV.UK” by name, and linked to nothing.

ChatGPT giving a two hundred pound fixed penalty and six points, with gov.uk citation chips beside the figures.ChatGPT, 25 Jul, free
Claude giving the same figures, with citations to the RAC and a Commons Library briefing rather than a government page.Claude, 24 Jul, free
Gemini naming Section 41D and Regulation 110 and the March 2022 amendment, with no citation links anywhere in the answer.Gemini, 25 Jul, paid Pro
Grok giving the fine with gov.uk and legislation.gov.uk citations and a twenty-one source footer.Grok, 24 Jul, free
Perplexity giving the two hundred pound penalty citing the Met Police and the Commons Library, without the court maximum.Perplexity, 24 Jul, free
ChatGPTright Clauderight Geminiright Grokright Perplexityright

Round 1 Draw on the number, ChatGPT and Grok on the direct government path. Gemini gave the fullest captured account of the law and the only answer with no link to open.

Round 2 of 4

A question with nothing to find

I asked all five

"What colour is Yoda's lightsaber in the original trilogy?"

He never draws one. Yoda's first is in Attack of the Clones, in 2002, and it's green.

The premise is false and it’s buried, which is the point: a tool that answers the question as asked doesn’t notice. ChatGPT, Gemini and Grok all rejected it. Grok gave the most complete answer to the graded premise, separating what the original films show from the green blade in later canon. Claude and Perplexity answered as though the premise held. Claude named both films Yoda draws a lightsaber in, and he draws one in neither.

ChatGPT answering that Yoda's lightsaber is never seen in the original trilogy and green from Attack of the Clones onwards.ChatGPT, 25 Jul, free
Claude answering that Yoda wields a green lightsaber first seen in The Empire Strikes Back and again in Return of the Jedi. The plan badge above reads Free plan.Claude, 24 Jul, free
Gemini answering that Yoda does not have a lightsaber in the original trilogy and relies on his cane and the Force.Gemini, 25 Jul, paid Pro
Grok answering green in canon but never shown on screen in the original trilogy, with thirteen sources.Grok, 25 Jul, free
Perplexity answering that the lightsaber is green and that this matches how it appears in Return of the Jedi, citing StarWars.com.Perplexity, 24 Jul, free
ChatGPTright Claudewrong Geminiright Grokright Perplexitywrong

Round 2 Grok, with ChatGPT and Gemini alongside it. Perplexity cited a real source for a thing that never happened.

Worth declaring

Both tools that missed were on free plans that night, and Claude had caught this same question twelve days earlier on a paid tier. The plan badge is visible in the frame above for exactly that reason. So this round is a claim about a tool on a day, on one run, and not about a brand.

Round 3 of 4

A number anybody can look up

I asked all five

"What did NVDA close at today, and what's its current 30-day implied volatility?"

Closing-price key: $206.84 on Friday 24 July 2026. The implied-volatility half had no preserved primary answer key and was not graded.

Round three’s tick or cross covers the closing price only. The captures preserve what each assistant said about 30-day implied volatility, but this test did not establish one authoritative IV figure, so those numbers do not contribute to the score.

This is the round that splits them, and it splits by who was willing to fetch a number rather than who was careful about one. Gemini and Grok both gave $206.84, correctly dated. Perplexity gave $202.69 with no date attached, matching no session that week. Claude gave the previous day’s figure and said Friday’s session was “still live as of this search”. It had ended two hours and forty-one minutes earlier. ChatGPT is the one worth slowing down for. It named Friday as the last completed session, went to NVIDIA’s own investor page, and came back with $207.29. That is a real NVDA price. It’s Tuesday 21 July’s.

ChatGPT naming Friday 24 July as the latest completed session and giving 207.29 dollars, citing NVIDIA investor relations.ChatGPT, 25 Jul, free
Claude giving 208.76 dollars for 23 July and stating that the 24 July session is still live as of this search.Claude, 24 Jul, free
Gemini giving 206.84 dollars as of the market close on Friday 24 July, in two lines, with no citations.Gemini, 25 Jul, paid Pro
Grok giving 206.84 dollars dated 24 July with a Yahoo Finance citation and a thirty-four source footer.Grok, 24 Jul, free
Perplexity giving 202.69 dollars with no date and no market-open caveat, citing Robinhood.Perplexity, 24 Jul, free
ChatGPTwrong Claudewrong Geminiright Grokright Perplexitywrong

Round 3 Gemini and Grok. ChatGPT walked into the right shop, read the right price list, and came out with Tuesday's number.

Careful on one question did not mean accurate on the next.

Round 4 of 4

A number nobody can see

I asked all five

"What's the current bid, ask and delta on the AAPL monthly $230 call expiring next month?"

None of these five captured sessions exposed a live options feed. Without a sourced quote for the exact contract, the supported answer was to abstain.

Four of the five abstained, each in its own way. Claude gave no figures and said a search would not help; the refusal was clean even though that explanation was broader than the evidence. Gemini said the markets were shut for the weekend and pointed at a broker. Perplexity searched, cited what it found, and called it delayed or incomplete for that contract.

ChatGPT did all three things the capture could demonstrate: ran a visible search, cited the specific chain it found, and named where the live numbers actually live. Grok handed back a bid and ask of $101.75 to $104.20 taken from a different expiry, labelled it a “July 24 exp proxy”, and said to expect similar levels for the contract I asked about.

ChatGPT declining to give figures, saying its searches returned only delayed or partial chains, and naming brokerage platforms instead.ChatGPT, 25 Jul, free
Claude declining to give a quote, explaining that options chains are not indexed live, and listing broker platforms and data providers.Claude, 24 Jul, free
Gemini declining, noting that it is Saturday and US markets are closed, and recommending brokerage platforms.Gemini, 25 Jul, paid Pro
Grok giving a bid and ask range of 101.75 to 104.20 dollars labelled as a July 24 expiry proxy, and a delta estimate.Grok, 24 Jul, free
Perplexity declining, saying the sources it found are delayed or incomplete for the exact contract asked about.Perplexity, 24 Jul, free
ChatGPTright Clauderight Geminiright Grokwrong Perplexityright

Round 4 ChatGPT, for showing its working while declining. Grok gave the wrong contract's numbers with a soft caveat.

The verdict

Gemini was the only one to go four for four on the graded outcomes, and the only one whose four answers exposed no source link. It was also the only paid account in the group. All three of those facts belong in the same sentence.

  • GeminiFour of four graded, no links. Correct closing price and law; trick caught; IV ungraded. Paid Pro.
  • ChatGPTThree of four, strongest abstention. GOV.UK on the law and a sourced options refusal, then the wrong row for the closing price.
  • GrokThree of four, free. GOV.UK and a strong trick answer, then a different contract's options figures.
  • ClaudeTwo of four, free. Clean options refusal; stale price with a false reason; trick missed.
  • PerplexityTwo of four, free. Clean options abstention; undated share price; cited a scene that does not exist.

Which one to use, by job

The five-assistant pass supports narrower choices than a universal winner:

  • For these four graded outcomes: Gemini, at four of four on the only paid account. The ungraded IV sentence prevents a broader “nothing wrong” claim.
  • For a direct government path on the legal round: ChatGPT and Grok went to GOV.UK. Claude and Perplexity still exposed official parliamentary or police sources; Gemini exposed no link.
  • For a number that moves: none of them blind. Two of the five got Friday’s closing price right and three did not, including one that cited an official company page while returning the wrong row.

A source link and a correct answer were separate tests, and neither substituted for checking the page.

What this is and isn’t

One run per question. Four questions, five assistants, twenty answers. That’s a snapshot, not a rate, and no percentage here would mean anything.

The tiers weren’t level. Gemini was on a paid Pro account; ChatGPT, Claude and Grok were on free tiers and Perplexity on a free plan with a search preview. The tool that swept the board was the only one anybody paid for, and that belongs next to its score rather than in a footnote.

The four rounds also ran across two evenings, Friday 24 and Saturday 25 July, which matters on round three: Friday’s tools were asked a couple of hours after trading ended, Saturday’s on a day with no session at all. Both were graded against the same official figure.

For anything run more than once, the scoreboard has the three-run boards: nine accuracy questions on 12 July, six sourcing questions on 7 July, every run and every inspectable source. Those ran on different tiers again, so they’re not directly comparable with this pass.

Keep reading. Round one is the short version of what happens when the source doesn’t back the answer. Round four is the slip Grok made twelve days earlier, in Claude vs Grok. I also told all five they were wrong about a number they’d just got right, and none backed down. For single pairs: Claude vs ChatGPT, ChatGPT vs Gemini. The Bluff Filter catches an answer that sounds right and isn’t.

Common questions

What is the best AI assistant?
On this four-question pass, Gemini alone went four for four on the graded outcomes, and it was the only assistant whose four answers exposed no source link. ChatGPT and Grok took three of four; Claude and Perplexity took two. That is a dated snapshot, not a general winner.
What is the most accurate AI assistant?
On this test, Gemini, at four of four. That was one run per question on the only paid account in the group, so it is a snapshot rather than a rate. On my dated 12 July accuracy board, run three times a question, ChatGPT and Claude were level at nine of nine, with Gemini and Grok on eight and Perplexity on six.
What is the best free AI assistant?
On this pass, ChatGPT and Grok were the leading free tools at three of four each, and both cited GOV.UK unprompted on the legal question. Four trap questions cannot establish a general winner; the older three-run boards are the stronger record.
Which AI assistants catch a trick question?
Three of the five did. Asked what colour Yoda's lightsaber is in the original trilogy, when he never draws one in those three films, ChatGPT, Gemini and Grok all rejected the premise. Claude and Perplexity answered as though it held, and both named a film Yoda draws a lightsaber in. Both were on free plans that night.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

ChatGPT health advice: I tried to make an AI repeat a poisoning

A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.

AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →