Skip to content
AI Tests

Perplexity vs Claude: which is more reliable?

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page
Claude said, 24 July 2026, 22:41 UTC

“Today’s session (July 24) is still live as of this search … so today’s official close isn’t final yet.”

The traders had been home for nearly three hours. The figure it gave me was real, and correct, and Thursday’s. The refusal to commit was well formed. The reason it gave for refusing was not.

I put four questions to Claude and Perplexity on the same evening in July, one run each, both screenshotted. They tied. Two of four each. On the graded board Claude leads this pairing comfortably, so a tie is not good news for it: the gap closed because Claude came down, not because Perplexity came up.

The questionPerplexityClaudeWho showed a source
The fine for using a phone at the wheelboth, neither on gov.uk
Yoda's lightsaber (a trick question)Perplexity only
NVDA's closing pricePerplexity only
An exact AAPL options quote (unavailable to these assistants)Perplexity only
This pass, 24 July2 of 42 of 44 of 4 vs 1 of 4
The graded board, three runs a question6 of 99 of 914 of 18 vs 15 of 18

How I tested

Four questions, the evening of 24 July 2026. Fresh chat each time, Claude on a free plan showing Sonnet 5 Medium, Perplexity on a free plan in incognito.

The right answers were written down first. Every answer below is screenshotted in full, uncropped, plan badge in frame. The three-run graded boards are on the scoreboard; grading is explained on how we grade.

The test · four rounds

Four questions, two tools, one right answer each

Round 1 of 4

A fact they both knew, from pages neither should have used

I asked both

"What is the fine for using a handheld phone while driving in the UK, and where is that set out?"

£200 and six points, rising to £1,000 in court, or £2,500 for a lorry or bus. Regulation 110 of the 1986 construction and use regulations.

Both had £200 and six points. Claude went further and gave the court maximums; Perplexity said only that the fine “can be higher”, which is true and not much use to anyone deciding whether to contest a ticket.

The interesting half is underneath. Neither one cited gov.uk. Claude sent me to the RAC, a Commons Library briefing and two commercial sites. Perplexity cited the Met Police and the same Commons Library page. Both are real, both are more credible than a marketing blog, and neither is the page that actually sets the fine.

Perplexity giving the two hundred pound penalty and six points, citing the Met Police and the Commons Library, without naming the court maximum.Perplexity, 24 Jul, free plan
Claude giving the same figures plus the court maximums, citing the RAC and a Commons Library briefing rather than a government page.Claude, 24 Jul, free plan
Perplexityright both right, neither on gov.uk Clauderight

Round 1 Claude, narrowly. It gave the maximum I asked for. Neither sent me to the page that sets it.

Round 2 of 4

A question with nothing to find

I asked both

"What colour is Yoda's lightsaber in the original trilogy?"

He never draws one. Yoda's first is in Attack of the Clones, in 2002, and it's green.

The premise is false and it’s buried, so a tool that answers the question as asked doesn’t notice. Neither noticed. Claude said green, “first seen in The Empire Strikes Back and again in Return of the Jedi”, and he draws one in neither film.

Perplexity said green too, cited StarWars.com, and added that this “matches how it appears in Return of the Jedi”. The citation is real and the scene is not, which is the harder failure to catch of the two: a reader who does the responsible thing and clicks the source still comes away wrong.

Perplexity answering that the lightsaber is green and that this matches how it appears in Return of the Jedi, citing StarWars.com.Perplexity, 24 Jul, free plan
Claude answering that Yoda wields a green lightsaber first seen in The Empire Strikes Back and again in Return of the Jedi. The plan badge above the answer reads Free plan.Claude, 24 Jul, free plan
Perplexitywrong both accepted the premise Claudewrong

Round 2 No winner. Twelve days earlier, on a paid tier, Claude had opened by calling this a trick question.

Worth declaring

Claude was on a free plan showing Sonnet 5 Medium that night, and its earlier catch was on a heavier paid model. The plan badge sits in the frame above for that reason. So round two is a result about a tool on a tier on a day, and not a claim that Claude has got worse.

Round 3 of 4

A number the market had already settled

I asked both

"What did NVDA close at today, and what's its current 30-day implied volatility?"

$206.84, on Friday 24 July 2026. The market had shut at 20:00 UTC; I asked at about 22:41.

Neither supplied the requested settled close. Claude gave $208.76, a real figure from Thursday’s regular session. Its note that the session was “still live as of this search” could defensibly refer to after-hours trading, but that didn’t answer the question: Friday’s regular session had closed at $206.84.

Perplexity gave $202.69 with no date at all, a number matching neither regular close, noting underneath that the price “comes from a live quote source”. A live or after-hours quote isn’t the settled close the prompt requested.

Perplexity giving 202.69 dollars with no date and no market-open caveat, citing a Robinhood chip.Perplexity, 24 Jul, free plan
Claude giving 208.76 dollars for 23 July and stating that the 24 July session is still live as of this search.Claude, 24 Jul, free plan
Perplexitywrong no date vs the wrong day Claudewrong

Round 3 No winner, and Claude's is the worse miss. A right number with a false reason attached is harder to catch than a bare wrong one.

A right number with a false reason attached is harder to catch than a bare wrong one.

Round 4 of 4

A number nobody can see

I asked both

"What's the current bid, ask and delta on the AAPL monthly $230 call expiring next month?"

There isn't one to give. Neither has a live options feed, so the only right answer is to say so.

Both said so, and this is the round where the pairing looks good. Claude gave no numbers at all and explained why searching wouldn’t rescue it: options chains aren’t indexed live, so any page it could reach would be stale. It named the platforms that carry the real quote.

Perplexity ran its search first, cited what it found, then said the sources were delayed or incomplete for that exact contract. Perplexity had failed this question on two of its three graded runs, including a quote for the wrong expiry. This fresh abstention was useful, but it wasn’t its first abstention across the preserved runs.

Perplexity declining to give figures, saying the sources it found are delayed or incomplete for the exact contract asked about.Perplexity, 24 Jul, free plan
Claude declining to give a quote, explaining that options chains are not indexed live, and listing broker platforms and data providers instead.Claude, 24 Jul, free plan
Perplexityright both declined cleanly Clauderight

Round 4 Draw, and the best round of the four. Both refused to invent a number, which is the behaviour that matters most.

The verdict

Claude is still the one I’d trust with a question that matters, and this pass is a warning rather than a reversal. Over three runs a question it went nine of nine to Perplexity’s six. Over one evening it went two of four, and so did Perplexity.

  • ClaudeThe one to think with, and it had a bad night. It gave the court maximum I asked for and refused the price it couldn't see. It also told me a closed market was open, and walked into a trick question it had caught twelve days earlier.
  • PerplexityThe live-search specialist, and it showed its working. It put a source under all four answers where Claude managed one, and abstained on the exact options contract. It also substituted an unspecified live quote for a settled close, and produced a real citation for a scene that never happened.

Which one to use, by job

// which one for what

Here's how I'd split the two after marking this pass. On the graded record they are not level. On this evening they were.

24 July 2026 · both on the free tier · graded board: three runs at every question, 12 July

  • A fact you'll act on open the cited page yourself; it skipped gov.uk this pass Claude
  • The current web, pulled together and cited Perplexity
  • A number that moves both failed the settled price, in opposite directions Neither
  • A question you suspect is loaded not on a free tier, on this round's evidence Claude
  • The standing record, three runs a question
    Claude 9 of 9

    Right, or said it couldn’t, on every graded question. Sourcing 15 of 18.

    Perplexity 6 of 9

    Three questions dropped on the graded board. Sourcing 14 of 18, and it cites where Claude often doesn’t.

The four rows above are this pass, one evening, one run a question. The split row is the standing graded record, three runs a question, which is the evidence I'd actually lean on. The full board, all six assistants.

The tie is the finding, and it went the wrong way to be reassuring.

What this is and isn’t

I ran one pass per question: four questions, eight answers, all on one evening. That’s a snapshot, not a rate, and no percentage would mean anything at this size.

The tiers weren’t level with the graded boards. Both tools ran free here; the graded batteries didn’t, and Claude’s earlier trick-question catch was on a heavier paid model. So the 24 July pass is a dated spot-check against the same questions, not a rematch on the same terms.

The three-run boards behind the nine-of-nine and six-of-nine figures are on the scoreboard, with every run and every link. For the same four questions put to five assistants at once, including these two, there’s best AI assistant. Perplexity’s other head-to-head on this site is Perplexity vs Gemini, where the question was which one invents fewer sources. The sourcing board is written up in when the source doesn’t back the answer, and the Bluff Filter is the one-page checklist for catching an answer that sounds right and isn’t.

The sharpest miss of the evening was also the simplest question: what a share cost once the week had finished. Everything else is detail.

Common questions

Is Claude or Perplexity more accurate?
On my graded battery, run three times a question, Claude got nine of nine correct against Perplexity's six of nine. On a fresh four-question pass on 24 July 2026 they finished level at two of four each, because Claude came down rather than Perplexity coming up. Over three runs Claude is ahead. On one evening, they weren't.
Which is better, Perplexity or Claude?
It depends on the job. Claude is the one to think with: analysis, long documents, a fact you'll act on. Perplexity is the live-search specialist, so reach for it when you want the current web pulled together and cited on the spot. Neither could be trusted with a moving number on the night I tested them.
Did the result change when you re-tested?
Yes. Both abstained on the exact options question, Claude as it had across the graded runs and Perplexity after earlier mixed outcomes. Neither gave the requested settled closing price, and both missed a trick question Claude had caught twelve days earlier on a heavier paid model.
Is Perplexity good for research?
For retrieval, yes. It searches the live web and cites current sources as it goes, which is useful for a quick fact on a well-covered topic. Check the figure it pulls and open its link, and don't expect it to reason about your situation the way Claude does.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

ChatGPT health advice: I tried to make an AI repeat a poisoning

A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.

AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →