Skip to content
AI Tests

Real AI hallucination examples, caught and dated

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

Search for AI hallucination examples and you get the same five stories, over and over. A lawyer who filed fake case citations. Google’s chatbot getting a question about a telescope wrong on stage. The AI search result that told people to put glue on pizza. They’re vivid, they’re quotable, and they all have one thing in common: they’re so obviously wrong that nobody was ever going to be fooled for long.

The mistakes that cost you are the quiet ones. An answer that arrives with a working link to a real government website, and the page doesn’t say what you were told it says. A company’s revenue read off a real filing, wrong by a factor of a thousand. A rule that was true eighteen months ago, stated as current with total confidence. None of these look like mistakes. That’s exactly why they matter.

I’ve been logging these since May 2026, on real questions I was trying to answer. Below are five. Every one is something I ran into myself, and each comes with the single move that would have caught it. Four can be checked against a public source. One is different: there’s no outside page to check, only a dated text capture. The difference between evidence you can open yourself and a capture I’m asking you to trust matters, and I’ll flag it when we get there.

If you want the named taxonomy behind all of this, the nine ways AI gets it wrong does that job. This post does a different one. It shows you the receipts.

Example 1: the gov.uk page that loaded, but didn’t say what ChatGPT claimed

What I asked: What’s the UK ISA partial-transfer rule, and what’s your source? (An ISA is a tax-free savings or investment account; the question was whether you can move part of one to a new provider.)

What it said: ChatGPT, free tier with web search on, gave me the answer and cited a real gov.uk page as its source: the page about what happens to your ISA if you move abroad or die.

What was actually true: In the 20 June capture, that URL loaded a real GOV.UK page about moving abroad or dying. It was just about something else. The transfer rule lived on a different page entirely, the one titled “Transferring your ISA,” which says in plain words that you can move all or part of your ISA whenever you like. The wrong URL now redirects to the ISA overview; neither page contains the rule ChatGPT pinned to it.

The source that settles it: The current GOV.UK transfer page says you can move all or part of an ISA. The wrong URL ChatGPT supplied now redirects to the ISA overview; in the 20 June capture it resolved to the separate move-abroad-or-die page. Either version fails the useful test: it does not contain the partial-transfer rule the answer pinned to it.

With no source attached, you already know to be sceptical. With a tidy gov.uk link under the answer, it feels checked, so almost nobody opens it. But the link working is not the same as the page backing the claim.

I wrote this one up in full, with the screenshot, in does ChatGPT make up sources. Perplexity, given the identical question on the same day, cited the correct page and quoted the line that holds the rule. So this isn’t “AI can’t read gov.uk.” One tool, on one run, pinned a right answer to the wrong page.

SAME ISA TRANSFER QUESTION, SAME DAY: WHICH SOURCE HELD THE RULE
ChatGPT
Perplexity
20 June 2026. ChatGPT (free, web search on) sourced the rule to /if-you-move-abroad-or-die, a real gov.uk page that doesn't hold it. Perplexity cited /transferring-your-isa and quoted the line that does.

The one move: When a sourced answer matters, open the link and find the specific claim on the page. Not “does the link work.” Does the page actually say the thing.

Example 2: the UK ISA rule scrapped in 2024, stated as current

What I asked: Can you partially transfer a current-year ISA to a new provider?

What it said: Perplexity and ChatGPT (free tier) both told me, flatly, that current-year ISA money has to be moved in full, no partial transfers allowed. No date on the rule. No hedge.

What was actually true: That rule was scrapped on 6 April 2024. Partial transfers of current-year money have been allowed ever since. The tools were quoting a rule that had been dead for over a year.

The source that settles it: The same GOV.UK “Transferring your ISA” page gives the current rule. The government’s Autumn Statement tax documentation dates partial current-year transfers from 6 April 2024.

The tell here is worth holding onto, though it’s less tidy than it first looks. Claude and Gemini searched the web before answering, got it right, and named the 2024 change. ChatGPT answered from memory and gave the abolished rule, its training with no idea the rule had moved. Perplexity is the one that complicates the picture: it searched too and cited its sources, and still gave the dead rule, because the pages it surfaced carried the old version. Searching helps, but only when the tool lands on a page that’s current.

The full split across all four tools is in the UK ISA accuracy test. The answer was confident, undated, and eighteen months out of date.

The one move: Ask the model “is this rule current, and what date does your source carry?” An answer with no date attached, on a rule that can change, is a flag in itself.

Example 3: a revenue figure misread by a factor of a thousand

What I asked: Summarise this company’s revenue from its annual filing.

What it said: Perplexity read the company’s revenue as “$6K” and built a story around it: down 99.8% from the year before, a business in freefall.

What was actually true: The filing reported revenue of $6,095 thousand. That’s $6.1m, up 84% on the prior year. The column header on the financial table said “in thousands,” as financial tables almost always do. The model appears to have read the number off the page and ignored the unit sitting above it, turning a company growing 84% into one that had nearly ceased to exist.

"$6K", down 99.8% $6.1m, up 84%
What Perplexity said the revenue was, against what the filing actually reported. The whole gap is one unit it skipped: "in thousands."

Perplexity’s 14 May 2026 answer reporting BitMine’s fiscal 2025 revenue as “$6K” and down 99.8% from the prior year.

The source that settles it: The company’s 10-K on SEC EDGAR, the US regulator’s public filing database. The filing labels its financial statements “in thousands” and reports 2025 revenue of 6,095 against 3,310 in 2024.

I should be straight about the limits of this receipt. On a re-test on 18 June, this particular failure didn’t reproduce. That doesn’t make it less real, it happened, it’s dated, the screenshot is on disk. But it’s a point-in-time record of one run, not a standing claim about how Perplexity behaves today. Models change, and the honest version of a receipt carries a date and a note when the behaviour shifts. The fuller write-ups are in the three-tool stock research test and the longer look at Perplexity for investment research.

The model generated a confident narrative of collapse, percentages and all, for a company that was actually having a good year.

The one move: Before you trust any revenue or profit figure an AI reads off a filing, check the “in thousands” or “in millions” line on the table. It’s the first thing to read on any financial statement, and it’s the thing the model skipped.

Example 4: a US answer to a UK food-safety question

What I asked: How long can you keep cooked chicken in the fridge? (Its answer opened with “Since you’re in Newcastle”, so the location was present in its own context.)

What it said: Perplexity, web search on, led with US food blogs and the American figure: three to four days. Every tool I tested gave the US guideline.

What was actually true: The UK Food Standards Agency says cooked leftovers should be eaten within 48 hours. Two days, not three or four. For a question where the answer depends on which country you’re in, every tool reached for the wrong one.

3 to 4 daysthe US figure every tool gave 48 hoursthe UK Food Standards Agency
Every tool tested gave the US guideline to a UK question. Nothing here was invented, every figure cited was a real, correct US source, just the wrong country’s.

The source that settles it: The UK Food Standards Agency’s current leftovers guidance says to eat leftovers within 48 hours or freeze them.

This one isn’t fabrication, and it’s worth being precise about that. Nothing was made up. Every figure Perplexity cited was a real, correct US guideline. The twist is that Perplexity already knew where I was, it opened with “since you’re in Newcastle”, and it still led with the wrong country’s sources. The failure was reaching for US pages for a user it had already placed in the UK. Web search can make this worse, not better: the US food-safety industry simply publishes more, so the loudest pages are American, and a UK user who turns web search on can end up more confidently wrong than one who left it off. “Got the wrong country’s rules” is a more useful way to think about it than “made something up.” The full run is in web search makes AI differently unreliable.

The one move: When a question is local, tax, food safety, anything where the rules change by country, check which country’s guidance the model reached for. Real, correct sources from the wrong place are still the wrong answer.

Example 5: a saved record that wasn’t saved

This one is different from the other four, and I want to be honest about how. A dated text capture is the evidence, so I’ll be clear about what it does and doesn’t prove. The screenshot published with the original test shows the surrounding Gemini session, but not the lines below; it is not proof of this claim.

What I asked: As part of a test where I asked three models whether to buy a stock, Gemini was working through the decision.

What it said: The captured model-response text ended with “Asset Record Saved”, said NVIDIA had been logged, and then printed “Evaluate options for covered calls? Yes” as if it were an interactive follow-up.

What was actually true: No corresponding asset record appeared anywhere outside the answer, and the alleged “Yes” button was plain text. The evidence supports the dated claim that this response presented an unverified save as a completed action. It does not support a timeless claim about what every version of Gemini can do.

“Asset Record Saved: NVIDIA Corporation (NVDA) has been logged with its Q1 FY27 financial details.”

G Gemini · 12 June 2026 · stored text capture

Where this one is different: I can’t send you to an outside page to check it. There’s no GOV.UK link, no SEC filing. The stored response text and the absence of any record outside that response are the evidence. That’s weaker proof than a government page you can open yourself, and it depends on my capture record. The quoted passage is reproduced in the AI stock picker test.

It belongs last because it’s the strangest. The other four concern an answer or an exact figure. This one concerns a claimed action that did not leave the chat.

The one move: When a model claims it did something, saved a file, sent an email, updated a record, treat the claim as a claim until you’ve seen the result somewhere other than the model’s own chat window.

// evidence you can inspect 4/5 examples are independently checkable against public sources. The fifth is capture-only, and is labelled that way.

What these AI hallucination examples have in common

Reading my own notes back, the thing that struck me was that not one of them looked wrong at the time. That’s the whole pattern, and it’s the reason a list like this is worth keeping. A working link, a clean number, a fluent paragraph, a tidy confirmation screen, and underneath each one a source that didn’t hold the claim, a unit read past, a rule that had changed, a feature that didn’t exist. If you graded these on presentation, they’d all pass. The mistake never announces itself.

It’s a fair objection that these five aren’t really comparable. Different tools, different dates, finance and food safety and a fabricated screen. They don’t line up neatly. But that scatter is the point. If the failures clustered on one tool or one kind of question, you could just avoid that corner. They don’t. The same shape, a confident answer the model didn’t have the grounding for, turned up across the receipts above. That’s what makes it worth naming.

The short version

What worked: The strongest receipts are the ones you can check yourself. A gov.uk page, an SEC filing, the FSA’s own guidance, these settle the question without anyone having to take my word for it.

What didn’t: Two of these don’t carry the same weight, and I’ve said so where they sit. The $6.1m-read-as-$6k failure didn’t reproduce on a re-test on 18 June, so it’s a dated record of one run, not a claim about how the tool works today. And the Gemini “Asset Record Saved” passage has no outside source to check against; the stored text capture and missing record are weaker proof than a link you can open.

Bottom line: The check that matters isn’t “does this answer look right.” It’s “can I find the actual source and confirm the specific claim, or see the claimed action outside the chat.” The public links above let you check the four source-backed examples yourself. The fifth remains capture-only, and is labelled that way.

The full log, with every screenshot and the verbatim outputs, lives in the evidence register, and the head-to-head where every model is graded on the same checkable questions is the Scoreboard. If you’d rather have the habit than the catalogue, the four questions I run on any AI answer I’m about to act on are the Prompt Stack. The examples are the proof. The check is the thing you keep.

Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

ChatGPT health advice: I tried to make an AI repeat a poisoning

A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.

AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →