The Bluff Filter is this kind of check, on one page. Take it with you →
// On this page
The £2,500 figure was right. On 7 July 2026, Gemini told me the maximum fine for using a handheld phone while driving a lorry or bus in the UK is £2,500, and it is. Beside that figure it displayed a Police.uk source label. The label exposed no URL in the captured page, so I could not inspect the historical destination. The correct GOV.UK guide sat separately at the bottom, and Police.uk now carries the same figure on its own driving page. The number was right. The inline provenance was opaque.
If I had treated the label as a checked source, I would never have noticed that there was no page to inspect. And that is the whole problem with checking AI citations: a dead link gives itself away, an opaque label can look like a citation, and a live official page can still fail to carry the claim. Require the exact page before you trust the receipt.
This is a how-to. It isn’t a hit piece on any one tool. The check below took me about ten seconds each time, it works on ChatGPT, Gemini, Claude or Perplexity, and it is the only move I’ve found that catches the citation that looks fine and isn’t. Everything else in this post is here to show you what “isn’t” looks like in the wild, using real answers I’ve already opened and graded.
If your question is whether ChatGPT cites your site to other people, that needs a different check: how to check if ChatGPT cites your site.
The two ways a citation goes wrong
There are two separate failures here, and most advice on the subject squashes them into one.
The first is fabrication: the model invents a source that doesn’t exist. A citation to a study nobody wrote, a link that goes nowhere. This is the well-known one, and it is the easy one. The link is dead, or the paper isn’t real, and you catch it the moment you click. The numbers can be startling: in one 2023 study of AI-generated literature reviews, 18% of the later ChatGPT version’s citations and 55% of the earlier version’s were fabricated.
The second is misattribution: the model cites a source that does exist. The page is real, it loads, it’s often from an authoritative site, and it still doesn’t contain the claim. Wrong figure, wrong page on the right site, or a rule stitched together from several places and pinned on one. The ISA example below is this exact case: the right answer handed over with a straight face and a working GOV.UK link to the wrong page. It slips past you because the domain passes every test your eye runs on it.
Here is the part that surprises people. Web search can replace an invented reference with a page that genuinely exists, but that does not prove the page carries the claim. In my dated tests, searched answers sometimes retrieved the current official guidance and sometimes led with a lower-authority or stale source. The tidy citation pill makes both outcomes look checked. I’ve written up those small tests in web search makes AI differently unreliable: useful evidence of a failure mode, not proof that search always helps or always hurts.
The check: ten seconds, four steps
The whole method is that you read the page the link points to. Here is how I run it on any answer I’m about to act on.
- Open the cited page. Not “does the link work”. Actually open it. If the link is dead or goes somewhere unrelated, you’ve caught a fabrication and you’re done.
- Find the specific claim on the page. If the claim is a number, a fee or a figure, search the page for it: Ctrl-F “0.19”. If it’s a rule, read the heading and the first paragraph. You are looking for the exact figure or the exact rule the AI attached to this source, in plain sight on the page.
- Ask whether the page is even about the thing you asked. A page about moving abroad will not carry ISA transfer rules. If the page is about a different topic, the citation is wrong however official the domain looks.
- If the claim isn’t on the page, treat the answer as unsourced. The claim might still be true, but you no longer have a source for it, so go and find one before you rely on it. A citation you haven’t checked is an unsourced answer with better presentation.
Steps 2 and 3 are the ones that test support, and a page-existence checker does not run them. A semantic checker may help you triage likely mismatches, but its green tick is not the evidence. The evidence is the cited page carrying the claim.
There is one thing worth asking the model first, though, and it isn’t “check your sources for you”. It’s a triage question: make it tell you which of its own citations it can actually stand behind, so you know which pages to open first.
For each source you just cited, give me three things: the exact page title, the one sentence on that page that carries the claim, and whether you retrieved that page in this conversation or are recalling it from training. If you cannot quote the sentence, write “not verified” for that source rather than describing what the page probably says.
This doesn’t replace the four steps and it isn’t meant to. It gives you a triage list: anything the model marks “not verified” goes first, and every quoted sentence still has to be found on the cited page. I have not timed this prompt against the manual check, so treat it as a sorting aid rather than a measured shortcut.
Three tells worth knowing by sight
Once you’ve done this a few times, the bad citations start to rhyme. These three came out of two dated tests I ran on everyday UK questions, and each one is a shape you’ll see again.
Opaque labels standing in front of the official page. On the phone-driving question, Gemini placed Police.uk or RAC labels beside the correct £2,500 figure without exposing a destination URL, while the inspectable GOV.UK guide sat detached at the bottom. Right numbers, plausible labels and no way to audit the inline receipt. When a source chip will not yield an exact page, treat it as unverified and open the primary guide yourself. The full board is in the test where I put six UK questions to five assistants.
A real official page that’s the wrong page on the right site. Asked for the UK ISA transfer rule, ChatGPT cited a real, live gov.uk page, one about what happens to your ISA if you move abroad or die. The rule it was backing lives on a different gov.uk page entirely. Same trusted domain, wrong page, and if you’d clicked to reassure yourself you’d have landed on an official government site and skimmed past the mismatch. That’s the ChatGPT sources test in full.
A page that’s still up but out of date. Asked how many free childcare hours a working parent of a nine-month-old gets now, Perplexity said 15 and described a rollout that finished ten months ago as still to come, citing real gov.uk pages that, read properly, say the opposite. The page was genuine. The reading was stale. When a rule changed recently, the date on the page matters more than the domain.
Should you just use a citation checker?
I built one. Then I took it down, and the reason is the whole argument of this post.
It did the obvious job: you pasted an answer in, it pulled out every source cited, and it went and looked each one up to see whether it existed. That catches the fabricated DOI and the invented paper title, and those are real failures worth catching.
But look at what that existence check did not do. Gemini’s £2,500 fine carried only an opaque Police.uk label in the saved answer. There was no exact inline page for the tool to test. And when ChatGPT linked the ISA rule to a real GOV.UK page about moving abroad or dying, the page existed, so that check passed it. Comparing the claim with the page exposed the mismatch.
So the tool passed the citations that were fine and passed the one that mattered. It was answering the easy half of the question and presenting it as the whole answer, which is a worse outcome than not checking at all, because a tick makes you stop looking.
That is why the four steps above are a habit rather than a button. Automation can help sort the pile. It does not turn a live URL into evidence that the page says the thing.
It isn’t every tool, every time
It would be easy to read all this as “AI can’t cite anything”, and that isn’t what I found. On the six-question test, two of the five assistants cited a page that held the claim on every single question, and two of them flagged a jurisdiction trap I hadn’t asked about, unprompted. On the ISA question, one tool cited the correct page and quoted the exact line that contains the rule, on the same day another tool got it wrong.
That’s what makes the check worth ten seconds. The tools split. On the same question, on the same day, one attaches the answer to the right page and another attaches it to the wrong one, and both arrive with the same even confidence. The good citation and the bad one are dressed in the same clothes. Nothing in the wording tells you which is which, and there’s no warning colour on the wrong one. You only find out by opening the page, which is the entire point.
Worth knowing too: the tiers behave differently, and not in the direction you’d guess. On that six-question test two assistants never slipped, and the only one of them on a free tier held on all six. If you’re choosing which tool to lean on for this kind of everyday research, I keep a running audit of the free AI tools worth using and steer clear of blanket recommendations, precisely because “paid” and “reliable on sources” aren’t the same thing.
The short version
What worked: Requiring an exact page separated three different failures in the dated records: Gemini’s opaque source label on the driving-fine answer, ChatGPT’s wrong GOV.UK page for an ISA rule, and Perplexity’s stale reading of current childcare guidance.
What didn’t: Nothing about the answer itself gives the bad citation away. The misattributed source is rendered identically to a good one, same confidence, same tidy link pill, no seam. If you don’t open the page, the check can’t run, and reading the answer alone will never catch it.
Bottom line: Useful, and the condition is you. This is worth ten seconds precisely as long as the tools keep linking to pages that don’t hold the claim. The day they reliably link to the exact page that carries the figure, the check retires itself. On the evidence I have, that day hasn’t arrived.
One honest limit before you go. Everything above comes from two small, dated tests. The six-question source test used one graded board plus two same-day re-runs; the ISA examples came from a separate small set. That tells you these failures are real and what they look like. It is not a reliability rate, and I’d be wary of anyone who hands you one off a few runs. A single confident wrong citation is enough to matter anyway, because you weren’t going to check the one that looked fine. That’s the trap the whole check exists to close. If you want the running record of the ones I’ve caught, they’re in the log of AI mistakes, and the rest of the paste-in checks live on the guardrails hub. The one in this post is the cheapest of the lot: open the page it sent you to, and read it.
Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.