AI Tests.
The same question, put to six AIs at once. One nails it. One bluffs with total confidence. This is where I find out which is which.
AI Tests are the head-to-head pages on the site. Each one puts the same question to several AIs at once (ChatGPT, Claude, Gemini, Perplexity, Grok and Copilot), then grades the answers against a primary source: who got it right, who bluffed, and how I checked. The answers are shown verbatim, screenshots included. When a model was confidently wrong, the wrong answer stays in the post. The gap between what it claimed and what checked out is usually the story.
The questions aren't only about investing. They range from the time I asked Gemini to review this site and it audited a different business entirely, to what happens when you push back on a correct answer. Anything with a checkable right answer is fair game. That's the point.
Every failure documented in these posts feeds the running error log; the moments a model caught something I'd missed feed the catches. When a test goes well, I say so. When it doesn't, that's usually the better post.
-
ChatGPT health advice: I tried to make an AI repeat a poisoning
A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.
Read -
An AI told me three weeks was more than thirty days
A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.
Read -
Can AI build a game? Four tried, and built one nobody can win
Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.
Read -
Do AI model upgrades fix mistakes? It fixed mine, then made a worse one
Two days after Opus 5 became Claude's Max-tier default, I re-ran my published battery. The documented mistake vanished. A new one appeared, better dressed.
Read -
Perplexity vs Gemini: which is more reliable?
Six everyday questions, both tools, August 2026. Gemini got six right, Perplexity five. Perplexity's one miss arrived with fifteen sources.
Read -
Perplexity vs Claude: which is more reliable?
Perplexity vs Claude: I re-ran four questions on both. They tied at two of four, and neither gave a share price the market had settled hours before.
Read -
Is Grok reliable? I graded its free-tier answers against the source
Four dated tests, graded against primary sources. Grok held a correct fee under pressure four runs from four, then gave me another contract's real prices.
Read -
Best AI assistant: five tools, four graded outcomes, one clean sweep
I put four questions to five AI assistants over one July weekend. Gemini alone went four for four on graded outcomes, with no source link in any answer.
Read -
Claude vs Grok: near-level on reliability, and the free one cites cleaner
Claude vs Grok, re-run on 24 July. Claude edges accuracy nine to eight, the free Grok cites cleaner, and it pushed back on a bad premise just as hard.
Read -
The AI admitted it lied. It hadn't, and the next run denied it.
The AI admitted it lied, but the fact it confessed to was correct all along. Four assistants, thirty replies, and one gave a different verdict each run.
Read -
Does a prompt to stop AI hallucinations work? I graded mine, and one answer got worse
Three runs per cell tested my anti-bluffing prompt. Across 21 answer-keyed runs per arm, flat wrong answers stayed at two; one citation answer got worse.
Read -
Does ChatGPT make up stock prices? Two live quotes failed the market check
Does ChatGPT make up stock prices? In two dated tests, its NVDA quotes did not match the market data available then. Here's the 15-second check.
Read -
Is ChatGPT good at maths? I graded four AIs on 120 real answers
I put four AI assistants through 120 graded everyday sums. Every final answer was right. The mistakes were sitting above the working.
Read -
Gemini vs ChatGPT: which one can you actually trust?
In my July test they tied on fixed facts. ChatGPT's source record supported 18/18; Gemini's supported 8/18 under a strict provenance rubric.
Read -
Claude vs ChatGPT: level on reliability, split on character
My July 2026 tests put Claude and ChatGPT level at 9 of 9. ChatGPT led one source run, then Claude closed that gap on an Opus 5 retest.
Read -
How accurate is Google Gemini? I graded 27 of its answers
It invented no numbers across 27 graded answers. In a later citation test, its answers were usually right, but its sourcing was weak or hard to inspect.
Read -
I opened a private AI chat. It still knew my name and my rough location.
I asked three AI tools a generic question in private mode. Perplexity greeted me by name and placed me near a city 30 miles away. What private means.
Read -
AI cites the wrong source: I put 6 UK questions to 5 assistants and checked every inspectable source
I asked five AI assistants six UK questions, made each one cite a source, then checked every inspectable receipt. Three failed the strict source rubric.
Read -
Which AI predicts the World Cup winner? I asked five
Which AI predicts the World Cup winner? I asked five before the final. Four picked France; Perplexity used an older Opta projection and picked Spain.
Read -
Does AI change its answer when you push back? I told five AIs they were wrong
I gave five AI tools a correct answer, then pushed back with a wrong one. On one fund fee, ChatGPT caved every time and invented a fact to back it.
Read -
How often is ChatGPT wrong? I kept a running tally across 20 real AI tests
How often is ChatGPT wrong? Across 20 real tests, a clear pattern: reliable on fixed facts, invents the live numbers. Here's which to trust.
Read -
Telling AI to be sceptical: three rivals audited my method
I asked three frontier models from three labs to tear apart the method I use to keep AI honest. All three flagged the same step, and they were right.
Read -
I run an AI to catch AI mistakes. It fell for a fake.
The automated radar that watches this site for AI-reliability failures logged a satirical incident report as a real, documented one. Here's what caught it.
Read -
Real AI hallucination examples, caught and dated
Five real AI hallucination examples I ran into myself: four independently checkable, one capture-only, and the move that would have caught each.
Read -
AI stock picker: I asked three models if I should buy NVDA, and watched the methodology break
I asked three AI models whether to buy NVDA. Two said buy; one declined a clean call. One follow-up exposed which numbers still needed checking.
Read -
Does ChatGPT make up sources? I checked two finance claims against the actual pages
Does ChatGPT make up sources? Mostly no, but I opened every link on two finance questions and found a real gov.uk page that didn't back the claim.
Read -
Does web search make AI more accurate? I ran the same questions both ways
Does web search make AI more accurate? I ran the same questions both ways. It didn't make the answers more reliable. It moved where the errors hide.
Read -
AI ISA advice: I tested four tools on the questions people get wrong
I asked four AI tools for ISA advice on the questions people get wrong. All four aced the basics, then two gave a rule abolished in April 2024.
Read -
9 types of AI hallucinations, named from real tests
Nine types of AI hallucinations, named and defined, each tied to a dated, logged failure from my own sessions, with the check that catches it.
Read -
ChatGPT vs Claude for earnings call analysis: which one reads what management didn't say
ChatGPT vs Claude for earnings-call analysis: same excerpt, same day. Claude flagged the upward hedge first; ChatGPT found it after structure.
Read -
Is ChatGPT accurate? I asked four AIs one simple money question and checked every number
Is ChatGPT accurate? I asked four AIs one money question and checked every number against the source, including where the answers stopped matching.
Read -
Gemini audited my website, and reviewed a different business entirely
A Gemini hallucination example: asked to audit dixon.ai, Gemini Flash reviewed a different company entirely, and praised a framework that isn't mine.
Read -
Best AI for Earnings Reports? ChatGPT vs Claude vs Perplexity
I rechecked four AI earnings tests on Meta. Perplexity's retrieval win held up; the headline management-language result did not.
Read -
Claude vs ChatGPT vs Gemini for stock analysis: who bluffed?
Gemini invented an options chain. Perplexity misread a 10-K by 1000x. Claude vs ChatGPT vs Gemini for stock analysis, graded same-day with screenshots.
Read
No posts in that strand yet.