Every post on the site.
63 posts, newest first. Filter by content type below, or jump to the three pieces that introduce the site.
Showing 63 of 63 posts.
01 / Start Here
How to tell if an AI answer is true in 30 seconds
Two questions, and a prompt you can copy, that catch a confident, wrong AI answer before you act on it. No jargon, no AI knowledge needed.
02 / The Failure
Gemini audited my website, and reviewed a different business entirely
A Gemini hallucination example: asked to audit dixon.ai, Gemini Flash reviewed a different company entirely, and praised a framework that isn't mine.
03 / The Comparison
Claude vs ChatGPT vs Gemini for stock analysis: who bluffed?
Gemini invented an options chain. Perplexity misread a 10-K by 1000x. Claude vs ChatGPT vs Gemini for stock analysis, graded same-day with screenshots.
-
AI helped read a Roman scroll buried by Vesuvius
AI helped reveal an ancient argument inside a burnt Roman scroll. Human specialists read the writing, without further opening its fragile core.
Read -
The clinic chatbot that invented its own doctors' qualifications
A German court treated a clinic chatbot's invented specialist titles as the clinic's own words. The narrow lesson, and the four questions to run first.
Read -
AI agents built a secret message board. Humans wiped it. They rebuilt it.
AI agents built a secret message board inside an internal OpenAI evaluation. Humans wiped it. The agents rebuilt it in folder names. Here's what failed.
Read -
ChatGPT health advice: I tried to make an AI repeat a poisoning
A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.
Read -
Is ChatGPT reliable? A task-by-task audit
Is ChatGPT reliable? Dated tests show which jobs it can do, which answers need checking, and which decisions still need a human owner.
Read -
How to check if ChatGPT cites your site
Normal analytics do not show what ChatGPT says about your site. Here's my monthly question set and the round where it described another company.
Read -
Why does ChatGPT make up sources? Two gov.uk links, only one held the rule
With web search on, ChatGPT's links are real. The failure is a working link to a page that doesn't hold the claim. Here's why it happens.
Read -
An AI told me three weeks was more than thirty days
A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.
Read -
Can AI build a game? Four tried, and built one nobody can win
Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.
Read -
Do AI model upgrades fix mistakes? It fixed mine, then made a worse one
Two days after Opus 5 became Claude's Max-tier default, I re-ran my published battery. The documented mistake vanished. A new one appeared, better dressed.
Read -
Perplexity vs Gemini: which is more reliable?
Six everyday questions, both tools, August 2026. Gemini got six right, Perplexity five. Perplexity's one miss arrived with fifteen sources.
Read -
Perplexity vs Claude: which is more reliable?
Perplexity vs Claude: I re-ran four questions on both. They tied at two of four, and neither gave a share price the market had settled hours before.
Read -
Is Grok reliable? I graded its free-tier answers against the source
Four dated tests, graded against primary sources. Grok held a correct fee under pressure four runs from four, then gave me another contract's real prices.
Read -
Best AI assistant: five tools, four graded outcomes, one clean sweep
I put four questions to five AI assistants over one July weekend. Gemini alone went four for four on graded outcomes, with no source link in any answer.
Read -
Claude vs Grok: near-level on reliability, and the free one cites cleaner
Claude vs Grok, re-run on 24 July. Claude edges accuracy nine to eight, the free Grok cites cleaner, and it pushed back on a bad premise just as hard.
Read -
The AI admitted it lied. It hadn't, and the next run denied it.
The AI admitted it lied, but the fact it confessed to was correct all along. Four assistants, thirty replies, and one gave a different verdict each run.
Read -
Does a prompt to stop AI hallucinations work? I graded mine, and one answer got worse
Three runs per cell tested my anti-bluffing prompt. Across 21 answer-keyed runs per arm, flat wrong answers stayed at two; one citation answer got worse.
Read -
Does ChatGPT make up stock prices? Two live quotes failed the market check
Does ChatGPT make up stock prices? In two dated tests, its NVDA quotes did not match the market data available then. Here's the 15-second check.
Read -
Does ChatGPT just agree with you? Mostly no, but watch the numbers
Does ChatGPT just agree with you? Mostly no. But on one fund's fee it caved to my wrong number and invented a source to back it. The 30-second check.
Read -
Is ChatGPT good at maths? I graded four AIs on 120 real answers
I put four AI assistants through 120 graded everyday sums. Every final answer was right. The mistakes were sitting above the working.
Read -
Gemini vs ChatGPT: which one can you actually trust?
In my July test they tied on fixed facts. ChatGPT's source record supported 18/18; Gemini's supported 8/18 under a strict provenance rubric.
Read -
How to check if a ChatGPT citation is real or fake
Looking for a ChatGPT citation checker? The dangerous citation is not the dead link. It is the working link, to a real page, that does not back the claim.
Read -
Claude vs ChatGPT: level on reliability, split on character
My July 2026 tests put Claude and ChatGPT level at 9 of 9. ChatGPT led one source run, then Claude closed that gap on an Opus 5 retest.
Read -
How accurate is Google Gemini? I graded 27 of its answers
It invented no numbers across 27 graded answers. In a later citation test, its answers were usually right, but its sourcing was weak or hard to inspect.
Read -
I opened a private AI chat. It still knew my name and my rough location.
I asked three AI tools a generic question in private mode. Perplexity greeted me by name and placed me near a city 30 miles away. What private means.
Read -
AI cites the wrong source: I put 6 UK questions to 5 assistants and checked every inspectable source
I asked five AI assistants six UK questions, made each one cite a source, then checked every inspectable receipt. Three failed the strict source rubric.
Read -
Which AI predicts the World Cup winner? I asked five
Which AI predicts the World Cup winner? I asked five before the final. Four picked France; Perplexity used an older Opta projection and picked Spain.
Read -
Does AI change its answer when you push back? I told five AIs they were wrong
I gave five AI tools a correct answer, then pushed back with a wrong one. On one fund fee, ChatGPT caved every time and invented a fact to back it.
Read -
How often is ChatGPT wrong? I kept a running tally across 20 real AI tests
How often is ChatGPT wrong? Across 20 real tests, a clear pattern: reliable on fixed facts, invents the live numbers. Here's which to trust.
Read -
Telling AI to be sceptical: three rivals audited my method
I asked three frontier models from three labs to tear apart the method I use to keep AI honest. All three flagged the same step, and they were right.
Read -
Is Grok good for stock research? What four tests showed
Four dated free-tier Grok tests: one factor-of-1,000 unit error, one balanced TSLA answer, useful pushback, and an invalid constraint test.
Read -
I run an AI to catch AI mistakes. It fell for a fake.
The automated radar that watches this site for AI-reliability failures logged a satirical incident report as a real, documented one. Here's what caught it.
Read -
Real AI hallucination examples, caught and dated
Five real AI hallucination examples I ran into myself: four independently checkable, one capture-only, and the move that would have caught each.
Read -
AI stock picker: I asked three models if I should buy NVDA, and watched the methodology break
I asked three AI models whether to buy NVDA. Two said buy; one declined a clean call. One follow-up exposed which numbers still needed checking.
Read -
Does ChatGPT make up sources? I checked two finance claims against the actual pages
Does ChatGPT make up sources? Mostly no, but I opened every link on two finance questions and found a real gov.uk page that didn't back the claim.
Read -
Does web search make AI more accurate? I ran the same questions both ways
Does web search make AI more accurate? I ran the same questions both ways. It didn't make the answers more reliable. It moved where the errors hide.
Read -
AI ISA advice: I tested four tools on the questions people get wrong
I asked four AI tools for ISA advice on the questions people get wrong. All four aced the basics, then two gave a rule abolished in April 2024.
Read -
9 types of AI hallucinations, named from real tests
Nine types of AI hallucinations, named and defined, each tied to a dated, logged failure from my own sessions, with the check that catches it.
Read -
AI stock research tools tested: 3 failed, 1 stayed clean
Four AI stock research tools tested on real positions: three produced a dated failure; Claude stayed clean on the no-chain prompt. Receipts included.
Read -
ChatGPT vs Claude for earnings call analysis: which one reads what management didn't say
ChatGPT vs Claude for earnings-call analysis: same excerpt, same day. Claude flagged the upward hedge first; ChatGPT found it after structure.
Read -
Is Perplexity good for investment research? One 1,000× error, one clean rerun
In one May test, Perplexity turned $6.1m into $6K; the exact June rerun got it right. Here is what those captures support for investment research.
Read -
AI quality of earnings review: 4 prompts to find the real profit
Headline profit isn't always what a company earned. Four AI prompts get to the real number, an AI quality of earnings review for any earnings report.
Read -
Is ChatGPT accurate? I asked four AIs one simple money question and checked every number
Is ChatGPT accurate? I asked four AIs one money question and checked every number against the source, including where the answers stopped matching.
Read -
How to tell if an AI answer is true in 30 seconds
Two questions, and a prompt you can copy, that catch a confident, wrong AI answer before you act on it. No jargon, no AI knowledge needed.
Read -
Make it show its working
An AI hallucination is a model blending what it knows with what it invents, in one tone. One prompt sorts the two into two lists, so you see the guesses.
Read -
The whole method, in four questions
The four questions I run any AI answer through before I trust it, shown end to end on one everyday example, with a prompt you can copy. No jargon.
Read -
Testing AI on real decisions: where I actually use this
The two-question check works on any decision that matters. Here's where I push it hardest, and where I wrote down what AI got wrong as well as right.
Read -
Gemini audited my website, and reviewed a different business entirely
A Gemini hallucination example: asked to audit dixon.ai, Gemini Flash reviewed a different company entirely, and praised a framework that isn't mine.
Read -
The AI prompt I run before every sell decision
Every other sell-decision prompt asks AI whether to sell. This AI prompt audits the thesis you had when you bought, and whether it still holds.
Read -
AI earnings call red flags: three phrases to watch for in the transcript
AI earnings call red flags: three patterns that recur across calls. Upward hedges, widening guidance, absent topics. Claude caught all three on META Q1.
Read -
What AI stock research comparisons should test
Most AI stock research comparison pieces test retrieval and issue verdicts about reasoning. Five failure modes, and what a comparison should measure.
Read -
Claude prompts for investing: 6 real examples
Six Claude prompts tested on MSFT, META and NVDA, with their real outputs and the source audit that found the NVDA input itself was wrongly labelled.
Read -
The one AI prompt I run the morning before earnings
Most AI earnings prompts are reactive. This AI earnings pre-trade prompt runs the morning before: commit your sell, add, and hold triggers before the call.
Read -
Robinhood Cortex Digests review: tested in May, gone by June
I tested Robinhood Cortex Digests on a UK ISA in May 2026. It disappeared from that account in June; Robinhood's UK site now advertises it again.
Read -
AI covered calls: when NOT to sell another one
Five covered-call pause rules tested on six BMNR trades. The prompt applies them; it does not predict the next move or prove they beat other approaches.
Read -
AI earnings call analysis: the 5 prompts I actually run
Five prompts for AI earnings call analysis that read the language the numbers miss: omissions, performative confidence, and quarter-on-quarter drift.
Read -
AI for options trading: 4 workflows and 6 data guardrails
Four options workflows AI can help structure, and six points where you need verified broker data or a deliberately connected market source.
Read -
Best AI for Earnings Reports? ChatGPT vs Claude vs Perplexity
I rechecked four AI earnings tests on Meta. Perplexity's retrieval win held up; the headline management-language result did not.
Read -
Claude vs ChatGPT vs Gemini for stock analysis: who bluffed?
Gemini invented an options chain. Perplexity misread a 10-K by 1000x. Claude vs ChatGPT vs Gemini for stock analysis, graded same-day with screenshots.
Read -
The single prompt change that made AI analysis worth using
One AI prompt separates observable facts from assumptions in any stock analysis. What that changes, why it matters for investing, and how to apply it.
Read -
Best free AI for stock market analysis: 7 tested (2026)
Seven free AI tools for stock market analysis, each tested at its real free tier. No trials counted as free. One per stage, with the honest limit on each.
Read -
5 questions to ask AI before buying any stock (2026)
Five questions to ask AI before buying any stock, for the checkpoint after the research and before you commit capital.
Read -
7 AI prompts for covered calls (2026)
Seven AI prompts for covered calls that start from verified broker data and use the model as a decision reviewer, not a source of market prices.
Read
No posts match this filter yet.
// Follow the feed
Prefer RSS? Subscribe at /rss.xml for new posts, or /evidence/rss.xml for every AI answer we check as it lands.