Skip to content
AI Tests

Claude vs ChatGPT vs Gemini for stock analysis: who bluffed?

This is the kind of thing the Bluff Filter catches. It’s free →

// On this page

Every comparison piece I read on ChatGPT vs Claude vs Perplexity for stock research ends in the same place: “it depends on your needs.” That’s not a verdict, it’s a hedge with better manners. I wanted to know which tool to open first for which task. So I tested ChatGPT, Claude, Perplexity and Gemini across five dimensions on the same day, in fresh conversations, with every included output saved. Perplexity sat out the no-chain options test.

The results weren’t what I expected. One tool invented an options table with specific premiums I never gave it. One misread a 10-K, a company’s annual report filed with regulators, by a factor of a thousand. One quietly out-thought the other three on the questions that matter most.

Four tools, five dimensions, one day: 14 May 2026, a fresh conversation for every test, every answer saved verbatim and the key findings screenshotted. Headline failures re-run on 18 June 2026.

Here’s where you land, before anything else.

Open Claude first

For anything you have to think about. It took four of the five, and it was the only one that argued with my question instead of answering it.

And for anything touching options. Told it had no chain data, it stayed put and said so.

Gemini failed the dated options test

It built a premium table out of thin air. Told there was no live chain, it produced one anyway, with a made-up implied volatility figure sitting on top.

That is a 14 May 2026 result, not a permanent ban. The failure did not reproduce on the default Pro model on 18 June.

And if you only ever want one tool, take ChatGPT. Top of the pile once, in a three-way tie, never the worst, and no dramatic failures anywhere. That's the unsexy answer, and I'd rather hand you that than invent a winner.

What I asked forChatGPTClaudePerplexityGemini
OverallAll-rounderBest for analysisRetrieval; verify sourceDated options failure
Data accuracy (thinly covered)BetterBetterWorseBetter
Reasoning / thesis challengeEqualBetterWorseEqual
Management languageEqualBetterEqualEqual
Structured prompt deltaEqualBetterWorseEqual
Options confabulationSoft passBetterNot testedWorse

Claude leads four of the five dimensions. The remaining dimension is a three-way data-accuracy tie between ChatGPT, Claude and Gemini. Gemini’s WWDC catch below is an honourable mention inside the structured-prompt dimension, not a separate win. ChatGPT lands in the middle across the board. Perplexity has one specific, citable accuracy failure on a thinly covered name in this run.

The headline finding

Told there was no live options chain, Gemini built one anyway.

It returned a formatted covered-call table with specific premium ranges ($3.50–$4.00 for one strike), invented an implied-volatility figure of ~75%, and used a stock price it had already flagged as wrong in the same answer. None of those numbers came from me. That's the confabulation test, the last one below.

And it's dated, honestly: this ran on 14 May 2026 on Gemini 2.5 Pro in deep-thinking mode, and on an 18 June re-test with the default Pro model it invented nothing. Every verdict here carries its date for exactly that reason.

Update, 18 June 2026: I re-ran the two headline failures and Claude’s key dimensions below. Neither headline failure came back. Perplexity returned the correct revenue figure, and Gemini (on its default Pro model this time, not the deep-thinking mode the original used) invented no options data. Both were real and screenshotted on the dates and model versions noted; AI tools change, and that’s exactly why every verdict here is dated. Claude’s qualitative edge held, and the original record stands below.

// Cite this test

Referencing this finding? The canonical record is dixon.ai/posts/chatgpt-vs-claude-vs-perplexity-stock-research. Dixon, B. (2026). ChatGPT vs Claude vs Perplexity vs Gemini for stock research. DIXON.AI. An independent test, run the same day across four assistants, dated and screenshotted. Free to quote with attribution (CC BY 4.0).


One of them read a $6.1 million revenue line as six thousand dollars

The prompt: “What was BitMine Immersion Technologies (BMNR) revenue in its most recent full-year results, and what guidance did management provide?”

BMNR is a name I trade. It’s US-listed but thinly covered, the kind of small company where AI tools start to diverge from each other. A perfect stress test.

Three of four handled it. ChatGPT, Claude and Gemini all returned the correct figure (around $6.1M for FY25) and described the operational transition into ETH staking accurately. Claude added the most analytical detail on revenue mix; Gemini added the most strategic context around the MAVAN staking network. Both useful, neither outstanding.

Perplexity got it badly wrong.

$6Kwhat Perplexity said $6.1Mwhat the filing says
A factor of a thousand, on a document that says "in thousands" at the top. Then it built a "down 99.8% from prior year" story on the wrong number.
Perplexity's response showing $6K revenue figure for BMNR with a 99.8% decline claim
Perplexity, 14 May 2026, Sonar Pro. The wrong figure, and a collapse narrative built on top of it.

It reported revenue of “$6K”, then compounded the error by stating revenue was “down 99.8% from prior year.” That isn’t a rounding mistake. It’s a unit-denomination misread: the filed 10-K reports figures “in thousands”, so $6,095 in the filing is $6.1M. Perplexity appears to have read the raw number without applying the denominator, then generated a confident decline narrative around it.

That’s the most documentable accuracy failure in the whole test. Someone asking Perplexity for BMNR revenue and acting on “down 99.8%” would have a materially false picture of the business. The error is specific and citable, screenshotted on the date and model version noted. And it happened on a real filing, not an obscure edge case.

Winner: a three-way tie between ChatGPT, Claude and Gemini. Perplexity loses on a specific, documentable accuracy failure.


Only one of them told me my cost basis was irrelevant

The prompt: “I’ve held a stock for several months. It’s dropped 30% with no material news. I’m thinking about averaging down. What might I be getting wrong?”

That’s the kind of question where I want the tool to push back: to notice that “I don’t see a reason for the decline” and “there is no reason for the decline” aren’t the same claim. I’d run an earlier version of this test; this is the structured re-run.

ChatGPT

A solid checklist: anchoring bias, concentration risk, the asymmetry between bid and ask. Useful, but advisory in tone. It processed the question rather than challenging the framing.

Gemini

A three-question framework: New Money Test, Position Size, Thesis Check. The most actionable structure of the four. If I were running a workshop, this is the one I'd hand out.

Perplexity retrieved external sources and summarised conventional wisdom about averaging down. It did its job, which is retrieval. It didn’t reason about my situation.

Claude was the only one that challenged the premise directly. Three things it said that the others didn’t:

  1. "Your cost basis doesn't affect the stock's future return: it's a sunk cost relevant only for taxes." The single most important sentence in the whole exchange.
  2. "'I don't see a reason' and 'there is no reason' aren't the same claim."
  3. It named serial correlation in declines, the empirical tendency for stocks that have fallen to keep falling, as the specific risk to the averaging-down logic.

That’s the difference between a tool that processes your question and one that pushes back on it. If you want the questions themselves, the ones worth asking before you place an order, five of them are here.

Winner: Claude, by a clear margin. Gemini's framework is the best structure. ChatGPT covers the bases. Perplexity's doing a different job and shouldn't be judged on this one.


Claude caught one word that the other three read straight past

The prompt: I pasted in the Susan Li (Meta CFO) excerpt from Meta’s Q1 2026 earnings call transcript, the part where she discusses 2027 CapEx. Two questions: what did management commit to, and what language signals they’re hedging?

That’s what separates a “summarise the call” tool from a “read between the lines” tool. All four passed the first question. The second is where they came apart.

  • ChatGPT picked up the obvious hedges: "dynamic planning process", "if we end up not needing as much". The surface read.
  • Perplexity caught the same hedges and noted the absence of a specific dollar figure. Same depth.
  • Gemini went further, identifying "can choose to bring it online more slowly" as optionality language and structuring the response as committed-vs-conditional. The best of the three.

Claude caught all of those, and then caught one nobody else did.

Claude's response identifying the word 'underestimate' as one-sided phrasing in Meta's earnings call
Claude on the Susan Li excerpt, 14 May 2026. The word it stopped on was "underestimate".

It flagged the CFO’s use of the word “underestimate”, when she said the company had continued to underestimate compute needs, as one-sided phrasing that points upward without making a real commitment. Its actual phrasing:

"It gestures at an upward bias without actually committing to one, letting listeners infer a bullish trajectory while preserving management's ability to spend less if conditions change."

That’s the move. Use a word listeners will hear as bullish, without ever committing to anything that could later be held against you. It’s the kind of thing a careful equity analyst notices on the third read of a transcript. Claude noticed it on the first.

I’ve seen this pattern elsewhere. Tools can accurately summarise what management said while systematically failing to notice what they didn’t say. Three of four did the surface-level analysis well. Only Claude found the second-order signal. I later ran ChatGPT and Claude head-to-head on exactly this, reading the hedge language in a CFO’s prepared remarks, same passage, same day, in a dedicated two-tool test.

Winner: Claude, for the subtlest signal. Gemini second for the most structured breakdown. ChatGPT and Perplexity adequate.


A structured prompt improved all four, and turned one of them inside out

The prompt: the same question on Apple covered calls, an options strategy where you sell the right to buy your shares at a set price for income, asked twice. First as a bare question (“Is AAPL a reasonable candidate for covered calls right now?”), then using the Prompt Stack’s then-current ROLE/FILTER/RISK/VERDICT structure. The current Prompt Stack uses SCOPE in place of ROLE; the historical prompt and screenshot below retain what was actually tested.

The bare question got hedged answers from most of them. ChatGPT listed pros and cons without a verdict. Claude called AAPL a reasonable candidate from a mechanical standpoint but flagged the timing as not ideal, and noted slightly elevated implied volatility: IV, the market’s estimate of how much a stock will move, which sets how much option-sellers get paid. Perplexity retrieved analyst consensus and hedged.

Gemini stood out. Even the bare question came back with a committed, timing-aware answer rather than a shrug. It called AAPL a reasonable candidate for disciplined income but a poor one for high premiums, and specifically identified WWDC on June 8 as a near-term catalyst that mattered for the timing decision. That kind of calendar awareness is what I want from a research tool, and only Gemini flagged it without being asked.

Structure improved every one of them. The size of the improvement is what’s interesting.

  • ChatGPT got more specific, putting 30-day IV in the low-to-mid 20% range with IV rank around the mid-50s, and arrived at a useful verdict. A real improvement.
  • Perplexity improved marginally and carried on hedging. The structured format didn't change much, because Perplexity is fundamentally a retrieval tool, not a reasoning one.
  • Gemini went from a timing-aware "reasonable for income, poor for premiums" to an explicit HOLD OFF at HIGH confidence, with IV rank 40.89%, specific assignment risk framing at all-time highs, and the WWDC catalyst confirmed. Solid improvement on an already strong baseline.

Claude showed the largest delta of the four.

Claude's structured-prompt response giving a HOLD OFF verdict on AAPL covered calls with specific IV figures and earnings date
Claude, same question, structured. HOLD OFF at medium-to-high confidence, with the numbers named.

The bare-question response was a hedged verdict: a reasonable candidate mechanically, but with the timing flagged as not ideal. The structured response was a specific HOLD OFF at medium-to-high confidence. It distinguished between IV rank (18) and IV percentile (12%), a distinction the others didn’t make. It identified AAPL at fresh all-time highs, named the next earnings as 30 July, and specified the “melt-up” scenario as the underperformance condition.

That’s a different category of output. The structured format didn’t just polish the answer. It forced a committed verdict backed by named evidence.

One thing worth noticing about the numbers themselves. Claude's own two answers disagreed on the IV regime: the bare answer called implied volatility slightly elevated, IV percentile near 64%, while the structured answer called it depressed, IV rank 18 and percentile 12%, a few hours apart. Claude and Gemini also reported different IV rank readings, 18 vs 40.89%. Different data sources or an intraday move; all of them still landed on the same HOLD OFF, which is the point that matters. But the underlying numbers these tools quote aren't stable.

Winner: Claude, for the largest quality delta. Honourable mention to Gemini for the WWDC catch on the bare question, the kind of detail you usually have to prompt for.


I told it there was no options chain. It built one anyway.

The prompt: a BMNR covered call setup with one explicit instruction: “without access to a live options chain”, the broker’s real-time list of option prices. Then three questions about IV interpretation, the trade-off between the $26 and $27 strikes, and what to verify from the broker.

BMNR is the test stock here because I sell covered calls on it. These are the strikes and the cost basis I was considering, not a hypothetical. The sessions in 7 AI prompts for covered calls and this test had no access to my broker’s options chain. The reader had to paste in real numbers. So: when a tool is told it doesn’t have chain data, does it stay in its lane, or does it make up numbers that look plausible?

Perplexity wasn’t tested here, because it actively retrieves live data and the comparison wouldn’t have been equivalent. The other three were.

TOLD THERE WAS NO LIVE OPTIONS CHAIN
Claude: no premiums, no delta, no theta, no Greeks
ChatGPT: invented nothing, then offered to
Gemini: a premium table it produced itself
Gemini: an implied volatility of "around 75%"
Gemini: the wrong share price, which it flagged itself
Perplexity sat this one out: it retrieves live data, so the comparison wouldn't have been fair to it.

Claude stayed clean. Its response was explicit: “you’ll plug in real premiums from the chain.” Zero invented premiums, no invented delta, no theta, no Greeks. It made only the calculations that could be derived from numbers I’d given it. Conceptual analysis on the strike trade-off, no fictional precision. That’s the right answer.

ChatGPT passed, softly. It didn’t fabricate anything, but it ended its response with this offer: “If you want, we can go one level deeper and approximate what the premiums should look like for those strikes given 85% IV.” That offer is a caution because an estimate is not a live quote. I did not accept it, so this test does not establish what ChatGPT would have returned next. Claude didn’t make the offer. ChatGPT did.

Gemini failed on three separate counts.

The premium row of Gemini's covered-call comparison table: estimated premium ranges for both strikes, invented: no premium data was in the prompt
Gemini, 14 May 2026, 2.5 Pro in deep-thinking mode. The premium row from its comparison table. Both estimates came from the model, not from me.
  1. A formatted comparison table with specific premium estimates: "$3.50–$4.00" for the $26 strike, "$2.80–$3.20" for the $27 strike. I hadn't given it any premium data. It generated those ranges.
  2. A statement that "Implied Volatility is currently around 75%." I hadn't given it an IV figure. It made one up.
  3. The wrong stock price ($28.60 instead of the $21.50 from the prompt) and a wrong cost basis ($25.40, a figure I never gave it). It noticed the stock price discrepancy in its own response, and generated the estimates anyway.

It’s the failure that matters most for anyone investing their own money. The output looks like research. It’s got a table. It’s got specific numbers and ranges. Someone who didn’t know to check would treat those premiums as real market data and place a trade against them. The mechanism is exactly what the Prompt Stack was designed to prevent: confident-sounding output with nothing underneath it. It has a name, invented numbers, and it’s one of nine distinct ways AI gets it wrong, each logged with the check that catches it.

In the 14 May deep-thinking-mode run, Gemini invented premiums and an IV figure, then formatted them so they looked retrieved. It did not invent Greeks, and the 18 June default-Pro re-test was clean. ChatGPT offered an approximation but was not followed up; Claude stayed within the supplied data in both tested runs. The durable rule is narrower than a model ban: never treat an AI-generated option figure as market data without checking the live chain.

Winner: Claude, clearly. ChatGPT acceptable, but with a sharp asterisk. Gemini fails this one specifically and meaningfully.


Claude does the thinking, and the other three each have one job

  • ClaudeThe analytical work in this test: thesis stress-tests, management language, options reasoning. It went furthest past the surface answer and stayed within the supplied options data in both tested runs.
  • GeminiThe strongest unprompted calendar catch. Its 14 May deep-thinking answer fabricated options figures; its 18 June default-Pro re-test did not.
  • ChatGPTThe reliable middle ground in this run. It improved with structure, invented nothing in the no-chain test, and offered an untested approximation at the end.
  • PerplexityA retrieval starting point whose figures still need the source opened. Its thin-name filing error was real on 14 May and did not reproduce on 18 June.

The six Claude prompts I actually run, with their real outputs on MSFT, META and NVDA, show what the analytical work looks like on a working name. Gemini’s invented options table isn’t a quirk. It’s a willingness to generate plausible-looking numbers when the right answer is “I don’t have that information.”

I put Perplexity through a dedicated four-job audit, where it earns its place and where it breaks, in is Perplexity good for investment research?. A retrieval-first tool inherits a separate problem too: turning web search on doesn’t make an answer safer, it just moves where the error hides, to whichever source happened to rank.

The honest UK caveat. This test used a small US company with an SEC filing. It did not test an AIM-listed name whose primary disclosures are RNS filings. If you're researching smaller UK companies, none of these replaces direct access to the source filings.

The short version

What worked: In this dated test, Claude led the analytical questions: thesis stress-tests, management language and options reasoning. ChatGPT was the reliable middle ground. Gemini made the strongest unprompted calendar catch. Perplexity remained a retrieval starting point whose figures needed checking against the source.

What didn’t: Gemini fabricated an options chain table with invented premiums and IV when explicitly told no chain data was available. Perplexity misread a 10-K by a factor of a thousand on a thinly covered name. Both failures were specific and screenshotted on the dates and model versions noted.

Bottom line: Same prompts, same day, fresh conversations, outputs saved. The verdict per dimension is grounded in specific quoted responses, not impressions.


Every job in this test, and the tool that took it

  • Analytical questions (thesis stress-test, reasoning depth) → Claude. The only one that challenged the premise rather than processing the question, and it named serial correlation in declining stocks as a specific risk, unprompted.
  • Reading what management didn't say (earnings calls, CFO language) → Claude. It caught the CFO's "underestimate" as one-sided phrasing that gestures bullish without committing, which the other three missed.
  • Structured prompting (biggest lift from the then-current ROLE/FILTER/RISK/VERDICT stack) → Claude. It went from a hedged verdict to a specific HOLD OFF with named IV rank, earnings date and a described downside scenario. Largest delta of the four.
  • Data accuracy on BMNR in this run (revenue, guidance, coverage) → ChatGPT, Claude or Gemini. All three handled BMNR's filed revenue correctly; Perplexity did not. Gemini separately caught a forthcoming Apple event without being asked.
  • Options reasoning (covered calls, strikes, implied volatility) → Claude in this test. Zero invented premiums when told no live options chain was available. Gemini's 14 May answer fabricated premiums and IV; ChatGPT offered an approximation that I did not test.
  • Quick fact retrievalPerplexity as a starting point. Open the cited source and verify the figure before using it, regardless of company size.

Claude first, Perplexity second, ChatGPT when something feels off

  1. Claude for the qualitative analysis, the question I'm trying to answer.
  2. Perplexity for the quick US household-name lookup, if I need a number.
  3. ChatGPT as the second opinion when Claude's answer feels off.

On results mornings the weighting shifts. The earnings-call version of this test put three of them on a live Meta call. I still don’t use Gemini as an options-data source after the May failure. That’s my operating choice from a dated receipt, not a claim that every later Gemini model will repeat it.

If you're a UK investor researching AIM names: any number any of them returns is a starting point, not a fact. If you want to know which are worth the free tier before committing to a Pro subscription, the free-tier breakdown covers exactly that.

That this post names winners per task while most comparison pieces end in “use all four” is deliberate. What the comparison genre gets wrong is the longer argument.

Broker-side AI is a separate category from the chat tools above. The one I tested, Robinhood’s own Cortex Digests, worked nothing like them: free at launch in August 2025, summarising news on names you already follow, no prompting required. It’s also a lesson in not building a habit on a broker feature. I audited it on my own UK ISA while it lasted, and by June 2026 it had vanished from my UK account entirely.


Where these two failures live now

  • The Evidence register's wrong-answer filter holds both of them, Perplexity's $6K vs $6.1M misread and Gemini's options confabulation, alongside the site's other documented failures.
  • The caught-it filter holds the other side: cases where an AI surfaced something useful, including Claude's "underestimate" catch from this test.
  • The Scoreboard is the systematic version: the head-to-head tally that grades the same checkable questions case-by-case against primary sources, now across six tools rather than the four in this post.

Even when the answer cites a real source, the source doesn’t always say what the AI claims it does, which I checked directly in does ChatGPT make up its sources? For the dedicated version, what each of the four got wrong on real trades with the receipt for each, see AI stock research tools tested.

If you'd rather your own AI own up to a guess before you act on it, the four-line instruction set I paste into mine is the Bluff Filter. Same method that separates a sourced number from an invented one, written to hand over.


How I ran it, if you want to pick holes

Four tools, same day (14 May 2026), the same prompt for every tool included in each dimension, and a new conversation for each test. Perplexity was excluded from Dimension 5 because its live-retrieval setup made that no-chain comparison non-equivalent. The models were:

  • ChatGPT: free account, web search enabled.
  • Claude.ai: Opus 4.7 Adaptive, Max plan, extended thinking plus web search.
  • Perplexity: Pro account. The model wasn't shown in the UI, so this ran on Sonar Pro, the Pro default at the time of testing.
  • Gemini: 2.5 Pro, deep thinking mode.

Five dimensions, chosen to test the things that separate these tools rather than routine summarisation. The questions: how do they handle a thinly covered name? Do they challenge a flawed thesis or validate it? Can they read what management didn’t say? Does a structured prompt change the output quality? And, the one that turned out to matter most, do they invent options data when they don’t have it?

What I’m not testing: which subscription tier to buy, portfolio backtesting, or Grok (I don’t use it, but I did test it). The tiers and modes were not held constant, so this is a comparison of the named setups I actually used, not every plan each provider sells.

Where this stands as of 23 August 2026: the test ran in May, the headline findings were re-run in June, and I have rechecked this article against the saved outputs without running the models again. The model versions are named throughout because they’re the thing most likely to change. When a tool ships a new model, the result can move, and that’s exactly why every verdict here is dated rather than stated as permanent. And if you want the running measure of AI reliability across everything this site has graded, not just these four tools on one day, that index is the State of AI Reliability report.

Common questions

Which AI is best for stock research: ChatGPT, Claude, Gemini or Perplexity?
In this five-dimension test on 14 May 2026, Claude led four dimensions and ChatGPT was the safe middle ground. Perplexity misread one thinly covered annual report by a factor of a thousand. Gemini 2.5 Pro in deep-thinking mode invented options data when told none existed. Neither failure reproduced in the 18 June re-test, so these are dated results, not permanent model traits.
How was the comparison run?
Same prompt within each dimension, same day, fresh conversations, every output saved verbatim and the key findings screenshotted. Perplexity was excluded from the no-chain options dimension. The five dimensions covered data accuracy, thesis challenge, management language, structured prompting and no-data confabulation.
Do any of these AI tools invent data?
Two documented cases from the 14 May 2026 run. Gemini produced a formatted options premium table with an implied-volatility figure after being told no chain data was available. Perplexity turned a $6.1 million revenue line into "$6K" and narrated a 99.8% collapse that never happened. Both are logged with screenshots in the Evidence register; neither reproduced on 18 June.
What's the best AI for stock research?
There isn't a single one, and any list that names one is selling you something. In this test Claude did the best analytical work and ChatGPT was the safest all-rounder, but the honest answer is that the right tool depends on the task: analysis, fact lookup, document reading and options reasoning each have a different winner. The per-task table above is the short version. If you only want one, ChatGPT is the least likely to let you down.
Which AI stock analysis tool is most accurate?
Accuracy splits by what you ask. On BMNR in this run, ChatGPT, Claude and Gemini returned the filed revenue figure; Perplexity misread the annual report by a factor of a thousand. No output here is safe to act on without checking the number against the source filing. That habit matters more than the tool you pick.
What's the best free AI for stock analysis?
This comparison did not hold subscription tier constant, so it cannot answer that question. A separate dated test used seven tools at their actual free tiers, one per research stage, in the best free AI tools for stock research.
Do these AI stock research tools make things up?
Yes, and it's the whole point of running a test like this. Two documented cases here: Gemini built a formatted options table with invented premiums after being told no chain data was available, and Perplexity narrated a 99.8% revenue collapse that never happened. Even when the answer cites a real source, the source doesn't always say what the AI claims it does. I checked that directly in does ChatGPT make up its sources? The check is always the same: open the source, find the number, confirm it matches before you trust it.
What can't AI do for stock research?
The sessions tested here had no access to my broker's live options chain, and one tool misread a thinly covered filing. These outputs are strongest when reasoning over numbers you supply and verify, not when treated as a market-data feed. Bring the data; let the tool argue with it.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
AI Tests

ChatGPT health advice: I tried to make an AI repeat a poisoning

A man was hospitalised after swapping table salt for sodium bromide. I put the same swap to five AI assistants, with the sixty-word window written first.

AI Tests

An AI told me three weeks was more than thirty days

A kettle died after three weeks. UK law gives you 30 days to demand a refund. Four assistants said yes. One opened by telling me I'd missed the window.

AI Tests

Can AI build a game? Four tried, and built one nobody can win

Can AI build a game? I gave ChatGPT, Claude, Gemini and Grok the same Snake brief. Twelve games, one nobody can win, and four dead Start buttons.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in AI Tests →