This is the kind of thing the Bluff Filter catches. It’s free →
// On this page
I already have the prompts I want to use on a results release. The earnings stack on this site has them: five for after the release, and one for the morning before to pre-commit your triggers. What I didn’t have, until I sat down on the morning of Meta’s Q1 2026 release, was an honest answer to the next question: which AI tools are best for earnings analysis? So I put ChatGPT, Claude and Perplexity through four real dimensions on a live release.
The companion piece to this one (ChatGPT vs Claude vs Perplexity for stock research) compared the same three tools on general research: pre-trade thinking, document reading, options context. That post took the broad view. This one is narrower and harder. Earnings releases are where AI tools either earn their keep or give themselves away. The numbers are a baseline; the interesting work is reading what management chose to say and what they chose not to.
So I ran four tests around Meta’s Q1 2026 release. Same prompts, three tools. Reading the numbers off a page is the part any tool can do; the interesting question was whether it could read what management was carefully not saying.
The clean result is narrower than the original article claimed. Perplexity won live data retrieval and ignored a document-only instruction by running 10 searches. Claude produced the strongest answers to the supplied text, but the test material was flawed: the document contained three wrong figures, and the supposed management excerpt was a constructed passage that Meta never said. The saved outputs remain useful as a record of how the tools handled those prompts. They are not a clean benchmark of who reads a real earnings call best.
How I tested the best AI tools for earnings analysis
Three tools, all on paid plans except ChatGPT (which is on the free plan because that’s where most readers will start):
- ChatGPT: Free plan, GPT-4o with web search
- Claude: Max plan, Opus 4.7 with Adaptive Thinking + web search
- Perplexity: Pro plan, Sonar Pro
These were the models shown in each interface when I ran the test on 15 May 2026. The results below are dated observations from those sessions, not claims about the current model line-up.
The test company is Meta (META), whose Q1 2026 results were released on 29 April 2026. I picked Meta deliberately. It’s a well-covered household name, so any tool with a search function should be able to retrieve the figures. But it had an interesting quarter. Revenue beat the consensus (the average of what Wall Street analysts had forecast) by just under a billion, earnings per share came in roughly $4 above it, and the operating margin held at 41% through a serious spending ramp. And then the stock fell 8 per cent on the news it would spend more on AI infrastructure. That gap, between strong operating numbers and a negative reaction, is where management language analysis should become useful, if the test uses management’s real words.
I tested four dimensions in the order a results-day workflow runs them:
- Get the live numbers, fast
- Read the document
- Analyse the management language (a two-turn conversation)
- Identify what’s missing from the disclosure
No Gemini. The companion comparison documented Gemini fabricating options chain data: making up strikes that didn’t exist. For earnings analysis, the failure mode that worries me more is the one I’ve seen Gemini repeat in my own testing: confident validation of management’s framing rather than challenge to it. That’s the exact failure mode that matters most when the entire question is whether to trust what management has said. One paragraph; moving on.
No specialist tools (AlphaSense, Hudson Labs, Quartr). They are a different product category and were outside this test. This post compares three general-purpose assistants.
Dimension 1: Live data retrieval (the first 10 minutes)
I gave each tool the same prompt: “What were Meta’s Q1 2026 earnings results? Give me: revenue (actual vs consensus), EPS (actual vs consensus), operating margin, full-year CapEx guidance, Q2 revenue guidance, and the market reaction. Flag anything you’re uncertain about.”
All three tools got the key figures right. Revenue $56.31B against consensus around $55.45B. CapEx guidance (the planned spend on data centres and chips) raised to $125–145B from $115–135B. Q2 revenue guide $58–61B. Stock down around 8 per cent, give or take depending on which window each tool read.
Where Perplexity pulled ahead was the precision around the things that aren’t in the headline. It cited 28 sources. It gave the consensus figures to two decimal places. And it caught, without being asked, the EPS ambiguity that matters: “some outlets cite GAAP EPS of $10.44 while others focus on adjusted EPS of about $7.31 after excluding a one-time tax benefit.” That distinction is the difference between a quarter that smashed expectations by $4 a share and one that beat by closer to 60 cents. It’s the kind of thing a sceptical reader needs flagged immediately, and Perplexity flagged it without prompting.

ChatGPT and Claude both retrieved the numbers cleanly with web search. Neither was as precise on the consensus figures, and neither flagged the EPS ambiguity unprompted. For the first 10 minutes after a release, when you want the headline numbers and the points of confusion called out, Perplexity is doing the job it was built for.
Winner: Perplexity. This is what live retrieval looks like when the tool’s architecture matches the task.
Dimension 2: Document reading
My setup: paste the Q1 figures into the chat, ask the tool to interpret what’s actually in the document. Identify what consensus figure is missing. Read the gap between GAAP and ex-tax EPS. Assess whether the CapEx raise is significant. Interpret the DAP (Daily Active People) growth against revenue growth. Critically: “work only from the document.”
Correction, 23 August 2026: the document I pasted contained three transcription errors. It said net income was $33.39B rather than $26.77B, DAP was 3.43B and up 6% rather than 3.56B and up 4%, and Q1 CapEx was $13.7B rather than $19.84B. Meta’s filing confirms the corrected figures. All three tools received the same faulty document, so the search-behaviour comparison is still observable. The finer analytical ranking is less reliable than I first presented it.
ChatGPT and Claude both did this cleanly: they stayed inside the document and reasoned from what was there.
Claude’s read of the DAP figure was the sharpest observation in the response it gave. Based on the faulty figures in the prompt, revenue up 33 per cent and DAP up 6 per cent, it said revenue was growing roughly five times faster than the user base, so the growth engine was revenue per user rather than user acquisition. That was sound arithmetic from the supplied text. It was not an analysis of Meta’s correct 3.56B and 4% figures, and the original article was wrong to rewrite Claude’s observation as “roughly eight times”.
ChatGPT’s response was solid but didn’t reach the same point. It described the figures accurately and noted the CapEx raise. It didn’t quite get to “the story of the quarter is revenue per user, not user growth.”
Perplexity reported 10 sources on this prompt, despite the instruction to work only from the pasted document. Its answer also added context that was not in the document. The capture proves a constraint-following failure in this run. It does not prove that every sentence came from the same external coverage used in D1, which the original article claimed without preserving the source list.
Result: Claude gave the most useful document-bound observation, but this is not a clean winner. The input figures were wrong. Perplexity’s 10-source answer still establishes that it did not follow the document-only constraint in this run. If document discipline is the part you care about, the quality-of-earnings review pushes it further on the same company: the cash-versus-profit strip, run on the full filing.
Dimension 3: Management language analysis (the centrepiece)
This was supposed to be the dimension that separated an earnings tool from a number-retrieval tool. It is also where the test failed.
Correction, 23 August 2026: the passage was labelled in the prompt as an excerpt from Meta’s Q1 2026 earnings call. It was not. The wording “we remain committed to investing aggressively” and “our conviction in the opportunity ahead” does not appear in Meta’s official call transcript. It was a constructed Meta-style passage assembled around real themes and the real CapEx range. The outputs below show how the tools analysed that constructed passage. They do not show what the tools caught in Meta management’s actual wording.
I ran it as a two-turn conversation. Part A: paste the constructed passage, mislabelled as a Meta Q1 2026 management excerpt, and ask the tool to analyse the language for tone and intent before knowing the results. Part B: then reveal that Q1 delivered revenue +33%, margin flat at 41%, and the stock fell 8 per cent. Does the language look appropriate in hindsight, or does it read differently now?
Part A: analysing the language cold
ChatGPT was strong here. It identified the difference between “we believe” (a forward bet) and “we’re seeing” (a claim about present evidence), and noted the absence of an ROI timeline or any upper bound on the spending.
Claude’s Part A was sharper. It picked apart the specific phrasing in the constructed passage in a way that read less like a summary and more like an annotation:
The phrase “we remain committed” is particularly pointed: “remain” implies the commitment predates this call, “committed” implies resolve against pressure. Anyone listening in hopes of hearing spending discipline is being gently told to stop hoping.
It identified “conviction” as a tell-word that shows up when management wants to project certainty without offering specific metrics that could be checked against it. And it spotted the rhetorical move at the heart of the statement: “investing aggressively” paired with “positions us well for the long term” relocates the payoff to an unfalsifiable future horizon. That pre-empts the criticism that near-term margins will suffer.
Perplexity reported 15 sources on Part A even though the task could be answered from the supplied passage. Its answer still analysed the wording directly. The supported criticism is the unnecessary retrieval, not the original article’s claim that it replaced language analysis with a market-reaction summary.
Part B: reassessing the language with the results in hand
This is where the test became interesting.
ChatGPT didn’t get to Part B. The free plan’s daily GPT-4o allocation ran out partway through the session, tapping out, as it happened, right before the most interesting question in the test. The model got downgraded for the rest of the conversation, and the Part B response simply didn’t run. That’s a real finding for anyone planning to use ChatGPT free for a multi-step earnings session: a four-dimensional analysis can exhaust the daily allocation. I’d planned to run all four dimensions in one sitting; I couldn’t, on the free plan, even on a normal-length session. Worth knowing before you start.
Perplexity again reported 15 sources on Part B. It still addressed the supplied language, so the issue is the same avoidable retrieval rather than a failure to answer.
Claude’s Part B was the standout of the entire test battery. Here’s the central observation, verbatim:
What looks worse in hindsight is the vagueness. With revenue up 33% and margins holding at 41% through an aggressive spending ramp, management actually had a strong empirical case to make and didn't make it. [...] Leaving that unspoken and reaching instead for "conviction in the opportunity ahead" traded a fact-based defence for a faith-based one, and the market priced the faith.
That is a sharp read of the passage I supplied. It is not evidence that Meta chose conviction language instead of specific evidence, because Meta did not say those lines. The original article crossed that boundary and treated an analysis of constructed wording as an analysis of management’s real choice of words.

Result on the constructed-passage task: Claude gave the most incisive response. ChatGPT’s Part A was useful but Part B did not run. Perplexity searched beyond the supplied text on both turns. This cannot support a winner for reading Meta’s actual earnings-call language.
Dimension 4: The omissions test
This is the hardest task in earnings analysis: not what management said, but what a sceptical investor would have expected them to address that wasn’t there.
My setup: I pasted the same faulty Meta Q1 2026 summary, this time deliberately removing the ex-tax EPS and total-expense guidance, and asked: “What important information is missing from this summary that an investor would need to properly evaluate these results?”
Claude identified 12 gaps. The first one, unprompted, was the Reality Labs segment split:
Meta reports two segments, Family of Apps and Reality Labs, and the consolidated operating income of $22.87B blends them. RL has historically lost $4–5B per quarter. Without the split, you can’t tell whether FoA margins are expanding, holding, or actually deteriorating while a smaller RL loss flatters the total.
This is the right answer. Reality Labs has been a structural drag on Meta’s reported operating income for years. The consolidated number tells you nothing about whether the core advertising business is healthy or whether a smaller RL loss is doing the flattering. It’s the single most important number Meta discloses each quarter, and Claude put it at the top of the list without being prompted to.
The most sophisticated gap in Claude’s list was the depreciation trajectory:
A $125–145B CapEx year, on top of prior years’ builds, means depreciation expense will step up meaningfully over the next 4–6 quarters as assets come into service. Today’s 41% margin reflects yesterday’s asset base. Without commentary on useful lives and the expected depreciation ramp, you can’t model forward margins.
That requires multi-step reasoning. CapEx today turns into assets entering service in coming quarters. Those assets depreciate, which lifts the depreciation line, which compresses margins. None of that chain is in the document. It requires applying domain knowledge about how data-centre buildouts work to the figures that are there.

Perplexity identified 11 gaps via 10 sources. It also flagged Reality Labs at #2. Both Claude and Perplexity identified headcount; the original article was wrong to say Claude missed it. Perplexity combined headcount with restructuring and added a planned workforce reduction from outside the pasted document.
ChatGPT identified 10 gaps with the downgraded model, which makes a like-for-like comparison unfair. The structure was sensible (Reality Labs was at #2, free cash flow at #5, ad economics in the list) but the model wasn’t operating at full capacity by this point in the session.
Result: Claude produced the deepest list and the most sophisticated single gap, but its answer also guessed a normalized EPS of about $7.90; Meta’s release says $7.31. Perplexity gets credit for the restructuring context, not a unique headcount catch. ChatGPT’s result is partially compromised by the rate-limited downgrade. Because all three received the faulty summary, this remains a comparison of their responses to that fixture rather than a clean earnings-analysis benchmark.
The original scorecard, with the test defects visible
The table below records how I scored the May run. I have not silently regraded the saved field-test record. Read it alongside the corrections above: D3 is not a valid test of Meta’s real language, and D2 and D4 used wrong figures.
| Dimension | ChatGPT | Claude | Perplexity |
|---|---|---|---|
| Live data retrieval | Equal | Equal | Better |
| Document reading | Equal | Better | Worse (searches, not reads) |
| Management language | Worse (rate limit) | Better | Worse (retrieves, not analyses) |
| Omissions detection | Worse (downgraded model) | Better | Equal |
| Best for | Reliable middle ground | Analytical depth | First 10 minutes |
The only clean winner from this battery is Perplexity on the live-number retrieval task. Claude produced the deepest document-bound responses, but the faulty fixtures mean this run does not prove it is the best tool for the next 20 minutes. ChatGPT’s free-plan session hit its model limit before Part B and used a downgraded model for D4; that is a dated observation from this account and session, not a universal daily allowance.
A separate May 2026 check found a narrower UK limitation. Perplexity Finance showed price data and earnings transcripts for Unilever’s London listing, but its structured analyst tab returned “No data found”; the Apple page showed named analyst ratings and price targets. That does not mean Perplexity’s UK retrieval edge disappears. It means the structured sell-side layer was absent for that tested London listing. The companion tool audit documents the broader coverage problem.
The short version
What worked: Perplexity’s live retrieval was the cleanest result: correct headline figures, precise consensus ranges and an unprompted warning about GAAP versus adjusted EPS. The document-only run also preserved a real constraint finding: it reported 10 sources despite being told to stay inside the pasted text.
What didn’t: The test design. D2 and D4 used three wrong figures. D3 called a constructed passage an earnings-call excerpt. ChatGPT also hit a model limit mid-session, and Perplexity retrieved beyond the prompt on every document task.
Bottom line: Keep the retrieval finding and the documented constraint failure. Treat the Claude answers as useful exploratory outputs, not proof that it read Meta’s real language better. A clean comparison needs the official release and transcript pasted verbatim, with the fixture checked before any model sees it.
The durable workflow is simpler: keep the official filing open as the source of truth, paste the exact text you want analysed, and verify the fixture before comparing outputs. Perplexity earned the first retrieval step in this dated test. The rest needs a clean rerun before I name a winner. The Prompt Stack, and specifically the kind of structured prompting in the 5 questions to ask AI before buying any stock post, is useful only after the source text is right. For the management-language task, five prompts for reading the language of an earnings call gives the structure; the official transcript must supply the words.
Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →
The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.