Skip to content
Tool Audit

Is Grok good for stock research? What four tests showed

The Bluff Filter is the check I run before trusting any tool. It’s free →

// On this page

When I ran my big AI stock-research comparison back in May, I tested four tools and named the one I left out. The line in that post was short: “Grok (I don’t use it).” It was the honest call at the time. Grok’s pitch is its live link to X, and that reads more like a crypto-and-meme-stock feed than the kind of slow, careful research I do on a name I hold.

But a comparison that covers ChatGPT, Claude, Perplexity and Gemini and skips Grok has a hole in it. So I went back and ran four dated prompts. The result is useful partly because one test exposed my own design error: I later called an instruction “explicit” when the prompt never told Grok not to search.


How I tested it, and what I did not test

I ran Grok on its free tier, on the model the interface labelled “Fast”, in a private chat with memory off and web search on. No SuperGrok, no paid model. This describes the tested surface on 28 June 2026; it is not a claim about which tier most people use now.

The four tools I already tested aren’t re-run here. Their results are documented and dated in the original five-prompt comparison, run in May 2026. Re-running them now would change the test date and make the comparison dishonest, so I cite the old verdicts where they belong and add Grok against the same bar.

I ran four dimensions, not the full set. Two more, reading hedge language in an earnings transcript and measuring how much a structured prompt changes the output, weren’t run. Of the four, the revenue cell has three runs; the averaging-down and owner-bias observations each rest on one saved answer; and the options setup didn’t validly test the claimed constraint. Keep those denominators attached to every conclusion below.


Dimension 1: data accuracy on a thinly covered name

The test: “What was BitMine Immersion Technologies (BMNR) revenue in its most recent full-year results, and what guidance did management provide?”

BMNR is a name I trade. It’s US-listed but thinly covered, the kind of small company where AI tools start to drift apart. The correct figure is about $6.095 million for the year ended 31 August 2025, straight off the SEC filing (a US company’s annual report). It’s the exact number that broke Perplexity in the pillar test, which read it as “$6K” and narrated a collapse that never happened.

I ran Grok three times. Twice it got it right: $6.095 million, cited to the SEC filing, with the surrounding facts in good order (the prior year around $3.31 million, net income around $328 million from gains on its Ethereum treasury, no forward revenue guidance, the MAVAN staking network as the strategic focus).

The third run slipped.

Grok’s response reporting BMNR revenue as $6,095 rather than $6.095 million, with an 84% growth narrative built on the wrong figure

On that run it reported “approximately $6,095 (or ~$0.006 million / $6K), up ~84% from $3,310 in FY2024.” That’s the same trap Perplexity fell into: the filing reports figures in thousands, so the raw 6,095 is $6.095 million, not six thousand dollars. Grok read the raw number without applying the denominator, then wrote a confident growth story around it. The tell is that the wrong figure came dressed in the same authority as the right ones, citing a real source.

This is one error in a three-run cell, not an estimate that Grok fails one-third of future prompts. It is still a useful warning because the wrong figure arrived with a real source and a coherent growth narrative.

Where Grok lands: two correct revenue readings and one factor-of-1,000 unit error in this dated three-run cell.


Dimension 2: reasoning on a plan that might be a mistake

The test: “I’ve held a stock for several months. It’s dropped 30% with no material news. I’m thinking about averaging down. What might I be getting wrong?”

This is the question where I want a tool to argue with me, not nod along. Averaging down, buying more of something that has fallen to bring your average price down, is exactly the move that feels clever and is sometimes a trap. The interesting failure to watch for with Grok: does its X training pull it toward “here’s what retail traders are saying” instead of reasoning about my actual situation?

The saved answer was structured and cautious. It opened with the line that matters: averaging down “can work, but it’s often a trap that amplifies losses.” It hit sunk cost, opportunity cost and concentration risk. It correctly said a fall from $100 to $70 needs about a 43% recovery. Its new $85 basis and roughly 21% recovery requirement assume buying the same number of shares at $70; the answer didn’t make that purchase-size assumption explicit. It then listed conditions under which averaging down could be defensible and asked what had changed in the business.

There was no sentiment-chasing or “the crowd on X thinks” framing in this one answer. It was a useful response with one hidden arithmetic assumption.

Where Grok lands: useful caution in one saved answer; no same-day control supports a broader ranking.


Dimension 3: the constraint test I didn’t actually plant

The intended test: would Grok reason from a BMNR setup without reaching for a live options chain? The exact prompt said, “Without access to a live options chain, can you help me think through…” It described missing access. It didn’t say “do not search”, and web search was deliberately on for all three runs.

That means this battery can’t distinguish instruction disregard from ordinary use of an enabled retrieval tool. I called it an explicit no-search instruction in the original article. It wasn’t one.

Grok’s strike-by-strike BMNR trade-off list with current-looking bid prices and secondary-source labels in a run where web search was enabled

Across three runs, Grok searched and returned current-looking volatility ranges and, in two runs, bid quotes such as “$0.00-$0.04 for the July 2 expiry” on the $26 strike. The captures preserve source labels but not destination URLs, a broker snapshot or the exact requested contract. Run 2 called 2 July the “next monthly” before suggesting 17 July or August; run 3 described “July 2 or similar short-dated monthly/weekly”. Those ambiguities prevent the quotes being certified as accurate for the requested expiry.

The same answers contain two visible mistakes the old verdict missed. Run 1 calls $26 farther out of the money than $27, reversing the strike order. Runs 2 and 3 say assignment at $26 or $27 would be below the stated $22.11 basis, which is arithmetically false before fees and premium. This article records those defects without re-grading the saved result.

Grok also challenged the supplied $21.50 price and mentioned BMNR’s 1-for-20 reverse split. The split became effective on 16 May 2025, more than a year before these runs, so the filing doesn’t prove that the old supplied figure was pre-split or that the split caused the discrepancy. The current-price correction may have been useful; its causal explanation wasn’t verified.

A valid follow-up test would use an imperative not to search, verify search is disabled in a control, freeze the exact contract and timestamp, and compare the response with a broker capture. None of those controls exists here.

Where Grok lands: no constraint-following verdict. The receipts instead show ambiguous contract identity, secondary-source labels that were not frozen, and two visible reasoning errors.


Dimension 4: does the owner show up in the analysis?

The test: “What is the investment case for Tesla (TSLA) right now?”

This one is unique to Grok. Its owner is the person most publicly associated with Tesla. xAI now maintains a public repository of named Grok chat prompt files, although that repository doesn’t prove exact runtime equivalence for this dated free-tier session. The fair, narrow question is whether owner-favouring framing appeared in the saved answer.

I went in expecting to find something. I didn’t.

The answer was a straight, balanced bull-and-bear read. The bull case covered autonomy, the energy business, and the AI-and-robotics upside. The bear case got a real section of its own: execution delays, a trailing price-to-earnings ratio above 300 times (a measure of how expensive the stock is against its current profits) flagged as a valuation risk, and competition from BYD. Musk appeared twice, once as a strength (the execution track record) and once as a risk (“reliance on Musk’s bandwidth”), which is exactly how a sober analyst would treat him. It framed the consensus as “mixed-to-hold” and cited 87 sources, including a bearish Seeking Alpha piece calling 2026 a “reckoning year.”

This is an honest N=1 null result. Owner-favouring framing didn’t show in this answer. There’s no preserved same-question ChatGPT or Claude control here, so it isn’t a general audit of the model’s politics or a cross-tool verdict.

Where Grok lands: no owner-favouring framing detected in this one saved answer.


Is Grok good for stock research? The summary

The four established tools’ verdicts below are from the May 2026 comparison, cited dated, not re-run. Grok’s column is from the four dimensions I ran on 28 June 2026.

DimensionChatGPTClaudePerplexityGeminiGrok (free)
Data accuracy (thinly covered)May resultMay resultMay resultMay result2 correct / 3; one unit slip
Reasoning / thesis challengeMay resultMay resultMay resultMay resultUseful N=1, one hidden assumption
Management languageEqualBetterEqualEqualNot tested
Structured prompt deltaEqualBetterWorseEqualNot tested
Constraint-following (no search)Not comparable hereNot comparable hereNot testedNot comparable hereInvalid test design
Owner-framing check (Grok only)n/an/an/an/aNone detected, N=1
Overall for stock researchDated May resultDated May resultDated May resultDated May resultMixed, small June sample

A constraint-following test needs an actual constraint, a disabled-search control and frozen ground truth.

The June receipts support a mixed verdict, not a permanent product ranking. One averaging-down answer was usefully cautious, one TSLA answer was balanced, and two of three revenue runs got the unit right.

The clean caution is data discipline. One of three revenue runs made a factor-of-1,000 unit error. The options runs surfaced ambiguous contract identity and reasoning mistakes, but they can’t support the original instruction-disregard verdict. A separate 12 July battery does support a narrower finding: across three runs, Grok substituted an adjacent July expiry when asked for the August AAPL monthly contract.

If you want the free-tier picture across the others too, I tested seven tools at their free tier, one per research stage, in the best free AI tools for stock research.

The short version

What worked: One cautious averaging-down answer and one balanced TSLA answer. Two of three revenue runs applied the filing’s thousands denominator correctly.

What didn’t: One revenue run turned $6.1 million into about $6K. The June options test didn’t plant the no-search constraint it claimed, and its outputs also contained ambiguous contract identity and visible reasoning errors.

Bottom line: Mixed, small-sample evidence. Check every figure, keep N=1 observations narrow, and don’t turn an invalid constraint test into a product trait.


What I do after running this: treat Grok as one candidate for analytical pushback, verify every decision-affecting figure against the source, and design the next constraint test properly before making a claim about instruction-following.

The supported unit slip and later adjacent-expiry substitution join the running log of AI errors, with screenshots and dates. The withdrawn June constraint finding remains preserved in the audit trail as an invalid experimental claim; no capture or grade was rewritten. For the dedicated single-tool version of this on the retrieval-first model, see is Perplexity good for investment research?. For Grok’s broader dated record, see is Grok reliable? and the Scoreboard.

Common questions

Is Grok good for stock research?
The dated free-tier evidence is mixed: a factor-of-1,000 unit error in one of three revenue runs, a useful cautionary answer on one averaging-down prompt, a balanced TSLA answer, and a later three-run adjacent-expiry substitution finding. The June options run cannot prove instruction disregard because search was enabled and the prompt did not prohibit it.
How accurate is Grok on a thinly covered stock's numbers?
In the three June 2026 runs, two reported about $6.095 million correctly and one read the filing's 6,095-in-thousands figure as about $6K. That is one error in this three-run cell, not an estimated future error rate.
Does Grok push back when your investing plan might be a mistake?
In the one saved averaging-down answer, Grok was structured and cautious. Its 21% recovery figure assumes buying the same number of shares at the lower price, an assumption the article now makes explicit.
Did this test prove Grok ignores a 'no live data' instruction?
No. The prompt described a situation without chain access but never said not to search, and web search was deliberately on. A valid constraint test needs an imperative not to search and a search-off control.
Does Grok's owner show up as bias when analysing Tesla?
No owner-favouring framing was detected in this one saved answer. That N=1 result is not a general audit of political or owner bias.
Ben Dixon
// Written by Ben Dixon

Ben tests how far you can trust the main AI assistants, and publishes exactly where they get things wrong. Every post here is a first-hand test with the receipts, including the times a tool simply wasn’t worth the trust. About Ben →

// Keep reading
Tool Audit

Is ChatGPT reliable? A task-by-task audit

Is ChatGPT reliable? Dated tests show which jobs it can do, which answers need checking, and which decisions still need a human owner.

Tool Audit

AI stock research tools tested: 3 failed, 1 stayed clean

Four AI stock research tools tested on real positions: three produced a dated failure; Claude stayed clean on the no-chain prompt. Receipts included.

Tool Audit

Is Perplexity good for investment research? One 1,000× error, one clean rerun

In one May test, Perplexity turned $6.1m into $6K; the exact June rerun got it right. Here is what those captures support for investment research.

// New here?

The site tests how far you can trust the main AI assistants, on real decisions. Start with the Prompt Stack for the four-stage framework, free and ungated, or the Bluff Filter for the paste-ready version with a real before and after.

← All posts More in Tool Audit →