SpaceXAI released Grok 4.5 on July 8, and the launch materials describe it — in Elon Musk's words — as "an Opus-class model, but faster, more token-efficient and lower cost." The benchmark tables tell a more specific story: on most headline evals, Grok 4.5 is not the best model available, and the launch doesn't really pretend otherwise. What it is, per the published pricing, is dramatically cheaper than everything it's compared against. That's the actual product here — and evaluating it honestly means looking at three things: the capability numbers, the economics, and two caveats buried in the fine print that deserve more attention than either.
What shipped
Grok 4.5 is a mixture-of-experts model that Cursor — the coding-tools company SpaceX is reportedly acquiring for $60 billion, per The Information — says it "trained jointly with SpaceXAI," incorporating what Cursor describes as trillions of tokens of its own developer-interaction data plus STEM and knowledge-work material. That joint-training arrangement is the release's most structurally interesting fact: a frontier model co-built with the application company that supplies its most valuable training distribution, announced while the acquisition is still in flight. Whatever else this launch is, it's a preview of what vertical integration between labs and tooling companies looks like.
The published pricing is $2 per million input tokens and $6 per million output — set against Opus-tier pricing at $5 and $25, GPT 5.5 at $5 and $30, and Fable 5 at $10 and $50. SpaceXAI also claims 80-tokens-per-second serving speed and 4.2 times fewer tokens consumed than Opus 4.8 on SWE Bench Pro. Those efficiency figures are the company's own; treat them accordingly until independent testing accumulates.
What the numbers actually show
On capability, the pattern across the published evals is consistent: competitive, rarely leading. On Terminal Bench 2.1, Grok 4.5 scores 83.3% against GPT 5.5's 83.4% and Fable 5's 84.3% — effectively a three-way tie. On DeepSWE 1.1 it trails meaningfully: 53% against 67% and 70%. On SWE Bench Pro it lands at 64.7%, behind Fable 5's 80.4% and Opus 4.8's 69.2%, though ahead of GPT 5.5. The one headline eval where it leads is SWE Marathon's pass@1 resolution, at 29.0% against Opus 4.8's 26.0%. Independent aggregation tells the same story: Artificial Analysis places it fourth on its Intelligence Index — a sixteen-point jump over its predecessor — and at 76 on its Coding Agent Index, tying GPT 5.5, one point behind Fable 5.
Then the economics, where the story inverts. On Artificial Analysis's agentic-coding cost accounting, a Grok 4.5 task runs $2.49 against $5.07 for GPT 5.5 and $11.80 for Fable 5 — driven not just by unit price but by token appetite, at a reported 1.9 million tokens per task against 6.2 and 7.2 million respectively. If those figures hold up in real workloads, the effective cost gap isn't the 2-to-5x the price sheet suggests; it's larger. For the large class of tasks where fourth-best intelligence is entirely sufficient — and honestly, that class covers most production work — that arithmetic is the whole pitch.
The benchmark tables say 'not the best model.' The price sheet says 'you might not care.' Both are true, and the second one is the product.
The two caveats that matter more
First, reliability. Artificial Analysis's independent testing found that while Grok 4.5's accuracy on its Omniscience Index rose from 35% to 52% generation over generation, its hallucination rate jumped from 25% to 54% — the model knows more, and it is also substantially more confident when it's wrong. For a model marketed on agentic work, where errors compound across steps instead of sitting quietly in a chat window, that is not a footnote. It's arguably the most important number in this entire release, and it points the wrong direction.
Second, contamination — disclosed, to Cursor's credit, in its own launch post. Cursor states that "an earlier snapshot of the Cursor codebase was accidentally included in training," giving Grok 4.5 an advantage on CursorBench, and that "the exact impact is unclear." The same post notes that third-party benchmark scores in the announcement are self-reported. Disclosing this voluntarily is better behavior than the industry norm, and it should be credited as such. It also means exactly what it says: at least one published result is known-inflated by an unknown amount, and the rest carry the standard self-reporting discount. Adjust your confidence intervals accordingly.
What's still unproven
Everything that matters most, as usual. The efficiency claims are vendor-published and the hallucination finding is one independent shop's measurement; the next two weeks of third-party testing will settle both, and we'll report what they find. The deeper open question is strategic: whether a model priced like a commodity and positioned as "good enough, much cheaper" can hold that position once competitors reprice — or whether the hallucination numbers surface in production and reprice it themselves. Musk's own internal calibration, per his launch comments, is that Grok 4.5 is "roughly comparable to Opus 4.7, but much faster." That's a strikingly modest claim by launch-day standards, and probably close to accurate. The honest summary: a genuinely competitive fourth-place model at a first-place price, shipped with one alarming reliability signal and one honestly disclosed asterisk. Watch the independent numbers, not the announcement.
- Grok 4.5: competitive but rarely leading on benchmarks, at $2/$6 per million tokens.
- Real agentic tasks run ~$2.49 versus $5.07 (GPT 5.5) and $11.80 (Fable 5).
- Cursor co-trained it on trillions of tokens of its own developer data.
- Independent testing found hallucinations jumped from 25% to 54% generation over generation.
- Caveat: one benchmark is known-inflated (disclosed contamination); most launch scores are self-reported.
