xAI -- now badged SpaceXAI after its February merger into SpaceX -- shipped Grok 4.7 on Sept. 21, two months after Musk's first delivery estimate and five walked-back timelines. The model was originally due "in about 4 weeks" by Musk's own July 24 count; the date slid to "a few weeks," then "3 to 4 weeks" once Grok 4.6 shipped Aug. 12, then "10 days" on Sept. 1, then "a few more days to cook" on Sept. 11, when Musk said reinforcement learning had penalized response length so heavily that the model was quitting solvable hard problems early rather than finishing them. The independent Artificial Analysis Intelligence Index (v4.3.2) scores the finished model at 46 -- two points above Grok 4.6's 44, and seven behind Claude Fable 5.1 and GPT-6 Astra, which both score 53.
Five walked-back dates, then a ship
- Jul 24 — Musk says Grok 4.7 is "about 4 weeks" out.
- Jul 28 — Timeline softens to "a few weeks."
- Aug 12 — Grok 4.6 ships; Musk pegs the successor at "3 to 4 weeks."
- Sep 1 — Estimate narrows to "10 days."
- Sep 11 — "Needs a few more days to cook" -- RL over-penalized response length, model quit solvable tasks early.
- Sep 21 — Grok 4.7 ships, scoring 46 on the independent Intelligence Index.
xAI's own account of the change is a larger base model, a reinforcement-learning run stretched toward tasks that take hours rather than minutes, and training aimed specifically at making the model check its own work before answering -- a direct response to the failure Musk named on Sept. 11. It appears to have partly worked: Artificial Analysis measured Grok 4.7's hallucination rate at 29%, down from 34% for Grok 4.6 on the same evaluation set. The model ships in two forms -- a standard tier at $2 per million input tokens and $6 per million output, and a faster variant billed at twice those rates inside Cursor and Grok Build only, not the public API. Context window holds at 500,000 tokens; xAI has not published a parameter count, and third-party trackers that estimated Grok 4 and Grok 4.6 at 2.1 trillion parameters expect no change for 4.7. xAI also says the model now natively understands the harness behind its Grok Build coding agent, which the company frames as a conversational-quality improvement rather than a raw-capability one -- a distinction that matters, because it is the kind of gain that shows up in a coding agent's throughput long before it shows up on a general reasoning benchmark.
The clearest gain shows up in a different benchmark than the headline index. Grok 4.7, run through xAI's own Grok Build coding agent, lifted Artificial Analysis's Coding Agent Index score from 47 to 56 -- moving it into 4th place among the agents the firm tracks. That is the workload xAI is actually selling Grok 4.7 into: Cursor integration and agentic coding, not a head-to-head swap for a general-purpose assistant, which is also where the model's steepest jump on Artificial Analysis's AA-Briefcase knowledge-work benchmark shows up -- a gain of 111 Elo points over Grok 4.6.
Cheap per token, not so cheap per task
| Grok 4.7 xhigh effort | GPT-6 Astra max effort | Claude Fable 5.1 max effort | |
|---|---|---|---|
| Intelligence Index score | 46 | 53 | 53 |
| List price, input/output per 1M tokens | $2 / $6 | $10 / $50 | $10 / $50 |
| What Artificial Analysis actually spent running its own eval, per task | $3.74 | $3.26 | $5.98 |
That $2/$6 list price undersells what a task actually costs. Artificial Analysis -- which runs the same fixed evaluation set against every model it measures and reports what it actually spent running it -- clocked Grok 4.7 at $2.73 per task at standard effort, against $0.82 for GPT-6 Astra's cheapest setting. The reason is verbosity, not the sticker price: Grok 4.7 burns roughly 81,000 output tokens per task at its highest setting, against 27,000 for GPT-6 Astra on the same benchmark -- nearly three times the tokens to reach a lower score. At maximum effort the picture is similar: Artificial Analysis spent $3.74 running Grok 4.7 (xhigh) per task, against $3.26 for GPT-6 Astra (max) and $5.98 for Claude Fable 5.1 (max) -- both of which list at $10 per million input tokens and $50 per million output, five times Grok 4.7's rate.
Musk's own read on where that leaves Grok 4.7 is unusually candid for a launch day. Asked to place it against Anthropic's latest, he wrote that it "should be roughly on par with Opus 5.0, not 5.1" -- better in some ways, worse in others -- and has already moved the AGI framing to the model after next. Grok 4.8 is described as finishing training within days and pitched as "a meaningful step up"; Grok 4.9 is aimed at Astra/Fable class. On Musk's own account, Grok 4.7 is a bridge release, not the arrival.
“It should be roughly on par with Opus 5.0, not 5.1 — better in some ways, worse in others.” — Elon Musk, on Grok 4.7 against Anthropic's latest models
What happens next is the more interesting number. xAI has shipped a numbered Grok release roughly every four to six weeks since Grok 4 -- on that cadence, both 4.8 and 4.9 could land before Grok 4.7 has been broadly benchmarked outside Artificial Analysis's own lab. That pace is itself the strategy: xAI has kept every Grok 4.x release at or near $2/$6 per million tokens since Grok 4.5, betting that shipping often at a fixed low price outcompetes shipping less often at a higher score -- a bet Grok 4.7's cost-per-task numbers complicate but do not obviously refute, since a buyer choosing on price alone still pays less than half of GPT-6 Astra's list rate even after the token-verbosity gap is priced in. The gap that matters now is the one between Grok 4.7 and whatever ships next, not the one between Grok 4.7 and the deadline it missed five times. (The 46 score is currently one benchmarking group's measurement. Whether an outside evaluator lands near the same number once third-party routers carry full traffic is the number actually worth watching.)
- xAI shipped Grok 4.7 on Sept. 21, two-plus months after Musk's first delivery estimate.
- Independent index score: 46, two points above Grok 4.6, seven behind Claude Fable 5.1 and GPT-6 Astra's 53.
- List price is a fifth of the frontier's, but per-task cost narrows because Grok 4.7 uses far more tokens.
- Musk says it's "roughly on par with Opus 5.0, not 5.1" and points the real AGI bet at Grok 5.
- Caveat: the score and cost figures so far come from one benchmarking group, not yet an independent second read.