Anthropic and OpenAI both shipped model updates on Sept. 22, and the two releases tell almost opposite stories. Claude Opus 5.5 scores 58 on Artificial Analysis's Intelligence Index -- a genuine capability jump that puts it outright first, five points clear of the prior tie between GPT-6 Astra and Claude Fable 5.1 (53 each) and seven ahead of the Opus 5 it replaces, which scored 51, at a 20% lower list price. GPT-6 Sol and Luna, by contrast, barely moved the capability needle -- Sol gained one point over GPT-5.6 Sol, Luna gained none -- while both models' list prices roughly halved. One lab shipped a better model; the other shipped a cheaper one. Both are real results, and neither headline works as a substitute for the other.
Opus 5.5 launched at 12:05pm PT with a $4/$20 per-million-token price (input/output), down 20% from Opus 5's $5/$25, and cache reads cut 60% to $0.20 per million tokens. Anthropic's own benchmark table shows the jump is broad, not a single-metric artifact: Terminal-Bench 4.0 up to 66.4% from 52.3%, CursorBench 4.0 up to 57.8% from 46.6%, GDPval-AA to 1846 Elo from 1708. Anthropic separately claims Opus 5.5 runs about 30% faster than Opus 5 and costs roughly 40% less on typical workloads once its lower token usage is counted alongside the list-price cut -- both company figures, not independently measured here. The release carries real breaking changes, not just a version bump -- forced tool use can now return errors under new conditions, and the older `computer_20251124` computer-use tool is discontinued on the API and Google Cloud -- and it ships alongside an expanded Cyber Verification Program and a Life Sciences Verification Program, part of the same dual-use-safeguard apparatus Anthropic has attached to every frontier release since Fable 5.1.
Anthropic also reports an 85% reduction in boundary-circumvention attempts succeeding against Opus 5.5 versus prior models on its own automated behavioral audit -- a 1,900-plus-scenario internal test, not an outside evaluation -- alongside EU AI Act watermarking compliance and a zero-data-retention option for enterprise customers. None of that is new machinery invented for this release; it is the same verification-program structure Anthropic built for Fable 5.1 and Opus 5, now extended to the model that just became the company's most capable. The pattern is consistent even if the specific numbers aren't independently checked here: capability keeps climbing on Anthropic's releases, and the safety-testing apparatus attached to each one keeps growing alongside it rather than lagging behind.
The Intelligence Index, before and after
| Claude Opus 5.5 max, new | GPT-6 Astra max | Claude Fable 5.1 max | Claude Opus 5 max, superseded | |
|---|---|---|---|---|
| Intelligence Index score | 58 | 53 | 53 | 51 |
| List price, input/output per 1M tokens | $4 / $20 | $10 / $50 | $10 / $50 | $5 / $25 |
GPT-6 Sol and Luna are a different kind of release entirely. Artificial Analysis's own published comparison states plainly that "Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others" -- Sol moved from 47 to 48, Luna held at 37. OpenAI's own framing matches that -- not claiming a capability leap, but an efficiency one. What actually changed is the bill: Sol's list price fell from $4/$20 to $2/$10 per million tokens, and Luna's fell from $0.20/$1.20 to $0.10/$0.50 -- roughly a 50-58% cut on both. Artificial Analysis's own measured cost to run its fixed evaluation set fell to match: $1.06 per task for Sol, down from $1.99, and $0.07 per task for Luna, down from $0.18.
“GPT-6 Astra introduced a new generation of intelligence; these models extend its benefits by making that intelligence more efficient and accessible.” — OpenAI, announcing GPT-6 Sol and GPT-6 Luna
OpenAI's own factuality claim for Sol -- roughly half as many mistakes as GPT-5.6 Sol, measured on the company's internal evaluation -- is worth flagging as exactly that: a company-reported figure, not one Artificial Analysis or any other independent evaluator has confirmed. It sits alongside the independently measured Intelligence Index score rather than inside it, and the two shouldn't be read as the same kind of evidence just because they arrived in the same announcement.
- Sol -- Intelligence Index score
- Sol -- price per 1M tokens (in/out)
- Sol -- Artificial Analysis's own cost per task
- Luna -- Intelligence Index score
- Luna -- price per 1M tokens (in/out)
- Luna -- Artificial Analysis's own cost per task
OpenAI's own comparative claim -- that GPT-6 Sol at xhigh effort beats Claude Opus 5 at max effort on AutomationBench for 9% of Opus 5's cost per task -- is real as far as it goes, but it was already dated the moment it published: Claude Opus 5 stopped being Anthropic's flagship the same day, superseded by Opus 5.5 hours after both companies' announcements landed. Neither company has yet published a same-day comparison of GPT-6 Sol against Opus 5.5 specifically, which is the actually current top-tier model this new pricing would have to compete against on capability, whatever it costs.
What each headline claim actually covers
- 58 · Intelligence Index score
- Claude Opus 5.5's new #1 position
Includes: An independent, third-party measurement (Artificial Analysis) against the same fixed evaluation set every other model on this board is scored against
Excludes: Anthropic's own "40% cheaper on typical workloads" and "30% faster" claims, which are the company's own figures, not independently verified here - $1.06 vs $1.99 · measured cost per Intelligence Index task
- GPT-6 Sol's real, confirmed price advantage
Includes: Artificial Analysis's own metered spend running its fixed evaluation set on both GPT-6 Sol and GPT-5.6 Sol
Excludes: OpenAI's own factuality claim ("about half as many mistakes"), which comes from OpenAI's internal evaluation, not an independent one
Read together, the two releases are the clearest one-day snapshot yet of how the frontier-model market is actually competing right now. Anthropic is competing on raw capability at the top of the board and using the resulting headline to also justify a price cut; OpenAI is competing on cost at the tiers below the flagship, holding capability flat and selling the savings directly. Both are legitimate strategies, and this same week has already produced two more data points in the same pattern -- xAI shipping Grok 4.7 mid-pack on score but priced a fifth of the frontier, and a StepFun model this desk found while researching this piece tying Kimi K3's score at roughly a third the measured cost per task. (None of this year's price cuts have come with a corresponding cut in what a real task actually costs to complete end-to-end once a real agent's tool calls, retries and context are counted -- the per-million-token rate and the per-task rate can move in different directions, which is exactly why Artificial Analysis publishes both.) The number worth tracking next isn't a fresh score -- it's whether Sonnet 5.5 and Haiku 5.5, still "coming weeks" out per Anthropic, extend the same capability-per-dollar gain down into the tiers most production traffic actually runs on.
- Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Sol/Luna both shipped Sept. 22.
- Opus 5.5 scores 58 on Artificial Analysis's index, a new #1, 5 points clear of the prior tie.
- GPT-6 Sol and Luna barely moved on capability but cut list prices roughly in half.
- Independent cost-per-task data confirms OpenAI's price story; Anthropic's is a real score jump.
- Caveat: OpenAI's own cost comparison used Claude Opus 5, already superseded hours later by 5.5.