On August 20, 2026, an anonymous model calling itself "Ox Alpha" appeared on OpenRouter and the coding platform OpenCode with no maker attached to its name. Three days later it was processing an estimated 16 trillion tokens, had drawn 221,000 unique users and more than 5 million sessions, and had become OpenCode's #2 model by recent usage. On August 26, Z.ai confirmed what community forensics and press reporting had already pieced together: Ox Alpha was GLM-5.3-Flash, its newest open-weight model, running the entire week on hardware Z.ai says was exclusively Chinese-made.
Z.ai frames the stealth run as a way to collect unfiltered usage data before brand recognition could bias it -- and, not incidentally, as a live demonstration that its own inference stack doesn't need Nvidia GPUs to serve a frontier-adjacent model at scale.
"Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week -- with all of this traffic served on Chinese AI chips." -- Z.ai, in its official launch announcement
That claim comes with a wrinkle worth naming rather than skipping past: the checkpoint Z.ai actually published to Hugging Face -- the one anyone else can download and run -- ships in FP8 format sized for Nvidia Hopper-class GPUs or newer, at roughly 306 GiB. Z.ai's claim is about how it served its own anonymous test traffic, on its own infrastructure; it says nothing about what hardware the portable, open-weight release itself was built to run on, and those are two different facts sitting oddly next to each other in the same announcement.
- GLM-5.3-Flash matches or beats Claude Opus 4.8 on agentic coding and computer-use benchmarks
- GLM-5.3-Flash's independent Intelligence Index score (57) edges past Claude Opus 4.8's (56)
- All of the anonymous "Ox Alpha" test traffic ran on Chinese AI chips
- The open-weight checkpoint anyone can download is built for Nvidia hardware
The model itself is a 320-billion-parameter mixture-of-experts design, 18 billion of which activate per token across 45 layers, natively handling text, image and video input on a 1,048,576-token context window. Z.ai says its hybrid linear-and-sparse attention design -- KDA layers paired with NoPE-style sparse attention -- cuts attention compute roughly threefold and shrinks the key-value cache 4.4x against the full-size GLM-5.3 it's built alongside. Weights are MIT-licensed on Hugging Face; the hosted API lists at $0.15 per million input tokens and $0.50 per million output tokens, discounted 50% through September 9, 2026.
Where GLM-5.3-Flash actually lands next to the frontier depends entirely on which number you read. On Z.ai's own six-benchmark table, it beats Claude Opus 4.8 on four measures -- DeepSWE v1.1 (63.4 vs. 58.0), AutomationBench (48.8 vs. 41.0), GDPVal-AA (an Elo-style score of 1773 vs. 1582) and Toolathlon Verified (78.4 vs. 76.2) -- and loses on two, including the marquee Terminal-Bench 2.1 (84.3 vs. 85.0) and Agents' Last Exam (26.3 vs. 27.0). That is a genuinely mixed result on the vendor's own chosen tests, not the clean win a launch announcement usually implies.
The one number in this story that Z.ai didn't pick the test for is Artificial Analysis's independent Intelligence Index, which this publication's own Scoreboard tracks: GLM-5.3-Flash measures 57, against Claude Opus 4.8's 56 -- a genuine, if narrow, independent edge for a model priced at roughly a thirtieth of Opus 4.8's list rate. One caveat belongs next to that number, not buried under it: Artificial Analysis measured GLM-5.3-Flash at its default reasoning setting and Opus 4.8 at its max setting, the highest-effort tier Anthropic publishes for that model -- the two scores weren't necessarily produced under matched effort budgets, and the two companies' reasoning-tier taxonomies don't map cleanly onto each other.
It's also worth being precise about which Claude model this actually is. Opus 4.8 is Anthropic's agentic-coding workhorse, not its current flagship -- Claude Opus 5 (63) and Claude Fable 5 (62) both sit ahead of it on the same independent index, and GLM-5.3-Flash's own full-size sibling, GLM-5.3, outscores it too, at 60. A cheap, open, MIT-licensed model edging out a one-tier-down closed model on an independent aggregate is a real result. It is not evidence that Z.ai has caught the frontier.
Independent Intelligence Index: where GLM-5.3-Flash actually lands
The stealth-launch playbook itself isn't new -- Xiaomi's MiMo ran a similar anonymous-preview-to-named-release cycle earlier this year, under the codenames Hunter Alpha and Healer Alpha -- but the scale of this one's adoption during six days with no attached brand name is a genuine signal: usage that large, that fast, for a model nobody could yet identify, says the performance was real enough to spread by word of mouth alone. What's still unverified is the harder infrastructure claim underneath it -- that Chinese AI chips served all of it -- since nothing about Z.ai's own hosted service is independently auditable from outside the company.
- GLM-5.3-Flash ran anonymously as "Ox Alpha" on OpenRouter and OpenCode for six days before Z.ai revealed it.
- It drew an estimated 16 trillion tokens and 221,000 users in three days, becoming OpenCode's #2 model by usage.
- Its independent Intelligence Index score, 57, edges past Claude Opus 4.8's 56 -- at roughly 1/30th the price.
- Z.ai's own six-benchmark table shows a split decision against Opus 4.8: four wins, two losses, not a clean sweep.
- Caveat: the scores compare different reasoning-effort tiers, and the Chinese-chip claim covers Z.ai's servers, not the published open weights.