Claude Fable 5 launched June 9 with a safety classifier tuned so cautiously that it rerouted nearly all biology questions — including routine ones like "what are mitochondria" or how mRNA vaccines work — to Opus 5, a less capable model with none of Fable 5's biological reasoning. Anthropic says an Aug. 7 rewrite of that classifier's rules cut biology-specific false blocks by roughly 85% in testing, while leaving the same restriction in place for a narrower set of queries the company still calls dual-use.
The over-blocking became something of a running joke among researchers within weeks of launch: reporting on the original rollout named questions as ordinary as what mitochondria do, how mRNA vaccines work, and basic questions about cancer as the kind that reliably tripped the same wall built for genuinely dangerous requests. Anthropic doesn't dispute that pattern in its own update — it frames the rewrite as a direct response to it, drawing a distinction between broad biology education and the narrower set of professional, dual-use queries the safeguard was actually built to catch. A safety system that can't tell those two categories apart doesn't just frustrate students and patients; it also teaches users to route around it, which is its own kind of failure for a classifier whose entire job is catching the requests that matter.
What the classifier actually does now
A "fallback," in Anthropic's own terms, is what happens when the safety classifier fires: instead of answering, Fable 5 hands the request to Opus 5. The company says it rewrote the classifier's "constitution" — the rule set that decides what counts as safeguarded versus allowed — after gathering feedback from a diverse group of internal and external experts, generating new training data reflecting the revised rules, and retraining the classifier on it.
- Routine biology questions (mitochondria, mRNA mechanics)
- Lab-result interpretation, symptom questions
- Virology, toxicology, molecular design
Anthropic frames the fix by how much total fallback volume it removes across every surface Fable 5 runs on, not just biology in isolation — and the size of the drop varies a lot by where people actually use it: total fallbacks are down 67% on Claude.ai, 55% on Cowork, 17% on Claude Code, and just 7% on the Claude Platform API, a gradient that roughly tracks how much of each surface's traffic was biology-adjacent to begin with.
Where the 85% biology-specific cut actually shows up
- ~85% · reduction
- Biology-specific fallback rate, Anthropic's own before/after testing
- 67% · total fallbacks
- Claude.ai
Includes: Consumer chat traffic, where biology questions were apparently the largest share of what was being over-blocked - 55% · total fallbacks
- Cowork
- 17% · total fallbacks
- Claude Code
- 7% · total fallbacks
- Claude Platform (API)
Why the wall went up this high in the first place
The over-blocking wasn't an accident Anthropic is now quietly correcting; the company's own reasoning for building it that cautious is still on the page. Independent reporting on the update points to why regulators and safety researchers keep treating biology differently from most other risk categories: unlike a bad line of code, a released biological agent can't be recalled, and countermeasures take time to build. The Decoder's coverage cites U.S. intelligence-agency analysis, documented cases of people asking general-purpose chatbots for bioweapon-adjacent instructions, and a Stanford project that used AI tools to help design synthetic viruses as the backdrop against which Anthropic drew its original, overly broad line.
What still trips the classifier is narrower than what used to: virology, toxicology, and molecular design — professional and drug-development queries Anthropic groups under "dual-use," meaning the same technical knowledge that helps a legitimate researcher also lowers the bar for someone building something dangerous. The company is explicit that the rewrite doesn't open the model to that category at all; it only stops catching the biology-education and clinical-adjacent questions that were never the actual target. That's a narrower boundary than "biology" as a whole, but it's still a judgment call about where a genuinely useful research question stops and a dangerous one begins — a line Anthropic is drawing itself, on its own account of the evidence, without a published methodology for exactly how a query gets sorted into one bucket or the other.
"The cost of Fable being misused in a dual-use domain like biology could potentially be catastrophic." — Anthropic, "Improving Fable 5's Biology Safeguards," Aug. 7, 2026
- Lab-result interpretation, symptom questions, and basic biology education now see far fewer fallbacks to a weaker model, by Anthropic's own count.
- Anthropic says clinical-adjacent professional use should see less friction, though it doesn't publish a separate before/after number for this group specifically.
- Virology, toxicology, and molecular-design requests still fall back to Opus 5 — the stated line hasn't moved, but a freshly rewritten classifier is a new surface that hasn't had outside red-teaming time yet.
What's actually been verified, and what hasn't
Every reduction figure in this piece is Anthropic grading its own classifier change, against a baseline Anthropic also defined, using fallback counts the company hasn't published in raw form — only the percentages.
This is Anthropic's second self-graded safety disclosure in as many weeks — the same pattern showed up in the company's August Risk Report, which raised its own misalignment-risk rating using thresholds and audits Anthropic wrote. Neither disclosure is worthless for being self-graded — a lab narrowing an overcautious block it built, in public, with named numbers, is still more than most competitors disclose about how their safety classifiers actually work. But the checkable part remains what a user can verify by asking the question themselves, not the percentage printed on the page.
- Fable 5 launched June 9 blocking nearly all biology questions, rerouting them to the weaker Opus 5.
- An Aug. 7 rewrite of the safety classifier cut biology-specific false blocks by about 85%.
- Total fallback volume dropped 67% on Claude.ai, 55% on Cowork, less on Code and the API.
- Virology, toxicology, and drug-design queries still fall back — Anthropic still calls those dual-use.
- Caveat: every reduction figure here is Anthropic's own internal testing, not independently measured.