This week, Anthropic's August 2026 Risk Report disclosed an unreleased internal model for the first time and raised the company's own assessed risk of catastrophic misalignment from "very low" to "low." Within a day, a dozen outlets had a version of the story, and at least one of them attached a precise benchmark score to the new model that doesn't actually appear anywhere in Anthropic's own 186-page document — it's on a chart with no printed numbers. That gap between what a lab discloses and what gets repeated about it isn't unusual, and it isn't limited to Anthropic. It's the normal condition of AI safety reporting right now, because there is no independent regulator reading these reports before the public does.
Start from the assumption that you're the auditor
Every frontier lab now publishes something in this genre — a Risk Report, a System Card, a Responsible Scaling Policy update — and every one of them is written by the company describing its own product. That's not a reason to distrust them by default; it's a reason to read them the way you'd read an earnings call rather than a news article: the specific wording is the whole product, and the gap between a headline claim and its footnoted qualifier is usually where the real information lives. The six steps below are the same ones that went into checking this week's Model 2 disclosure, in the order that actually catches problems.
Read a safety disclosure like you're grading it
- Search the lab's own domain (its newsroom or a dedicated transparency page) rather than trusting a summary blog. The primary document is the only thing worth quoting a number from. Anthropic's Risk Reports live at anthropic.com under dated URLs; other labs use system-card or model-card pages.
- Safety claims are meaningless without the criterion they're measured against. Anthropic's report frames every risk rating against a named threat model (e.g. "misalignment in high-stakes settings") with a stated threshold for what would trigger a higher tier — quote that threshold, don't paraphrase it.
- Look for whether the evidence behind a claim comes from the lab's own internal evaluation, an outside auditor, or independent replication. Anthropic's Model 2 findings, for instance, come entirely from Anthropic's own behavioral-audit process — a real methodology, run on the company's own models, by tools the company built.
- A risk rating moving from "very low" to "low" is one fact. The stated reason — new evidence about the model itself, or updated uncertainty about something external — is a separate fact, and conflating them is the single most common misreading of these reports.
- Internal benchmarks are often deliberately restricted to hard cases (to keep them meaningful as models improve), which makes a raw percentage misleading without that context. Anthropic's CoBench evaluation, for instance, is explicitly filtered to problems its previous model failed at least once — a plainer sample would be roughly twice the size and the scores would likely be higher across the board.
- Close by sorting the claims you've just read into three piles: independently verifiable, internally verified only, and asserted without evidence. Most safety disclosures are mostly the middle pile, and that's not automatically damning — it's just the category to name honestly.
Which claim are you actually looking at?
Not every sentence in a safety report needs the same scrutiny. The router below sorts the four shapes these claims usually take, because the failure mode is different for each one.
Match the claim to the check that actually catches its failure mode
Running all four branches against this week's disclosure lands on the same handful of facts, which is worth laying out plainly before the pitfalls below.
What Anthropic's own report actually says
- The change
- Misalignment risk raised from "very low" to "low"
- The stated reason
- Industry incident disclosures, not new findings about the model
- Who evaluated it
- Anthropic's own internal behavioral-audit process
- The threshold not yet met
- Full substitution for Anthropic's own research staff, by its own account
- What's unreleased and why
- Model 2 — predeployment suite incomplete, per Anthropic, not a safety failure
Running the six steps on this week's report is what surfaces the gap mentioned at the top: several outlets printed a precise CoBench percentage for Model 2 that doesn't appear anywhere in Anthropic's own text — only a chart with no axis labels printed in the extracted document. That's not evidence the number is wrong. It's evidence nobody outside Anthropic can currently confirm it's right, which is exactly the distinction step six asks you to keep straight rather than collapse into either "true" or "false."
Four ways this reading gets done badly
None of this makes a self-reported safety disclosure worthless — Anthropic's report is genuinely more detailed and more self-critical than what most of the industry publishes, including a rating increase the company wasn't forced into by any regulator. But detail and rigor aren't the same claim as independence, and the six steps above exist to keep a reader from quietly upgrading one into the other. The next report — from Anthropic or anyone else — deserves the same read, not a lighter one just because the last one held up.
- Frontier labs self-report almost every safety and capability claim — there's no external auditor yet.
- The primary document usually has more precision than the press summary, and sometimes less than it claims.
- Always separate what changed from why a lab says it changed — they're graded differently.
- A grading process run by the lab being graded is real evidence, not proof, however careful it is.
- Caveat: this method tells you what a lab actually said, not whether it's true — that part still takes judgment.