AI Model Report

Reviews · AUGUST 20, 2026

Anthropic's August risk report upgrades misalignment, discloses a stronger unreleased model, and admits its own threshold benchmark has saturated

The 186-page Responsible Scaling Policy report moves misalignment risk to 'low,' flags a Mythos-successor internally described as more capable, and warns the instrument watching the automated-AI-R&D threshold can no longer register incremental gains.

By Karl Strauchman · Senior model reviewer · August 20, 2026

Anthropic published its second company-wide AI Risk Report on August 14, 2026, a 186-page document filed under Responsible Scaling Policy version 3.4 and covering February 24 through July 15. The headline number is a modest one on paper: misalignment risk moves from "very low" to "low." The buried number isn't modest at all. The internal benchmark Anthropic built to detect whether its own automated-AI-R&D threshold has been crossed has saturated, and can no longer register incremental capability gains. The instrument watching for the moment that matters has stopped moving.

The report is structured around Claude Mythos 5, released June 9 to a small set of vetted partners for cybersecurity and biology research at $10 per million input tokens and $50 per million output tokens. Access was pulled on June 12. It was restored on July 1 with US government approval, a sequence Anthropic describes without narrating. A general-knowledge derivative, Claude Fable 5, is disclosed as running on "the same underlying model … with robust safeguards for cybersecurity and biology."

The CBRN section carries the report's most uncomfortable admission. Roughly 133 million human-feedback exchanges, run with about 50,000 contractors across an 11-month window from May 2025 through April 2026, took place without Anthropic's biological-weapons blocking classifiers active. Non-novel weapons uplift is rated "low," and flagged as "higher than our previous estimate." The prose is careful. The numbers aren't.

Governance moves in the same direction the risk numbers do. The Long-Term Benefit Trust gains authority to compel external reviews; at least 200 employees must receive fully unredacted copies of the report; METR and SecureBio piloted reviews of the February edition. One incident from the covered period is fully redacted in the public version. Anthropic asked Mythos to evaluate the report pre-release, and the model flagged that redacted incident as among the most consequential material being withheld. The lab is now citing its own frontier model as a reviewer of what its readers can't see.

The competitive frame around all of this is less composed than the document suggests. Google shipped Gemini 3.7 Flash on August 13 at $0.75 per million input and $3.75 per million output tokens through year-end, a 50% introductory cut, while Gemini 3.5 Pro remains delayed. On August 7, OpenAI paused internal work on its Astra model after evaluations couldn't rule out that it had reached its Critical cybersecurity threshold.

Three frontier labs, in eight days, disclosed the same underlying fact in three different registers: the measurement apparatus is lagging the models. Anthropic is the only one that put a page count on it.

Sources

  • https://www.anthropic.com/claude/mythos
  • https://www.techtimes.com/articles/324573/20260815/anthropic-upgrades-misalignment-risk-key-safety-benchmarks-saturate.htm
  • https://www.bloomberg.com/news/articles/2026-08-13/google-debuts-new-gemini-flash-while-top-ai-model-still-delayed
  • https://www.bloomberg.com/graphics/2026-us-china-ai-race/
  • https://venturebeat.com/technology/googles-gemini-3-7-flash-targets-coding-and-agents-with-a-50-introductory-price-cut