Anthropic Raises Its Own Risk Rating
A lab publishing a worse number about itself is unusual enough to be worth reading carefully. The reason for the change is not what most coverage implied.
Founder, Automation Squad ·
The short answer
Anthropic published its second company-wide Risk Report on August 14, 2026. It raises the company's assessment of the risk of catastrophic harm from misalignment in high-stakes settings from "very low" — the rating in its first report in February 2026 — to "low". The report also discloses an unreleased internal model called Model 2, described as somewhat more capable than its frontier model Mythos 5, which Anthropic says it has no current plans to release externally. The report states that Model 2's internal approval surfaced no new or more concerning form of misalignment beyond the profile already discussed for Mythos 5.
Reading the report without the headline distortion
Two separate facts landed in one document and most summaries welded them together. Keeping them apart is the whole value of reading it properly.
Read the primary document, not the aggregation
This story was widely rewritten into "Anthropic's secret model is scary", which is close to the opposite of what the report says about Model 2 specifically. The report is public and short enough to read yourself.
Separate capability disclosure from risk assessment
The interesting fact is that a lab disclosed an unreleased frontier-class model at all. The separate interesting fact is that it revised its own risk rating upward. Conflating them produces a story neither fact supports.
Note what a rating change like this is for
Self-published risk ratings are a governance artifact. Their value is that they create a public record you can hold a company to later — which only works if people read the second one against the first.
| The claim | What the report actually says |
|---|---|
| "Model 2 is dangerous, so the rating went up" | No. The report says Model 2 surfaced no new form of misalignment. |
| Catastrophic misalignment rating | Raised from "very low" (Feb 2026) to "low" |
| Model 2's capability | Somewhat more capable than Mythos 5 on many internal tasks |
| Model 2's release status | No current plans to release it externally |
| Who publishes a rating like this | Very few labs. That is part of why it is worth reading. |
Anthropic published its second company-wide Risk Report on August 14, 2026. The headline change is that it now rates the risk of catastrophic harm from misalignment in high-stakes settings as "low", up from the "very low" it assigned in its first report in February.
The facts: the report also discloses an unreleased internal model called Model 2, described as somewhat more capable than the company's frontier model Mythos 5 for many tasks relevant to internal work, and states that Anthropic has no current plans to release it externally. On the question everyone jumped to, the report says Model 2's internal approval surfaced no new or more concerning form of misalignment beyond the profile already discussed for Mythos 5. Reporting has linked the rating change to uncertainty around recent cybersecurity evaluation disclosures rather than to Model 2 itself.
Automation Squad's take: two things landed in one document and most of the coverage welded them into a story neither supports. "Lab hides scarier model" is a better headline than "lab publishes a slightly worse number about itself and separately mentions an internal build", but only the second one is what happened. The genuinely notable thing is the direction of the revision. Self-assessments published on a schedule tend to drift toward reassurance, because nobody enjoys writing down that their own risk estimate went up. This one went up. Whatever you make of the underlying number — and "low" is still low — a company creating a public record it can be held to next time is the part worth encouraging.
Run this now: read the report itself rather than the summaries, and read it beside the February edition, because the value of a recurring self-assessment is entirely in the comparison. If you are the person in your organisation who gets asked whether any of this matters, this is a short, citable, primary document from a frontier lab — which is a more useful thing to have in your pocket than another round of secondhand commentary.
Questions people are asking
- What exactly changed?
- Anthropic's assessment of the risk of catastrophic harm from misalignment in high-stakes settings moved from "very low" in its February 2026 report to "low" in the August 14, 2026 report.
- What is Model 2?
- An unreleased internal model, described in the report as somewhat more capable than Mythos 5 for many tasks relevant to internal work. Anthropic states it has no current plans to release it externally.
- Did Model 2 cause the rating change?
- The report says Model 2's internal approval surfaced no new or more concerning form of misalignment beyond the profile discussed for Mythos 5. Much of the coverage implied a direct link that the document does not make.
- Why should I care if I just use these tools for work?
- Directly, very little changes today. Indirectly, this is one of the few places a frontier lab publishes a structured self-assessment on a schedule, which makes it one of the few documents you can compare against its own previous edition.
Sources
Further reading
- Anthropic — Risk Report: August 2026 — The primary document. Short, and it says something different from most of the coverage.
Last checked August 18, 2026 against the primary sources above, by Automation Squad Research. Spot an error? [email protected].
