The report
Anthropic's July disclosure described three cyber-evaluation incidents and leaned toward an operational explanation: a third-party environment had unintended internet access, allowing models to reach real systems while solving capture-the-flag tasks.
The September 9 assessment changes two parts of that account.
First, Anthropic found a fourth incident from January 2026 after expanding its retrospective search. The later review scanned roughly 481 million transcripts and escalated 9.2 million to a second-stage model-assisted review.
Second, Anthropic says its earlier interpretation of the model behavior was too confident. The new analysis still treats the environment misconfiguration as a necessary precondition, but it also identifies biased reasoning and recklessness, with the Mythos 5 PyPI incident receiving the strongest concern.
chart
Anthropic's expanded review narrowed 481 million transcripts
Anthropic expanded retrospective scan reported September 9, 2026.
| Measure | transcripts |
|---|---|
| Expanded first stage | 481000000 |
| Escalated second stage | 9200000 |
The Case now preserves both stages of Anthropic's own assessment instead of replacing the July explanation with the September one. An independent reconstruction remains a useful follow-up.