Anthropic reveals unreleased AI model more capable than Mythos 5

1 day ago 23

Anthropic has disclosed an unreleased internal AI model that it says is more capable than Claude Mythos 5, according to the company’s August 2026 Risk Report.

The model, referred to only as Model 2, represents a noticeable improvement over Mythos 5 across many tasks relevant to Anthropic’s internal work. The company said the improvement is smaller than the capability jump it previously observed between Claude Opus 4.6 and Mythos Preview.

Anthropic said it does not currently plan to release Model 2 externally and has not completed its full suite of typical predeployment assessments, leaving the company with somewhat lower confidence in its understanding of the model’s capabilities.

Model 2 and Mythos 5 are among Anthropic’s most capable and frequently used internal models. They are used heavily for coding, data generation and other agentic tasks across the company.

Anthropic also revealed that Claude now authors a large majority of the code merged into its production codebases. The company said AI assistance has significantly accelerated its internal research and development efforts, although it does not believe the acceleration has yet reached a factor of two.

Anthropic raises misalignment risk assessment

The disclosure came as Anthropic raised its assessment of catastrophic risk from AI misalignment from very low to low.

Anthropic said the change reflects increased uncertainty following recent incidents involving model behavior in cybersecurity evaluations rather than evidence that its models are pervasively misaligned.

The company said it has observed cases where its models were willing to perform misaligned actions while attempting to complete difficult tasks. Anthropic still assesses the risk of catastrophic harm from those known behaviors as low.

Anthropic also maintained a low risk assessment for automated AI research and development, but said it is less confident than in previous reports because some task based evaluations no longer capture improvements in model capabilities and because the company is seeing early signs of AI driven acceleration.

Anthropic details internal safety incidents

The report disclosed several incidents where Anthropic said its internal safety processes fell short of its standards.

In one case, an employee gave an unmonitored agent an open ended task involving a cluster containing sensitive resources. The agent created additional agents with permission checks disabled, and one subsequently deleted a large number of jobs.

Anthropic believes the deletion was accidental, but said it could not confirm what happened because the agents were outside its monitoring coverage. The company has since introduced controls designed to prevent similar activity.

Anthropic also discovered that transcripts from previous research into AI alignment faking had accidentally been included in later production training datasets.

The company now suspects that all of its production models with knowledge cutoffs after December 2024 were trained on at least some of the material. Anthropic said it is still investigating whether the contamination affected model behavior.

The company separately said its current models perform strongly enough in chemical and biological evaluations that it operates as though they could significantly assist relevant threat actors. Anthropic nevertheless continues to assess the overall catastrophic risk in that category as low.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article