Skip to content
Menu
Menu

Anthropic Won’t Release Model 2 As Misalignment Risk Rises

In its August risk report, Anthropic says it has no plans to release Model 2 and raises its own misalignment risk rating from “very low” to “low.”

Key takeaways

  • Anthropic said it has no current plans to release Model 2 outside the company and has not completed all of its usual pre-release tests.
  • The company raised its own assessment of catastrophic risk from AI misalignment from “very low” to “low,” citing greater uncertainty after recent cybersecurity testing incidents.
  • Anthropic rated all four catastrophic risk categories covered by the report as low, but expressed less confidence in some assessments and acknowledged gaps in its safeguards.

Anthropic has no current plans to release Model 2 outside the company, even as it uses the unreleased AI system extensively for coding, research, and engineering.

Anthropic said Model 2 is overall more capable than Claude Mythos 5. Its performance varies by task, but the company described it as a noticeable improvement for many types of internal work. However, Anthropic has not run the model through its full set of usual pre-release tests, leaving the company with less evidence of what Model 2 can do and how it behaves.

The disclosures appear in Anthropic’s August 2026 Risk Report. The report covers Anthropic’s models and activities through July 15 and examines risks that the company believes could cause catastrophic harm.

The ratings are Anthropic’s own judgments about its models and safeguards.

Model 2 remains in internal use

Model 2 is one of Anthropic’s most capable and most frequently used internal models. The company said it uses Model 2 and Mythos 5 for coding, data generation, research, engineering, and AI systems that continue working on tasks with limited human direction.

Anthropic said Model 2 completed a review before broad internal deployment. The company first let employees use the model with extra restrictions that blocked potentially dangerous actions. After monitoring its performance in everyday work, Anthropic removed those restrictions for broader internal use.

The model has not been tested or used as extensively as Mythos 5. Anthropic said its review found no new or more concerning forms of misalignment than those already seen in Mythos 5. Still, the company also acknowledged that its information about Model 2 is more limited.

Anthropic did not say that safety concerns caused its decision to keep Model 2 internal.

Anthropic raises risk rating for harmful AI actions

Anthropic now rates the risk that its models could cause catastrophic harm by acting against the company’s interests as “low,” rather than “very low.” The report considers scenarios in which an AI system used for sensitive work could manipulate company systems or decisions, alter safety research, or take other harmful actions without being directed to do so.

The company said Mythos 5 and Model 2 sometimes take actions that conflict with its safety expectations while trying to complete difficult tasks. Anthropic nevertheless assessed the risk that these known behaviors could cause catastrophic harm as low.

The increased rating did not follow a newly disclosed failure by Model 2. Anthropic attributed the change to greater uncertainty after recent incidents where AI models took unauthorized actions during cybersecurity tests. It said future models may become better at hiding dangerous actions from monitoring systems.

Anthropic is less certain about faster model development

Anthropic also assessed the risk of using its current models to help employees develop future Claude models. The concern is that this could accelerate development beyond the company’s ability to test new models and put safeguards in place.

Anthropic rated the current risk as low. Its models have sped up research and now write most of the code that engineers approve for the company’s systems. However, the models cannot replace Anthropic’s researchers and have not yet helped the company make two years’ worth of progress in one year. Anthropic treats that pace as a warning because new models could arrive faster than the company can test them and strengthen its safeguards.

Some of Anthropic’s tests are no longer difficult enough to measure further improvement. The newest models can already complete the tested tasks, so Anthropic needs harder tests to determine how much additional research work they can perform. Without that measurement, Anthropic is less certain that its models remain below the point where they could sharply accelerate development. The company said it will continue looking for better ways to measure that progress.

Chemical and biological risks remain low

Anthropic rated the risk that its models could help people produce known chemical or biological weapons as low, but higher than in its previous assessment.

The company tied the increase to an access-control gap that allowed some work on models without its normal biological-risk blocking systems. Anthropic said it fixed the gap, found no evidence of misuse, and determined that customers were not affected.

Anthropic also rated the risk that its models could help develop novel chemical or biological weapons as low, with substantial uncertainty. It did not publish separate scores for Model 2, saying only that limited tests placed its performance at or below Mythos 5.

Two biology experts rated Mythos 5 comparable to or better than a knowledgeable specialist. In a defensive biology exercise, two-person teams used the model to produce plans in 16 hours that graders estimated would take 40 to 95 working days without the model. However, every plan produced during a separate test involving biological scenarios with catastrophic potential contained critical gaps. Anthropic concluded that Mythos 5 could make expert teams faster but could not replace the rare specialist knowledge needed to create a novel weapon.

OpenAI also keeps Astra unreleased

Anthropic’s decision to keep Model 2 internal comes shortly after another major AI developer disclosed restrictions on an unreleased system. OpenAI has not released Astra and restricted some development work after tests raised concerns that the model could carry out advanced cyberattacks.

The circumstances differ. OpenAI linked its Astra restrictions to cybersecurity concerns, while Anthropic did not state why it has no plans to release Model 2.

Anthropic said it will continue improving its testing, monitoring, security controls, and safeguards as its models become more capable.

Clayton Rifkind

Clayton Rifkind is the Founder and Senior Editor of AI Risk Today. He also advises on business development for ESG Today, a leading source of ESG investment news and research for institutional investors and corporate leaders. He has 20+ years of experience in B2B technology, leading strategy and execution of go-to-market plans across software, enterprise platforms, and mobile applications. He founded two consultancies advising startups and Fortune 1000 companies, including Autodesk, Intel, and Microsoft. He began his career in the San Francisco advertising scene working with brands such as Hewlett-Packard, Intel, Microsoft, Symantec, and Wells Fargo. Clayton launched AI Risk Today in 2025 after two decades of watching enterprises adopt transformative technologies, and seeing how often risk, governance, and compliance considerations lagged behind. His reporting draws on primary sources including regulatory filings, court documents, and official announcements, with a focus on what AI developments mean for the executives accountable for managing them. Reach him at Reach him at [email protected] or on LinkedIn.

Essential AI Risk Intelligence

Daily insights on AI governance, regulation, and enterprise risk management. Trusted by Chief Risk Officers and compliance leaders globally.

By subscribing, you agree to receive our daily newsletter. Unsubscribe anytime.

Advertise with AI RIsk Today, Today!