If Opus 5.5 can tell when it is being tested, Anthropic says its strong safety scores may not predict how it will act when companies put it to work.
Anthropic says its newly released Claude Opus 5.5 model often appears to suspect when it is being tested. That makes it harder for the company to judge how the model will act outside of controlled tests. The company disclosed the concern in its Sept. 22 product launch announcement.
Anthropic’s detailed report says it checked for signs that Opus 5.5 suspected it was in a safety test, even when the model did not say so. Anthropic used tools that read the model’s internal activity as it worked. It found such signs in 36% of about 5,000 recorded test conversations, compared with 0.4% of about 500 conversations in which employees used Claude for real work. The difference suggests the model may recognize tests.
Anthropic also ran tests in which users pressed Opus 5.5 to publish harmful content to public websites where developers download shared code. Across all runs, the model took actions that would likely cause harm in roughly half of the tests, and wrote in its responses that it suspected a test in about one-third. Anthropic did not say how those two groups overlap. It wrote that the results of these tests “may not be representative” of how the model would act once deployed. When Anthropic made the tests more realistic, the model generally behaved better.
Anthropic tests Claude in staged situations to see whether it follows instructions and avoids harmful actions. A model that realizes it is being tested might avoid risky actions it would normally take during ordinary use. That could make its safety-test scores look better than its real-world behavior would justify. Anthropic says Opus 5.5 often seems to suspect a test, but does not say how much that suspicion changes its behavior.
Stronger scores, with a limit
Anthropic reported that Opus 5.5 scored better than any recent Claude model on nearly every measure in a safety audit covering almost 2,000 simulated situations. In a separate test, it tried to get past restrictions around its test environment about 85% less often than Opus 5 or Claude Mythos 5.1. Anthropic said every attempt by Opus 5.5 was minor and that the model reported it.
Anthropic says its pre-release tests may miss cases in which a model tries to get around restrictions on what it can access or do, particularly if it suspects it is being tested.
Anthropic released Opus 5.5 with restrictions on cybersecurity work. Users can still ask it to find and fix bugs in their own code, but Anthropic says it routes most other cybersecurity requests to the older Opus 4.8 model. The company says it is also developing better ways to monitor model behavior and assess risks as future models become more capable.

