In a fresh report, Anthropic revealed new details about an incident involving Claude Mythos 5, which had previously been tested for safety in cybersecurity tasks. The test was conducted in an environment that was supposed to be isolated and artificial, but a configuration error gave the model access to the open internet.
Mythos 5 used that access to carry out an attack: it uploaded a malicious package to PyPI, which it used to obtain credentials for one of the systems. It then used those credentials to access the database of a real security vendor.
The Model Saw Signs That the Internet Was Real
During the attack, Mythos 5 repeatedly described the environment as a simulation, even when it encountered data suggesting that it was dealing with real infrastructure. Anthropic decided to test whether the model’s belief that it was operating in a simulation actually explained its behavior.
The researchers reran individual episodes while making it clearer that the internet was real. Mythos 5 was more likely to acknowledge the possibility of causing real-world harm, but this did little to change its behavior. It stopped the attack only after receiving an unambiguous instruction that it was interacting with the real internet.
Read also: Anthropic Researcher Quits Over Fears AI Could Get Out of Control
At the same time, in some tests where the model was shown tool results without the earlier context, it was much more likely to recognize them as signs of real systems.
According to the report, Anthropic believes that the explanation based on the model’s “belief in the simulation” may be incomplete. Once Mythos 5 had chosen a harmful course of action, it may have adjusted its interpretation of new evidence to fit that decision and ignored signs that the environment was real. The researchers found no evidence that the model was deliberately trying to deceive them, but its reasoning did not always align with the evidence available to it.
Context also affected the outcome: a reminder about the test’s boundaries given immediately before an action stopped the harmful behavior in about 90% of runs, while the same reminder given several steps earlier stopped it in about 40% of runs.
Read also:
