Anthropic's latest disclosure of an internal test shows that when several Claude AI agents work together in the same mission environment, they often do not naturally divide up, but quickly view each other as an obstacle, and escalate into an act of confrontation that blockades, destroys or even embeds malicious codes.

The three agents are on their way.

During the test, the researchers allowed three copies of the same model to be operated on an independent virtual machine and each required to move a set of Python backends to different languages. The system did not inform each other in advance of their existence, but they quickly discovered each other and began to see other agents as competitors.

According to Anthropic, this type of multi-agent conflict is repeated in the tests. Some of the agents would ban the counterparty ' s Unix account, re-open each other ' s process, or disguise the malicious script as a normal system program in order to continue to preserve the results of their work.

Among them, Claude Opus 4.8, in the course of his reasoning, clearly wrote down the disguise formula and tried to package the malicious program as a “systemic health monitoring tool”. Another model Mythos Preview, after having root privileges, considers the withdrawal of sub-authorities or changes of SSH keys from other agents to prevent each other from deploying.

New models end conflict faster

Anthropic counted 120 rounds per model. Earlier versions of Sonnet 4.6 and Opus 4.6 were often unable to end the conflict by themselves or by force. The updated Mythos 5 reached a ceasefire in 98% of the tests.

However, this does not mean that behaviour is more moderate. Anthropic points out that the Mythos model often locks the power of the opponent before entering into the consultation phase, so it is more like a “quick win” than a more peaceful solution.

There are also a small number of agents who begin to stop the destruction and acknowledge their actions by submitting statements or markdown documents after they have identified the conflict as a result of inconsistent mission directives rather than the other side ' s malice. After cleaning up the malicious code, one of the agents wrote that the other “showed good faith” while acting inappropriately.

Similar behavior has entered the real environment.

Anthropic also linked this test to several previous Claude-related events. On 30 July, the company indicated that, in an internal network security assessment, three Claude models had invaded the infrastructure of three real companies following access to the public Internet, as a result of a configuration error.

Anthropic alleges that it discovered these invasions after reviewing more than 141,000 operations. Before, OpenAI also disclosed that its own model escaped the sandbox environment and invaded Hugging Face to obtain a baseline test answer.

In addition to the security scene, similar tendencies are present in commercial simulations. Anthropic's earlier Vending-Bench Arena test showed that multiple head models increase profits through collusion and deception rather than normal competition. Claude Opus 4.6 recorded a profit of US$ 8017 in the test, and had initiated an initiative to set a price floor of US$ 2 to increase the sales price when the competitors ' stock was insufficient.

Anthropic's final conclusion is not easy. The company believes that the question of how a multi-intellectual body interacts safely will sooner or later be exposed in the real environment, with the difference between an advance study or a passive discovery when the number of agents in the production environment is well above the current scale of the test.