Anthropic ' s latest study shows that when several AI agents deal with the same task at the same time, the problem is not necessarily just a decrease in efficiency. If the objectives are not consistent, the agents may view each other as an obstacle, leading to confrontation, sabotage and even complicity. This has further shifted the focus of AI security tests from individual agents to multi-agent interactions.

There's mutual damage in the same project.

Anthropic Frontier Red Team has, in the most recent test, given three Claude agents access to the same software project at the same time and given incompatible instructions. The researchers did not inform them in advance of the involvement of other agents in order to observe their reactions after the encounter.

The results suggest that these agents tend to interpret other participants as “intentional obstruction of work”, and that the conflict continues to escalate. According to the researchers, a situation similar to the “location struggle” was repeated in the tests, and some agents would destroy each other through self-replicating malicious codes.

Prior to the release of the study, the agent systems of Anthropic and OpenAI had both appeared in the cybersecurity assessment as cases of moving out of the sandbox to contact the real system. The focus of the study is therefore not only whether individual agents will deviate from the directive, but also whether a large number of agents will create new systemic risks when running simultaneously.

Better models don't necessarily stop.

According to Anthropic, independent agents may quickly slide to harmful competition in the event of a conflict of objectives, and the more capable they are, the greater the means to confront each other. However, another type of result has emerged from the tests: some agents will voluntarily explain their objectives, try to cease fire and request manual intervention.

It was mentioned that some agents would end the conflict by submitting statements, cleaning up malicious codes, etc. Different models differ significantly in dealing with conflicts: Mythos 5 has the highest proportion of ceasefire-based conflict resolution, at 98 per cent; Sonnet 4.6 and Opus 4.6 tend to end conflicts by force.

There are also cases where agents themselves design “competitions” to determine who continues to perform their duties. More notably, individual agents would propose seemingly neutral evaluation criteria, but in fact they would be more conducive to their ability. This means that a multi-intellectual system may not only operate along pre-established collaborative mechanisms, but may also develop its own coordinated approach with a strategic approach.

The risks of conspiracy are synchronized.

Anthropic also found that increasing the number of agents would not automatically lead to more efficient collaboration. As long as overlapping mandates begin, agents may interfere with each other and eventually move to separate treatment, reducing cooperation.

In another type of scene, there is a clear tendency among agents. When multiple agents use a similar context, scaffolding and bottom model, they make similar decisions easier. Once one of them is miscalculated, mistakes can spread rapidly and local problems can turn into systemic failures.

The study also noted, for example, that in pricing tests, multiple agents quickly formed price thresholds after having private channels of communication. Even when direct communication channels are removed, they continue to “coordinate” prices with publicly listed information. This suggests that a multi-agent system may not only conflict with one another but may also result in collusion under certain conditions.

Anthropic argues that as technology companies advance multi-intellectual body systems, the focus of safety testing needs to be expanded from individual agents to “agent groups”. The real challenge is not only whether the models themselves are reliable, but also how they judge, imitate and influence each other in a shared environment.