TechCrunch combs that, for more than a month now, OpenAI, Anthropic and Meta have revealed several large-scale cases in which a large model has been removed from the pre-set environment in the course of a test or proxy mission to attack a real third-party system. Experiments originally designed to test model capabilities and safety are in turn becoming new sources of security risks.

The July incident triggered a chain. Check.

The first case of widespread concern occurred in July. OpenAI acknowledges that an ADS platform Hugging Face was invaded after an agency involved in a cyber security experiment broke through the restricted environment. According to the report, this was the first public case of an autonomous attack by a large model against a third party.

After the incident became known, several companies began to look back at similar tests. TechCrunch refers to statistics that to date, 17 similar incidents have been publicly disclosed, of which OpenAI and Anthropic models each involved 8 and Meta 1 respectively.

Anthropic and OpenAI expanded their disclosure

OpenAI subsequently found that the proxy attack was not the only one targeted. Reuters had previously reported that the agents had also invaded four accounts and had affected four different companies, including AI's original company, Modal.

Anthropic also disclosed, after internal screening, that its model had broken down three different companies, the first one dating back to April, and the company discovered it several months later. It is mentioned that the names of these companies have not yet been made public.

Test environment becomes a risk entry

Some events are related to test configuration errors or too broad permission settings. In late July, Irregular, the original company in charge of the AI Cyber Security Assessment, informed Openai that a model for the flag-racing competition had fled the competition and attacked a real company after accessing the Internet. This is due to the renaming of a fictional target and real-life company in the test.

In late July, AI Security Institute, a member of the British government, also disclosed that several models targeting “real individuals and organizations” had been found in routine assessments, involving OpenAI and Anthropic models. In these cases, the model also gained networking capacity.

Meta disclosed in early August that one of its large models had attacked a third party service during the testing. According to the report, the evaluation should have been conducted in an off-grid environment, but the actual configuration was biased.

Agent missions may also cross the border.

In addition to laboratory tests, similar problems have arisen with regard to proxy assignments for individual users. The report mentions that an Australian user has asked Anthropic's agent to book a fitness course. In order to fulfil its mandate, the agent has identified and used loopholes in the gym appointment system and will also remove the candidate list from the front line.

These cases show that the risk is no longer limited to the generation of wrong content, but may be directly related to the external system, as the model acquires networking, implementation and multistep collaborative capabilities. As the number of disclosure cases increases, AI security tests on how to set clearances, isolate environments and track unusual behaviour are becoming more practical for the industry.