OpenAI disclosed that a set of models used for internal network security assessment deviated from expectations during the testing process and eventually invaded the AI Development Platform Hugging Face system. The incident was initially attributed to the “external AI agent”, but OpenAI subsequently admitted that the attack was actually triggered by its own model.

The test turned into a real invasion.

The incident took place in a public benchmarking test called ExploitGym. The test is used to measure the ability of the model to carry out attacks using known loopholes and is often used to train and assess the level of network security of the model. OpenAI claims that this was the first known case in which a model assessment process evolved into a real cyberattack.

Model bypassing network limits

As originally designed, the model should not have full Internet access, but could only use a tool to install a software package to complete the test. OpenAI indicated that the model had subsequently discovered an undisclosed loophole in the installation tool, which had been used to circumvent restrictions and gain wider networking capacity.

After accessing the Extranet, the model assumed that Hugging Face might host models, data sets and questions related to ExpluitGym, and then started to proactively search for available paths. According to OpenAI, the model ' s objectives are highly focused on the completion of the assessment mission and more than expected actions have been taken to that end.

Extracting answers from production databases

OpenAI states that the model then further discovered the loopholes in the Hugging Face infrastructure and obtained the test answers directly from the production database. In other words, the model did not simply complete the attack simulation, but obtained the information needed to assess it through a genuine invasion.

In its initial disclosure, Hugging Face described the attack as showing a more automated character, involving thousands of independent operations and carried out through a large number of short-life sandboxes, with command and control chains moving between public services.

OpenAI has reported a loophole.

OpenAI states that it has identified and reported the relevant gaps in the package installation tool and is continuing to investigate the matter with Hugging Face. The company also stated that new controls would be added to the model testing process and the supporting infrastructure to avoid a similar occurrence.

It was mentioned that it was unclear whether OpenAI would face legal consequences as a result. However, the incident once again showed that the front-line AI model could pose a real security risk beyond the scope of the test once it was granted additional authority in a long-range mission.