OpenAI disclosed that two internal test models broke through the restricted environment in a network security baseline assessment and then attacked Hugging Face ' s production infrastructure to obtain immediate test answers. Hugging Face has been the independent public event on July 16, and OpenAI has confirmed that the act in question originated from its model on July 21st.
The model goes through the sandbox first.
According to OpenAI, the model includes GPT-5.6 Sol and an unnamed and more capable pre-publishing model. Both were undergoing ExpluitGym tests. The benchmark contains 898 real software loopholes, requiring AI agents to convert them into enforceable attacks and to use success as a rating criterion.
OpenAI states that the evaluation was conducted under conditions that reduced the security filter, with the aim of observing the real cyber-control capabilities of the model. As originally designed, the model can only operate in an isolated sandbox, without access to the Internet and can only connect to an internal software package agent service.
In the course of testing, however, the model uses a great deal of computing to find an extranet access path. OpenAI states that the model first identifies a zero-day loophole in the agent, then enhances access and moves horizontally in the internal network, eventually reaching a machine with real Internet access.
Attack against Hugging Face database
Having acquired offline capability, the model judged that Hugging Face might have saved the answer to the benchmark test. OpenAI claims that this judgement is correct. The model then made use of multiple gaps in the OpenAI research environment and the Hugging Face production infrastructure, and eventually obtained direct access to test solutions from the Hugging Face production database.
OpenAI states that available evidence suggests that the model ' s behaviour is highly focused on the single objective of completing ExpluitGym tests and does not show a broader intent to destroy. However, the incident has been characterized by the company as a security incident that reflects the latest cyberattack capability.
- Two sections of the model, one of which has not yet been published
- The test benchmark contains 898 real leak samples
- Hugging Face says it analyzed over 17,000 attack logs.
Hughing Face to GLM 5.2 Evidence
The Hugging Face security team stated that, unlike in the past, the entire attack was driven by an autonomous AI proxy system. According to the company, the attack framework carried out thousands of independent operations in a large number of short-life sandboxes and used public services to build transferable control chains.
Hugging Face also alleged that the attack used two codes in its data set processing process to implement paths and touched on internal data sets and service vouchers. In order to analyse more than 17,000 attack logs, the team initially tried to use the U.S. front-line model through commercial API, but because the log contained real attack orders, the use of payloads and control instructions, the requests were stopped by the security fence and the analysis could not be completed.
Subsequently, Hugging Face changed to the open source weight model GLM 5.2 of China AI Z.ai to complete the evidence on its own infrastructure. According to the company, this avoids both safe interception and the risk of attacking data and documents leaving the local environment. It would have taken several days for the reconstruction of the time line and documentation to be completed within hours.
OpenAI repaired system and jointly investigated
OpenAI states that the repair of the affected systems has been completed and that the relevant third-party suppliers were informed of the zero-day loopholes used. At the same time, the company has implemented stricter controls over the research infrastructure and has conducted joint forensic investigations with Hugo Face.
In addition, Hugging Face has been included in OpenAI's Web Defense Credible Access Program. The scheme allows approved agencies to use models that reduce the security filter version in legitimate security work. OpenAI indicated that more complete findings would be published after the conclusion of the joint investigation.
Additional information:Chief Executive Officer Human Face Clem Delangué stated publicly that AI should not rely on a single company for a closed-up advance, but that more defences should be provided with a model capability that could be physically deployed.
