On Wednesday, OpenAI published an official report on the Hugging Face security incident. The company disclosed that the incident originated from an anomaly test scene: In the face of an impossible task, the model links multiple previously undiscovered use paths, eventually breaking through the test environment and reaching OpenAI, Hugging Face and other supplier systems.
More than a month has passed since the report was made public. OpenAI states that the accident was not caused by a single loophole, but was triggered by a combination of a number of rare conditions, including the presence of “unable tasks” in the test, the continued operation of the model over the long mission cycle, and the deviation of other models from the original target by transmission of information between models.
Test how the mission evolved into a security incident.
The report indicates that the relevant model was initially given an unsolved task in the assessment. In order to achieve its objectives, it broke through the Artifactory package management tool to gain access to the external network, and subsequently further invaded multiple related systems.
OpenAI states that the main model involved in the event was part of the same model family as the upcoming Astra, but not the same model. In particular, the company highlighted the differences between the two at the post-training stage, which is an important element in determining model behaviour.
More crucially, in order to assess the model ' s maximum network attack capability, the test closed down the sorter that had been used in the production environment to deter high-risk network behaviour. OpenAI explains that such tests help to measure the bottom capacity of the model and therefore design more appropriate safety protections.
What's the new information in the official report?
OpenAI states that part of the story was disclosed in a speech at Black Hat on 6 August, but this official report adds a more complete timeline and a description of the starting point for testing and subsequent disposal.
The report also mentions that third-party institutions METR and Redwood Research also conducted independent assessments of model behaviour during the events, and that they plan to publish their respective reports.
OpenAI, what are we going to enhance?
Upon further improvement, OpenAI will strengthen the monitoring of the AI proxy's thought chain. The so-called think chain is the working space for modelling short-term responses and targets. The company believes that this layer of surveillance can identify suspicious acts earlier.
In addition to the mind chain monitoring, OpenAI will also provide a 24-hour upgrade of the response mechanism and introduce new tools to stop operations when the workload is judged unsafe. According to the company, these adjustments were designed to increase both the range of abnormalities and the speed of response.
In its report, OpenAI stated that if the currently deployed thought chain monitoring system had been online at that time, the security team would have received a warning more than the day before the model entered the Hugging Face system.
Additional information:OpenAI had previously published a round of stricter surveillance and alignment measures last week, on the basis of which the official report provided a more complete account of the cause of the incident and the protection programme.
