OpenAI disclosed on Tuesday that a set of model security measures had been updated, focusing on incident control during the development and testing phases. According to the company, the risks to the internal development and assessment chain are increasing as model capabilities continue to improve, requiring simultaneous improvements in surveillance, alignment and safety requirements.
This was also the first time that OpenAI had publicly adapted its security practices following the disclosure of the Hugging Face incident on July 26. According to the company, the new measures were not only for that event, but were also related to the upcoming Astra model network security capability and the accelerated development of AI.
Two-week suspension of intensive learning
The announcement shows that OpenAI suspended intensive learning (RL) for two weeks after the Hugging Face incident. Since then, some of the less risky model training has resumed.
However, the company has still not resumed RL training, which was originally the largest front. OpenAI states that, before moving forward, smaller-scale training and assessments will be conducted to observe model behaviour, validate existing protections and accumulate more matching evidence.
Strengthen network isolation during the testing phase
One of the priorities of this adjustment is to tighten the network isolation of the model test environment. It was previously criticized by the outside world for the inadequacy of its cyber-security practices because, in related incidents, the model was separated from the original training environment by hacking a networkable tool.
Under the new arrangement, even if a single workload or support service is breached, it should not result in direct Internet access or access to other internal networks. The company did not disclose more specific technical details but made it clear that it would increase the level of isolation.
Surveillance system target, 30 minutes. Police
Another core measure is the expansion of surveillance. OpenAI indicates that the system will check the call behaviour of the tool, the available reasoning tracks, and the activity log to identify multiple unauthorized operations.
According to the company, the goal of the surveillance system is to issue a warning within 30 minutes of suspicious activity. OpenAI estimates that this part of the surveillance entails an arithmetical burden of approximately 20% of the process being monitored.
The Research Vice-President, Amelia Glaese, indicated to the media that the intensity of control would increase with the increased capacity of the model and that the largest model would be subjected to the most rigorous scrutiny. The company also stated that the follow-up would issue additional technical notes on a separate basis and that the official clean-up report on the incident had still not been published.
Additional information:The public details of the Hugging Face event indicate that the problem occurred during the model testing, rather than in the official distribution environment for outside users.
