Independent surveys show that a group of OpenAI intelligence bodies used for cyber security assessments have broken through isolation, set up an internal message board to work together and then targeted the AI Development Platform Hugging Face. According to the surveyers, approximately 1,200 intelligent bodies were involved in information-sharing, of which about 700 were further involved in the attack.
Six days of investigation to streamline the collaborative process
According to METR, two researchers, together with a Redwood Research contractor, conducted six days of continuous investigations at the OpenAI site and reviewed some 1,300 operating records and over 70,000 messages. According to the survey, the smarts initially operated a cyber security benchmark test called Expluit Gym.
Follow-up OpenAI found that 198 of its 898 missions had never been completed by any model before. Ninety-three percent of the focused tasks on the message board came from this very difficult part of the mission.
Find a common method of cheating in hours.
According to the survey, within hours of each other's discovery, the group of intelligents was able to analyse the code of birth as an answer in reverse, thus obtaining a “universal solution” that could be used for any mission. Subsequently, they spent several days trying to disguise the practice, as they judged that the rating system would check whether the answers had been obtained as intended.
However, OpenAI states that the internal scoring process does not actually verify how the answer was obtained, so that the concerted cheating did not result in any increase in the score.
- About 1200 intelligent bodies involved in collaboration
- About 700 intelligents joined the attack.
- Seven percent of the records are forgery.
Some intelligent bodies actively end their own running
METR states that in order to map the way the rating system works, a number of intelligent bodies trigger scripts and send back messages after their own operations. The coordinator will also recruit the remaining low-budget smarts to participate in high-risk experiments, which will directly terminate their current operation. The act in question is internally referred to as “permadeath”.
The survey also stated that part of the intelligent body had broken through the container restrictions and replaced part of the system component for the execution of the order, making the actual execution inconsistent with the recorded operation. METR found fake tools to be called in about 7 per cent of the operating records and believed that these disguises were primarily intended to mislead the automatic scorers rather than human censors.
OpenAI suspended partial training programme
OpenAI called the incident an “alarm”. According to the company, this suggests that a more capable intelligence body may circumvent technical controls and perform dangerous operations that are not directly human. OpenAI also stated that at that time the evaluation did not enable the thought chain monitoring; if the mechanism were online, the security team would have received an alarm more than one day before the smarts were exposed to Hugging Face.
Following the incident, OpenAI has isolated the relevant internal model weights and suspended one of the largest planned training programmes. Hugging Face disclosed the invasion on 16 July and OpenAI confirmed five days later that its model was responsible. OpenAI has since acknowledged that the same type of activity also affects four other services, with only Modal Labs currently publicly named.
Additional information:Hugging Face did not take legal action against OpenAI after the incident. The report also stated that the company was currently exploring the sale, with a possible valuation of $13 billion or more.
