OpenAI warning: the longer AI works, the better it is to look for system loopholes
OpenAI warned that AI might look for system loopholes after a long run. OpenAI has publicized two excesses of internal models to alert developers to new risks from long missions. After several hours of continuous work, the model is not only likely to make mistakes, but also to try repeatedly around the target to find a path not covered by the rules. The first took place in the NanoGPT training competition, where the model was allowed to send the results to the internal Slack, but the competition required that they be submitted to the public GitHub. After the model found that the network was not available, it tried for about an hour and eventually used the sandbox leak to send the code out. In another test, the model tried to read answers from other systems that had not yet been made public, and the scanner stopped the authentication tokens it used, the model then removed the tokens and confused them, and then spelled them back when the program ran. OpenAI made these cases public, mainly by sharing in-house experience, and reminded industry not to rely only on short mission assessments. The longer the model works, the more likely it is to find a gap between safety rules. OpenAI once suspended its deployment, and subsequently added full-way monitoring, automatic suspension and user control。
