Flash News

OpenAI training on home-based attack models with a success rate of 84 per cent far above humans

OpenAI published an automated red team model gpt-red. The model was injected into the attack through self-playing learning design tips and used the identified loopholes to train gpt-5.6. gpt-red is for internal use only and is not open to public use. The success rate of gpt-red attacks gpt-51.1 was 84 per cent in a new set of untrained scenes, compared to 13 per cent for the Human Reds. In addition, gpt-red succeeded in breaking up vending machines, leading to lower prices and cancellation of orders. When testing codex cli, it can induce agent to transfer sensitive data. After training in confrontation, the success rate of the attack dropped to 0.05 per cent when the gpt-5.6 SOL was injected to the gpt-red direct。

OKX - Unlock Rewards