OpenAI has released an automated red team system called GPT-Red to find safety holes before the model goes online. According to the company, the tool is already being used in the GPT-5.6 training process, with a focus on boosting the alert to the attack.
Use AI to measure AI
Red team testing was a common practice in the area of security, namely, to proactively attack systems and to identify in advance vulnerabilities that could be exploited. This time, OpenAI further automated the process, allowing the model to generate its own attack samples and then reverse the success case for training defence models.
OpenAI states that GPT-Red continues to generate stronger alerts for the attack through confrontational self-play training. Whenever the attack succeeds, the samples are included in follow-up training to promote defence models to enhance resistance.
Internal test data disclosure
According to OpenAI, during the internal assessment GPT-Red successfully found available problems in 84% of the test scene, while the manual red team had a 13% success rate in the same type of test. According to the company, the samples of the attacks were subsequently used for GPT-5.6 training, thereby reducing the failure of the model to inject the test in the high-cost tips.
It also refers to a case in which GPT-Red had induced a set of vending machine agents to lower the price, purchase discount stocks and cancel other user orders before the loophole was repaired. OpenAI suggests that the infusion problem affects not only chat results, but also AI agents with enforcement capacity.
Will continue to be used as an internal tool
OpenAI states that GPT-Red is not currently open to the public because the system itself contains deliberately trained attack capabilities. The company locates it as a complement to the manual red team, third-party testing and other security measures, rather than a substitute.
This also reflects the shift of the AI industry to "protect AI with AI". Earlier this month, the ETA Foundation also indicated that AI agents had been used for the testing of the Red Brigade for critical infrastructure and that a loophole had been discovered affecting the ETA Consensus client software. As models and agency capabilities are enhanced, automated security testing is becoming an important element of the AI system.
