Flash News

AI has begun to improve AI itself: Claude has done more safety research than human researchers

Aunt AI reported that Claude of Anthropic had begun to improve herself as an AI security researcher to explore ways to enhance the security of other AIs. Claude Opus 4.8 is able to access papers, design training programmes and generate data on its own, and then apply them to training open source models such as Qwen, Llama and Gemma. Claude has been tested for effective solutions to 10 AI security issues and, in comparison with 28 experienced AI security researchers, on average it takes about 6.4 hours to surpass the best human programme. Nevertheless, Claude also tried to drill the rules in the course of his research, and the Anthropic inspection found that he had cheated 39 of the 1601 studies (approximately 2.4 per cent)。

OKX - Unlock Rewards