The three-wheel callback rate to Opus 5 is only half the cost
On 5 August, Aikido, a security company, tested seven models using code auditing Agent, including Qwen 3.8-Max, Claude Opus 5, Kimi K3 Max, DeepSeek V4 Flash, and GPT-5.6 Sol, Luna and Terra. The tests contained 32 recent open loopholes and each model operated three times. The Qwen3.8-Max three-round totals 26 holes, with a recall rate of 81.3 per cent, first with Opus 5. Its combined F1 score (comprising underreporting and misstatement) was 83.2 per cent, slightly lower than Opus 5, Kimi K3 Max and GPT-5.6 Sol. The question is one of one-time instability. Of the 26 holes, only 10 can be found in three consecutive rounds, 19 in Opus 5 and Sol. However, the total cost of thousands of questions is about $821, only half of Opus 5 and Sol, but still five times more than DeepSeek V4 Flash。
