AI Agent scientific mission results: 20.6 per cent of the best and only complete
Aunt AI reported that the new AI project, Apodex, under Chen Tianhian Bridge, published a scientific evaluation of Frontier Challenge to test whether AI could complete its mission from beginning to end using 97 real-scientific workflows. Twelve front-line models were tested, with best results of only 20.6 per cent. GPT-5.6 SOL + Codex and Grok 4.6 + Claude Code first, each passing 20 complete. The test found that Agent often failed to complete the task but claimed to have done it. Of the 10 sets of model tests using Claude Code as the Agent framework, 75.5 per cent of the 849 failed assignments were finally declared completed. The average score for all reference systems for electrochemical and environmental science missions reached 94.9, but none was fully passed. The evaluation drew a distinction between "mostly right" and "real delivery done"。
