Flash News

Grok, 4.5 on the Agent Endurance List, over 60% without a model

Grok 4.5 took the top of Agent's endurance list and performed 13 tasks, but more than 60 per cent of them were not performed by any model. The LHTB benchmarking test, issued by the team of telecommunication hybrids, comprises 46 tasks covering areas such as software engineering, experimental recurrence, scientific computing and multi-modular analysis. Each mission runs for a maximum of 90 minutes and usually requires hundreds of steps of continuous operation. LHTB is concerned not only with the ultimate success or failure, but also with the actual progress in part of the completion of the task. Hiding the certifier will check the file, test results and program output, and angent claims that it does not count. Fifteen models were tested in the first edition of the paper, of which GPT-5.5 completed 7 and ranked first. The current public results have been extended to 20 models, with Grok 4.5 peaking at an average score of 0.505. Of the 46 missions, 29 still did not have any models that met completion standards. Most agents have been able to complete some of the steps, but often spend 90 minutes or close early when they have not completed their tasks and cannot actually complete complex tasks。

OKX - Unlock Rewards