Flash News

Agent's best researcher still has only one full search for every seven objects

A study published by OneMillion AI shows that only one out of every seven objects can be fully surveyed for Agent ' s task. The study is based on the Wandr benchmark published by perplexity and is dedicated to testing the performance of Agent in large-scale information collection. The mandate requires that a full set of objects be identified and that each object be provided with a reliable source of information. A total of 500 missions involved over 170,000 records. In the six reference systems, the top ranking was given by the universal search as code, but only 13.3 per cent of the subjects met both the information and the evidence. In conversion, only about one in seven objects could be delivered in its entirety. The results show that Agent has been able to find a number of web pages in the existing study, but there is still a risk of missing objects, missing information or missing evidence when processing long lists and multi-field assignments。

OKX - Unlock Rewards