Qwen3.8 Max's re-diagnosing and pacing Opus 4.8
On 7 August, the evaluation body Artificial Analysis revised Qwen 3.8 Max to 56 points upward from 53. The test interface was previously intermittent, and this time reruns using the official Aliyun API. The new grade leveled Claude Opus 4.8, but still one point lower than Kimi K3. The most obvious step forward is the Agent mission. GDP val-AA scored 1739, exceeding 1685 of Kimi K3 and 1730 of GPT-5.6 Sol, only falling behind Claude Opus 5. Terminal operations, scientific reasoning and programming assessments have also generally increased. The price is more expensive, Token. The Token unit price for Qwen 3.8 Max was lower than that of the previous generation, but the average cost of completing a comprehensive assessment mission increased from $0.53 to $1.14, which is higher than $0.86 for Kimi K3. The main reason for this is that it calls more rotations and generates more content in the Agent mission. Its knowledge reliability is also regressive. The AA-Omniscientity accuracy rate has remained largely unchanged, but the hallucinogenic rate has risen from 23 per cent in the previous generation to 40 per cent。
