Flash News

Terminal-Bench 4.0: GLM-5.3 to third, exceeding GPT-5.6 Sol

On 29 August, Terminal-Bench issued 4.0 for the purpose of recalibrating Agent ' s time, CPU and memory on mission and repairing 19 missions and removing 8 missions with saturation, denial, open resolution or quality problems. The maximum duration of all assignments was harmonized to eight hours, mainly to reduce the disruption of achievement by time overruns and environmental problems. In the last list, Opus 5+ClaudeCode ranked 51.8% first and Fable 5 44.5% second. GLM-5.3+Claude received 41.8 per cent, reaching the third, exceeding 37.3 per cent of GPT-5.6Sol+Codex. GLM-5.3 is the only model that is not Anthropic. In Terminal-Bench3.0 the GLM-5.3 is still the fourth largest in 32,4%, behind 34.6% of GPT-5.6Sol; by 4.0, the GLM-5.3 is third, which in turn leads Sol 4.5 percentage points。

OKX - Unlock Rewards