Flash News

Kimi Open Source PerceptionBench Visual Perception Benchmark, GPT-5.6-Sol tops with no model breakthrough of 60%

Kimi Team declared the Open-source Large Model Visual Sensory Assessment Benchmark Perception Bench, which aims to decompose visual sensory capacity into 10 atomic-grade capabilities for independent assessment, covering dimensions such as visual relationships, counting, properties, depth and 3D, positioning, comparison, fine particle recognition, contextual integration, OCR and hallucinogenic recognition. The benchmark is based on 42 existing centralized model failures, with a total of 3,000 manually validated questions, each of which looks at a single visual capability without the need for reasoning or external knowledge. The results of the assessment showed that none of the 16 front-line large multi-modular models had overall accuracy of over 60 per cent. GPT-5.6-Sol ranked first at 59.7 per cent, with Kimi K3 (58.5 per cent), Claude-Fable-5 (57.2 per cent), Gemini-3.1-Pro (56.2 per cent) and GPT-5.5 (55.8 per cent) ranked fifth. The report notes that visual hallucinations continue to be the weakest performance of models and that there is still significant room for improvement in overall sensory capacity。

OKX - Unlock Rewards