parallelquant
August 15, 2026 · The Decoder

New benchmark shows top AI models still struggle to 'see'

Moonshot AI's PerceptionBench tests multimodal AI models on visual perception, separate from logical reasoning. No frontier model scores above 60% accuracy, with GPT-5.6 Sol leading by a narrow margin, and many apparent reasoning errors actually trace back to misreading the image.

Why it matters: This suggests a meaningful share of AI 'reasoning' failures on visual tasks are really perception failures, pointing labs toward a different bottleneck than commonly assumed. It's a useful check against continued frontier-model benchmark claims covered elsewhere.

Related updates