parallelquant
July 19, 2026 · The Decoder

Kimi K3 tops frontend coding benchmark but trails badly in math

Moonshot AI's open Kimi K3 became the first Chinese model to lead the Code Arena Frontend rankings, beating Claude Fable 5 and GPT-5.6 Sol by a wide margin. On the FrontierMath Tier 4 benchmark, however, it scores only about 39%, versus roughly 90% for top OpenAI and Anthropic models.

Why it matters: The split result complicates the simple narrative that open Chinese models are closing the gap uniformly with Western labs. Kimi K3, already noted here as a large 2.8-trillion-parameter open release, is genuinely competitive on applied coding tasks but still far behind on the hardest formal reasoning, suggesting the remaining capability gaps are becoming task-specific rather than across-the-board.

Related updates