parallelquant
July 24, 2026 · The Decoder

Kimi K3 lags US models on cyber exploit tests, fueling distillation claims

The UK AI Security Institute and the US Center for AI Standards and Innovation tested Moonshot AI's Kimi K3 on offensive cyber tasks. It scored 32% on ExploitBench versus 76% for leading US models, and its safeguards failed to block exploit development or simulated attacks. The gap between Kimi K3's strong general benchmarks and weak cyber results fits allegations that Moonshot AI distilled Anthropic's models.

Why it matters: This is independent evidence feeding the US Treasury's sanctions threat against Moonshot over alleged distillation of Anthropic's models: a model that scores well generally but poorly in a specialized domain like cyber offense is a plausible signature of training on another model's outputs rather than from-scratch capability. It's also a reminder that general benchmark leaderboards can mask real safety-relevant capability gaps.

Related updates