August 4, 2026 · MarkTechPost
Cursor open-sources MoE training kernel, 2.37x faster on GB300 racks
Cursor Research open-sourced Mixture-of-Kittens (MoK), the mixture-of-experts (MoE) training megakernel behind its Composer models. It fuses all MoE communication and computation into a single deterministic kernel and runs up to 2.37x faster than the strongest public baseline, but requires Nvidia Blackwell SM100/SM103 GPUs.
Why it matters: Concrete, reproducible training speedups are rare in open-source releases, and the hardware requirement — GB300 NVL72-class racks — illustrates how frontier training optimizations are increasingly designed for, and gated by, the newest Nvidia hardware generation.