parallelquant
August 11, 2026 · The Decoder

Nvidia's open-weight Nemotron 3.5 Lightning trades size for speed

Nvidia released Nemotron 3.5 Lightning, an open-weight model with 3.6 billion active parameters that reportedly matches OpenAI's gpt-oss-120b on the Intelligence Index despite being roughly four times smaller. It runs at nearly 670 tokens per second, making it the fastest model in its comparison set.

Why it matters: The release reflects a broader industry shift toward efficiency-optimized open models rather than chasing ever-larger parameter counts, useful for developers who need low-latency inference on cheaper hardware. It also puts competitive pressure on other open-weight releases, like OpenAI's gpt-oss line, on the speed-versus-intelligence tradeoff rather than raw benchmark scores alone.

Related updates