Thinking Machines Lab releases open-weight Inkling-Small MoE model
Thinking Machines Lab released Inkling-Small, a 276-billion-parameter mixture-of-experts (MoE) model with only 12 billion active parameters per token, as open weights. It reportedly matches the performance of its larger sibling Inkling at a quarter of the size, and its NVFP4 quantized checkpoint runs on a single Nvidia B300 GPU.
Why it matters: Matching a larger model's performance with far fewer active parameters, and fitting it on a single GPU, continues the industry-wide push to make frontier-adjacent capability cheaper to run -- a trend that matters more than raw benchmark scores for who can actually deploy these models. Coming from Mira Murati's well-funded Thinking Machines Lab, it's also a concrete signal of the company's research direction after months mostly known for hiring and funding news.