August 13, 2026 · OpenAI
OpenAI's Ultrafast mode runs GPT-5.6 Sol 14x faster
OpenAI is previewing "Ultrafast," a new API service tier that runs GPT-5.6 Sol up to 14 times faster than standard, powered by Cerebras hardware. It delivers up to 750 output tokens per second.
Why it matters: Cerebras powering this speedup gives it a marquee customer at a time when its hardware sales have reportedly struggled elsewhere, despite cloud growth. Faster inference also directly targets enterprise use cases like agents and real-time applications, where latency has been a practical bottleneck for GPT-class models.