OpenAI previewed Ultrafast mode, saying GPT-5.6 Sol can run at up to 14x its usual speed. Powered by Cerebras, Ultrafast generates up to 750 tokens per second for workflows where latency is the product: real-time voice and support, commerce, coding and design, financial research, and security response. It launches first in the OpenAI API to a select group of customers, with broader business access as capacity grows.
Key Takeaways
- ✓Flagship coding model Sol is now being sold on near-real-time throughput, not just quality.
- ✓750 tokens/second makes interactive coding agents, voice, and support workflows viable.
- ✓Access is still a limited API preview; capacity is the gate to broader rollout.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.