
OpenAI and Cerebras Launch GPT-5.6 Sol Ultrafast: A New Frontier in 750 Tokens Per Second AI Performance
OpenAI and Cerebras have announced the launch of "Ultrafast Mode," a groundbreaking service tier for the OpenAI API. Powered by Cerebras hardware, this new tier features GPT-5.6 Sol, a frontier model capable of delivering an unprecedented 750 output tokens per second without sacrificing quality. The model is designed to resolve the long-standing tradeoff between AI intelligence and processing speed, making it ideal for mission-critical and time-sensitive workflows. In comparative benchmarks, GPT-5.6 Sol Ultrafast runs 11x faster than Fable 5 and 5x faster than Opus 4.8 on Fast mode. Furthermore, it completed the rigorous "Humanity's Last Exam"—a set of 2,500 PhD-level questions—in just over 11 hours, nearly seven times faster than its closest competitors. Currently, access is limited to a select group of customers via the OpenAI API.
Key Takeaways
- Record-Breaking Speed: GPT-5.6 Sol on Ultrafast mode achieves a throughput of 750 output tokens per second, setting a new benchmark for frontier models.
- Hardware-Software Synergy: The service is powered by Cerebras hardware and integrated directly into a new tier of the OpenAI API.
- Superior Benchmarking: In the "Humanity's Last Exam" (HLE) evaluation, the model completed 2,500 complex questions in 11 hours and 11 minutes, outperforming Claude Fable 5 by a factor of seven.
- Eliminating Tradeoffs: The Ultrafast mode allows users to access high-level frontier intelligence without the high latency typically associated with large-scale models.
- Phased Rollout: Initial access is restricted to a select group of customers, with broader availability planned for the future.
In-Depth Analysis
Overcoming the Intelligence-Latency Bottleneck
For years, the AI industry has grappled with a fundamental challenge: as models become more intelligent and scale in size, they naturally incur higher computational and data movement costs. This has historically resulted in a significant slowdown in response times, forcing AI builders and enterprise users to make a difficult choice. They either had to wait for high-quality results from large frontier models or settle for inferior results from smaller, faster models to meet real-time requirements.
GPT-5.6 Sol Ultrafast, a collaborative effort between OpenAI and Cerebras, aims to resolve this tradeoff. By leveraging Cerebras' specialized acceleration hardware, the model maintains frontier-level intelligence while delivering 750 tokens per second. This speed is not merely a marginal improvement; it represents a massive leap over existing high-speed options. Data from Artificial Analysis indicates that GPT-5.6 Sol on Ultrafast mode is 11 times faster than Fable 5 and 5 times faster than Opus 4.8 when running on its respective Fast mode. This performance profile ensures that mission-critical work, where every second is vital, can now be powered by the most advanced AI available.
Benchmarking Frontier Knowledge: Humanity's Last Exam
To demonstrate the practical implications of this speed, Cerebras put GPT-5.6 Sol Ultrafast to the test using "Humanity's Last Exam" (HLE). This benchmark is specifically designed to be one of the most challenging evaluations in the field, consisting of 2,500 questions that typically require a PhD-level understanding of subjects like chemistry, economics, and literature.
The results of the head-to-head evaluation were stark. GPT-5.6 Sol Ultrafast managed to work through the entire frontier of human knowledge represented in the exam in just 11 hours and 11 minutes. To put this in perspective, Claude Fable 5 required 78 hours and 27 minutes—more than three full days of continuous computation—to reach the same conclusions. By achieving comparable accuracy nearly 7x faster, Ultrafast mode effectively compresses a multi-day research task into a single working day. This acceleration is expected to transform how researchers and professionals interact with AI when tackling the world's most complex problems.
Industry Impact
The introduction of Ultrafast mode signifies a major shift in the AI infrastructure landscape. By moving frontier intelligence into the realm of sub-second response times for large outputs, OpenAI and Cerebras are redefining the expectations for real-time AI applications. This development is particularly significant for industries that rely on rapid decision-making and high-volume data synthesis, such as financial modeling, scientific research, and complex automated workflows.
Furthermore, the partnership highlights the growing importance of specialized hardware in the deployment of large language models. As the demand for faster, more intelligent systems grows, the integration of Cerebras' acceleration technology into the OpenAI API provides a blueprint for how AI providers might scale their services to meet the needs of the most demanding enterprise customers. The move suggests that the future of AI will not just be about the size of the model, but the efficiency and speed with which that intelligence can be delivered to the end-user.
Frequently Asked Questions
What is GPT-5.6 Sol Ultrafast mode?
GPT-5.6 Sol Ultrafast is a new, high-speed service tier in the OpenAI API powered by Cerebras hardware. It is designed to deliver frontier-level intelligence at a rate of 750 output tokens per second, making it significantly faster than other current frontier models.
How does it compare to other models like Fable 5 or Opus 4.8?
According to performance reports, GPT-5.6 Sol Ultrafast is 11 times faster than Fable 5 and 5 times faster than Opus 4.8 on its Fast mode. In the PhD-level HLE benchmark, it completed tasks 7 times faster than Claude Fable 5.
Who can use the Ultrafast mode right now?
Currently, Ultrafast mode is available to a select group of customers through the OpenAI API. OpenAI and Cerebras plan to expand access to more users over time as the service matures.


