Back Original

Accelerating GPT-5.6 Sol Ultrafast

Today, Cerebras and OpenAI are sharing an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras. Ultrafast is available initially to a select group of customers, with access expanding over time. Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second and without any quality compromise, allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical work.

Frontier Intelligence at Unprecedented Speed

AI builders have always needed to choose between speed and intelligence. As models scale up in size and intelligence, they incur higher computational and data movement costs, slowing down response times. Users often need to wait for high-quality results or accept inferior results within a shorter timeframe.

GPT-5.6 Sol Ultrafast resolves this tradeoff, bringing frontier intelligence to products and workflows where every second matters. Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.

At Cerebras, we put Ultrafast to the test by running it head-to-head with popular models on Humanity's Last Exam. HLE is a challenging model benchmark that consists of 2,500 questions typically answerable only by those holding PhDs in fields such as chemistry, economics, and literature.

In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.

Benchmarking was performed by Cerebras using GPT 5.6 Sol Ultrafast with Codex on xhigh reasoning on July 10 and Claude Fable 5 with Claude Code on xhigh reasoning on July 13-15.

As model capabilities continue to advance, the range of applications for fast inference expands. GPT-5.6 Sol is OpenAI’s best model yet for legal briefs, financial models, and engineering reports. On GDP-Val, a benchmark for economically valuable knowledge work tasks, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation, showing how faster inference can accelerate economically valuable work.

Benchmarking was performed by Cerebras on July 31 2026 using GPT 5.6 Sol and GPT 5.6 Sol Ultrafast on medium reasoning within Codex.

High-Speed Intelligence Powers High-Stakes Work

Faster intelligence changes what’s possible for individuals and organizations. With Ultrafast, you can now put agents on the critical path of problems where every second counts.

"With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate. We’re excited to see how workflows and applications are transformed by Ultrafast inference."

Rohan Varma

Product at OpenAI

Ultrafast is a persistent edge for organizations using frontier AI to quickly respond to incoming information. Companies operating web services can leverage Ultrafast to root-cause and address production outages, preserving customer trust, preventing lost revenue, and saving downtime minutes against their SLAs. And in adversarial, high stakes cyberattacks, Ultrafast is an invaluable tool for security teams who must quickly detect and respond to bad actors to contain catastrophic losses.

More broadly, Ultrafast enables entirely new modes of working with agents, it delivers real-time insights and updates, so you don’t have to context-switch across multiple parallel sessions to get the most out of your agents.

"Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive."

Jeffrey Wang

OpenAI Researcher

With Ultrafast, researchers and engineers can reserve their attention for going deep on select problems that matter most, while continuing to use Standard processing for parallelizing commodity tasks. Cerebras is excited to power the next wave of AI innovation, raising the ceiling for what individuals and organizations can accomplish with responsive AI.

Breakneck Speed is Enabled by Breakthrough Innovation

GPT-5.6 Sol on Ultrafast mode is powered by Cerebras’ revolutionary Wafer-Scale Engine architecture, purpose-built for frontier AI workloads. Fast frontier inference is a data movement problem: on GPUs, inference on large models is bottlenecked by memory bandwidth, as model weights must be repeatedly transferred between on-chip memory and off-chip storage to generate successive tokens within a model response.

Cerebras takes a contrarian approach to eliminating this inefficient data movement: we pack 44 GB of SRAM on each wafer-sized chip. Weights stay on-chip, and tokens flow uninterrupted through model layers pipelined across wafers. This technical approach scales smoothly with model size, paving the way for a continued speed advantage on future frontier models.

Ultrafast: Now in Limited Preview

GPT-5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. Access will expand as capacity grows. Sign up for updates.