OpenAI Ultrafast Pushes GPT-5.6 Sol to 14× Speed as Real-Time AI Race Accelerates

· · Views: 2,016 · 3 min time to read

OpenAI has introduced Ultrafast, a new service tier capable of running GPT-5.6 Sol as much as 14 times faster than standard processing, targeting businesses that need frontier-level artificial intelligence without the delays normally associated with more capable models.

750 Tokens Per Second Changes the Speed Equation

Frontier AI models typically require more computation because they perform more sophisticated reasoning and processing, creating a trade-off between capability and response speed.

OpenAI said that getting real-time performance previously often meant choosing a “smaller or more specialized model,” while Ultrafast is intended to deliver “more useful work per second.”

TechCrunch reported that the maximum 750 output tokens per second represents the pieces of generated text produced by the language model as it responds to a user.

India Today said Ultrafast is aimed at users who want GPT-5.6 Sol to respond faster without having to move to a less capable AI model.

Cerebras Powers OpenAI’s Speed Boost

The dramatic performance increase comes from OpenAI’s expanding hardware partnership with AI chip company Cerebras.

Cerebras provides the “ultra-low-latency inference” underlying Ultrafast and is now supporting GPT-5.6 Sol at speeds reaching 750 output tokens per second.

The preview is powered through OpenAI’s Cerebras partnership and is initially being made available only to a small group of customers.

Cerebras hardware enables Ultrafast’s low-latency AI inference and noted that the service is currently available through the API rather than as a broadly released ChatGPT mode.

OpenAI Targets Outages, Finance and Customer Support

Speed becomes particularly valuable when information is changing faster than a slower AI system can analyze it.

OpenAI identified incident response, financial research and security, customer support, commerce, and live research among early Ultrafast use cases. Its developers have used the mode to read logs, analyze traces, synthesize conversations and help prepare or validate fixes while outages are still unfolding.

TechCrunch highlighted incident response, customer service, financial-market analysis and e-commerce as corporate workflows where OpenAI expects faster inference to be particularly useful.

Ultrafast can analyze application logs and recent code changes during an outage, examine suspicious financial activity while market conditions change and resolve complicated customer issues in real time.

Jane Street and Other Companies Test Ultrafast

OpenAI is using the preview to determine whether raw speed changes how companies design AI-powered products.

OpenAI named Jane Street, Podium, Basis and Rogo among early customers testing GPT-5.6 Sol with Ultrafast across coding, commerce, financial research, support and other interactive applications.

India Today also identified those four companies and reported that businesses seeking access can join the Ultrafast waitlist by providing information about workloads, latency requirements and expected usage.

Pricing remains an unanswered question. OpenAI has not yet disclosed token costs for Ultrafast.

Ultrafast therefore represents more than a quicker response animation. OpenAI is testing whether frontier AI can become fast enough to participate in work while events are still happening—turning model latency from a technical limitation into a competitive battleground.

Share
f 𝕏 in
Copied