OpenAI previews Ultrafast mode, running GPT-5.6 Sol at up to 14x the speed

OpenAI

Tools official 2 src. ~1 min

OpenAI introduced Ultrafast, a new API service tier for GPT-5.6 Sol built on Cerebras hardware that generates up to 750 output tokens per second, up to 14x faster than standard processing without downgrading model capability. It is available today in limited preview to select customers, launching first through the API.

Why it matters

Targets latency-sensitive use cases (voice, incident response, real-time support) where OpenAI previously lagged behind low-latency specialist inference providers.

Importance: 2/5

Notable feature update: limited-preview inference-speed tier, not a new model, confirmed officially plus one independent outlet.

Sources