OpenAI rolled out "Ultrafast" mode for GPT-5.6 Sol, its latest flagship model, delivering output speeds of up to 750 tokens per second. The infrastructure upgrade leverages Cerebras hardware from OpenAI's $10 billion hardware partnership announced last year.

The speed gain represents a 14x acceleration over baseline performance. Ultrafast mode joins two existing tiers—Standard and Fast—creating a three-tier inference pricing structure. This positions inference speed as a standalone product variable rather than bundling it uniformly with model capability.

The deployment marks a shift in how AI providers monetize access. Enterprises face a trade-off matrix now: pay less for slower inference, or pay premium rates for near-instantaneous responses. Real-time applications like customer service, live translation, and trading algorithms benefit most from the fastest tier. Batch processing workloads can stay on Standard.

Cerebras hardware provides the computational backbone. The chip manufacturer specializes in large-scale AI inference, with massive wafer-scale processors designed to handle parallel token generation. This partnership, formalized in late 2024, gave OpenAI dedicated capacity to roll out performance tiers without congesting shared infrastructure.

The move reflects growing market competition in LLM inference. Companies like Together AI, Groq, and Hugging Face have positioned speed as a competitive edge. OpenAI's approach—offering speed as a paid premium rather than free standard—converts performance gains into revenue. This strategy works only if Ultrafast mode remains meaningfully faster than alternatives.

Token-per-second throughput matters for streaming applications. Higher throughput means shorter latencies between user input and first token appearance, then complete response time. At 750 tokens per second, Ultrafast can deliver a 2,000-token response in under three seconds, a significant improvement over Standard's 50-150 token per second baseline.

Cost structure remains unclear from available details. OpenAI typically charges per million input and output tokens. Ultrafast likely commands a multiplier—perhaps 1.5x to 3x Standard pricing—to justify infrastructure costs and manage demand. Early adoption clusters in latency-sensitive sectors: financial services, customer support, and real-time content generation.

The Cerebras partnership resolves a prior constraint. OpenAI's training runs on NVIDIA chips, but inference diversity allows vendor flexibility. Cerebras wafers pack more compute density and lower power consumption than alternatives, making high-throughput inference economical at scale.

Competitors watch closely. Anthropic's Claude and Google's Gemini must decide whether to match speed tiers or emphasize capability depth. Groq, which already sells speed-optimized inference, faces pressure from an OpenAI Ultrafast offering backed by OpenAI's dominance in user trust and model quality.

GPT-5.6 Sol itself represents iterative progress over prior models. The exact capability improvements remain undisclosed, but OpenAI hints at better reasoning and code understanding. Speed gains plus refined reasoning create stacking benefits for professional users.

The three-tier model standardizes across OpenAI's product line. Researchers, hobbyists, and cost-conscious builders use Standard. Power users and small businesses adopt Fast. Enterprises and latency-critical systems deploy Ultrafast. This segmentation mirrors cloud computing's reserved-capacity pricing: pay more upfront, get guaranteed performance.