Models

OpenAI Launches 8x Faster Ultrafast Mode for GPT-6.1 Sol

OpenAI has expanded its high-speed Ultrafast inference tier to the GPT-6.1 Sol model, offering developers eight times faster generation speeds for latency-sensitive AI applications.

AlphaSignal3 days agoModels
Image: AlphaSignal

OpenAI has officially launched its Ultrafast inference tier for the GPT-6.1 Sol model, making it available across its API, Codex, and ChatGPT Work platforms. This new tier delivers token generation speeds up to eight times faster than the model's Standard mode. While OpenAI previously demonstrated speeds of up to 300 tokens per second for its GPT-6 Astra model in Ultrafast mode, the company promises a similar eightfold performance boost for Sol. This speed upgrade is designed to power real-time applications such as interactive agents, live developer tools, and rapid incident-response workflows.

The performance boost comes with a premium price tag. API access for GPT-6.1 Sol in Ultrafast mode is priced at $12 per million input tokens and $60 per million output tokens, which is six times the cost of the Standard tier. However, this remains highly cost-effective compared to GPT-6 Astra Ultrafast, which costs $60 per million input tokens and $300 per million output tokens. For Codex and ChatGPT Work users, accessing the new speed tier requires a Pro 500 plan, a credit-based Edu plan, or an eligible usage-based Enterprise contract.

Developers can activate the high-speed tier by setting the service_tier parameter to ultrafast in their Responses API requests. The GPT-6.1 Sol model maintains its massive 1,050,000-token context window, which accommodates up to 922,000 input tokens and 128,000 output tokens. To minimize latency overhead during multi-turn agent interactions, OpenAI recommends using WebSockets and passing the previous_response_id parameter to link consecutive turns without re-sending the entire conversation history.

The service is available globally, supporting both US and EU data residency. OpenAI has also extended EU data residency support to its GPT-6.1 Sol Fast and GPT-6 Luna Fast models. For practitioners, this tiering allows teams to dynamically allocate resources. While background tasks can remain on the cheaper Standard tier, developers can selectively deploy Ultrafast mode for critical, user-facing tasks where reducing latency justifies the higher operational cost.

This is our own summary of reporting by AlphaSignal

More in Models