THE CRUNCH

OpenAI has released GPT-6 Astra Ultrafast, a faster inference mode for its GPT-6 Astra model, now available through the API and to eligible ChatGPT Work and Codex users. The Ultrafast mode runs on NVIDIA Blackwell GPUs and leverages inference optimisations to deliver up to 8x faster token generation compared to the Astra Standard mode. This acceleration is designed to improve the responsiveness of coding agents and other interactive applications by shortening edit-test-debug cycles and reducing latency between tool calls.

OpenAI’s inference lead, Philippe Tillet, noted that NVIDIA’s investment in tooling and documentation has enabled the team to create high-performance kernels for Blackwell and Rubin GPUs, making the hardware compelling across latency, throughput and cost metrics.

The acceleration is not a one-off gain; OpenAI is using its own models to refine the inference software running on NVIDIA GPUs, allowing for continuous performance improvements over time. This approach also allows developers to reuse infrastructure across training, inference and reinforcement learning as models evolve.

Chief technology officer of compute at OpenAI, Uday Ruddarraju, highlighted that the programmability of the NVIDIA platform was key to delivering the acceleration behind Astra Ultrafast.

WHAT HAPPENS NEXT

Developers can access GPT-6 Astra Ultrafast through the API today, with implementation details available in the Ultrafast guide.