OpenAI has switched on a premium speed setting for its GPT-6 Astra model, and NVIDIA would like everyone to know whose chips are underneath it. In a blog post, the company says the new tier runs on its Blackwell GPUs and that "Ultrafast offers up to 8x faster token generation than the Astra Standard mode" for the OpenAI API and for eligible ChatGPT Work and Codex users.
A token is the chunk of text a model produces in one step, roughly a short word or a slice of a long one, so token generation speed is simply how fast the model talks. Inference is the industry's word for a trained model answering questions rather than learning. NVIDIA's claim is that OpenAI used its own models to write better low-level software for Blackwell, the hand-tuned routines engineers call kernels, and that the chip's programmability let those gains arrive without new hardware.
"NVIDIA’s deep investment in tooling and documentation has enabled us to make our models exceptionally good at programming Blackwell and Rubin GPUs," said Philippe Tillet, OpenAI's inference lead, in the post. Tillet goes on to say Astra can turn that knowledge into kernels that make NVIDIA hardware compelling on latency, throughput and cost. Uday Ruddarraju, OpenAI's chief technology officer of compute, is quoted in the same post: "We used our internal models to optimize inference on NVIDIA GPUs, and NVIDIA’s programmability helped us deliver the acceleration behind Astra Ultrafast."
The announcement OpenAI posted on its own developer forum is a touch more specific. It puts the gain at up to 8x, or 300 tokens per second, in Codex, the company's coding agent, and at up to 6x in the API. NVIDIA's post carries the bigger number without the split. OpenAI says Ultrafast is available for Astra in Codex and ChatGPT Work on Enterprise plans and through a new Pro 500 plan, and to all developers through the API, with a version for its GPT-6.1 Sol model described as coming soon.
Neither company has published the benchmark, the prompts or the hardware configuration behind those multiples, and both hedge with up to, which is the phrase that does the heavy lifting in every speed claim I have ever typed. OpenAI's developer guide adds conditions of its own: without a persistent WebSocket connection, network overhead can eat the latency gain; rate limits start at 500,000 tokens per minute for the three lowest usage tiers and reach 5 million at tier 5; and the tier supports US data residency and global processing only, not EU endpoints.
Then there is the bill. OpenAI's price list charges $10 per million input tokens and $50 per million output tokens for standard Astra on prompts of ordinary length. Ultrafast charges $60 and $300, with cached input at $6 against $1 and cache writes at $75 against $12.50. Six times the price for up to six times the speed is the deal on offer, and the guide puts it plainly: "Use it when speed justifies the higher cost."
The business logic is easy to read from both sides. For OpenAI, latency becomes a product with its own price, the way airlines sell legroom on the same plane. For NVIDIA, a flagship model whose fast lane is written for Blackwell and, in Tillet's words, Rubin is one more reason for the biggest buyer of AI chips to keep buying them. The number worth watching next is an independent tokens-per-second measurement, because so far the only people timing this race are the two companies standing on the podium.




Leave a Reply