Community
Developers can now access significantly faster model inference for latency-sensitive applications, while organizations evaluating or deploying OpenAI models should factor improved performance and cost-per-token metrics into procurement and infrastructure planning decisions.
OpenAI has released GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, offering up to 8x faster token generation compared to Astra Standard mode. The model is available now via OpenAI API and to eligible ChatGPT Work and Codex users, leveraging NVIDIA's hardware-specific inference optimizations.
Read the full article at NVIDIA
CoFabrix summarises and comments on this story. The original reporting belongs to NVIDIA.
Find out where your organization stands -- and what to do about it.