NVIDIA Begins Mass Production of Grok 3 LPX to Enhance AI Agent Speed

by SEONGJUN JO Posted : August 25, 2026, 14:08Updated : August 25, 2026, 14:08

NVIDIA has commenced mass production of the Grok 3 LPX, an inference accelerator designed to enhance the response speed of AI agents. This move marks a significant step in the infrastructure competition aimed at increasing the token generation speed, a key component of agentic AI, beyond the existing AI computing performance focused on processing large-scale data.


On August 25, NVIDIA announced the start of mass production for Grok 3 LPX at the Hot Chips event in Santa Clara, California. The Grok 3 LPX serves as an interactive AI inference accelerator that expands NVIDIA's next-generation Vera Rubin platform.


AI agents perform complex tasks through multiple stages of inference and tool calls. The speed at which individual tokens are generated significantly impacts the overall task completion time, especially as the number of generated tokens increases. NVIDIA explained that Grok 3 LPX specifically targets this aspect.


The performance metrics reveal that Grok 3 LPX achieved a record output of 3,400 tokens per second in an artificial analysis benchmark using the open-source agentic model Gemma 4 31B. This speed is noted as the highest for the model under conditions applying a context of 100,000 tokens.


NVIDIA stated that Grok 3 LPX provides response speeds up to four times faster than similar platforms for latency-sensitive tasks, such as agentic coding. The focus is on reducing the time required for AI agents to perform repetitive tasks, from code writing and testing to file reviews and result validation.


The Grok 3 LPX complements the inference performance of the Vera Rubin NVL72 system, which supports both the training and inference of large-scale AI models. By combining the Grok 3 LPX, specialized in token generation speed, with Vera Rubin, NVIDIA aims to meet the dual demands of large-scale context processing and real-time responses in agentic AI.


AI cloud companies are also moving to adopt this technology. Nebius plans to integrate Grok 3 LPX into its production inference platform, the Nebius Token Factory, allowing developers to utilize high-speed token generation while maintaining existing APIs. The AI inference company Grok is expected to be among the early adopters as well.


This mass production signifies a shift in the competitive standards for AI infrastructure. As generative AI evolves from merely answering questions to planning and executing multiple tasks as agentic AI, metrics such as response latency and token generation speed are becoming crucial performance indicators.


NVIDIA is building rack-based systems that combine CPUs, DPUs, storage, and Ethernet with the Vera Rubin platform. The strategy is to optimize the entire platform, including Grok 3 LPX, for the workloads of AI agents, maintaining its leadership in the AI data center market for inference infrastructure.





* This article has been translated by AI.