Nvidia announces full-scale production of Groq 3 LPX, Vera Rubin’s inference efficiency increases severalfold, SpaceXAI increases investment in agents and space AI
Nvidia stated that, based on SemiAnalysis's AgentX workload and utilizing the DeepSeek V4 Pro model, the throughput per megawatt of Vera Rubin is up to 30 times that of the previous generation GB300 NVL72, and the cost per token is reduced by up to 35 times. SpaceXAI will use Vera CPU to accelerate next-generation agent AI applications and expand the platform construction based on Vera Rubin. SpaceXAI plans to extend the optimized Vera Rubin NVL72 into space for use in the first generation of Starmind AI satellites.
Nvidia is accelerating the competition in AI infrastructure from traditional model training and inference performance to a new stage focused on Token generation speed, energy consumption, and cost surrounding Agentic AI.
On Monday, the 24th (Eastern Time), Nvidia announced that Groq 3 LPX, designed for interactive AI inference, has officially entered full-scale production. As an important component of the Vera Rubin platform, Groq 3 LPX is optimized for ultra-low latency Token generation. Nvidia stated that in Artificial Analysis tests, equipped with the Gemma 4 31B model and 100,000 Token context, Groq 3 LPX achieved 3,400 output Tokens per second, setting a new record for this model. For low-latency workloads such as agentic programming, its response speed is up to four times that of the latest competing platforms.
At the same time, Nvidia released new test results for Vera Rubin NVL72 under real agent workloads: Based on SemiAnalysis's AgentX workload and using the DeepSeek V4 Pro model, Vera Rubin's throughput per megawatt (MW) is up to 30 times higher than the previous generation GB300 NVL72, while the cost per Token is reduced by up to 35 times. Nvidia emphasized that agentic tasks differ from traditional chats or summarizations, as task context accumulates through hundreds of steps, even reaching hundreds of thousands of Tokens.

On the same day, Nvidia also announced that SpaceXAI will adopt the Vera CPU to accelerate next-generation agentic AI applications and further expand the construction of platforms based on Vera Rubin, as its computing power scales up to several GW. In addition, SpaceXAI plans to extend an optimized version of Vera Rubin NVL72 into space for use in the first generation Starmind AI satellites.
Groq 3 LPX Fully Launched: Nvidia Completes the “Last Mile” for Agents
Nvidia’s main intention in bringing Groq 3 LPX into full production is not to replace GPUs with dedicated inference chips, but to supplement Vera Rubin with the ability to perform low-latency Token generation that traditional GPU architectures lack.
For agents, a task is often no longer just a "question-and-answer": agents need to read files, call tools, execute code, and check results, then continue inference based on results, potentially looping through hundreds or even thousands of steps. During this process, the speed at which each step generates Tokens impacts the total completion time.
Nvidia says Agentic AI brings two distinct computational challenges: on one hand, efficiently handling growing large-scale contexts, and on the other, continuously generating Tokens with extremely low latency. Vera Rubin NVL72 handles broader training, inference, and context processing, while Groq 3 LPX is optimized for single-user interaction speed, enhancing Token generation rate.
Nvidia's latest tests show that Groq 3 LPX reaches 3,400 output Tokens/second under Gemma 4 31B and 100K context. For agentic tasks and other latency-sensitive workloads, responsiveness is up to four times that of recent competitor platforms. According to Nvidia, this speed enables agents to rapidly complete sequential tasks such as coding, testing, tool invocation, and result validation.
Commercially, AI cloud provider Nebius will be the first cloud vendor to adopt Groq 3 LPX, planning to deploy it in the Nebius Token Factory inference platform. Groq also intends to be among the first adopters for the platform.
This means that as Groq 3 LPX enters full production, the capabilities acquired by Nvidia’s previous Groq technology purchase are now truly being integrated into Vera Rubin’s data center product system.
30x is Not “Speed” 30x, But Many More Agent Tasks Per Megawatt
Compared to standalone performance metrics like 3,400 Tokens/second, what’s more noteworthy from Nvidia this time is the focus on throughput per megawatt and Token cost.
Nvidia’s published test results show that under the SemiAnalysis AgentX workload, Vera Rubin NVL72 achieves up to 30x the throughput per MW of the GB300 NVL72, with Token costs down by as much as 35x. The test uses real agent programming trajectories, maintaining realistic context growth, tool invocation, and sub-agent generation steps.

Therefore, the 30x improvement in throughput per MW reflects a huge leap in AI inference efficiency, not simply a 30x increase in chip inference speed.
This metric is particularly significant for AI data centers.
Traditional large-model services focus more on single-request latency and throughput, while agent workloads continuously consume Tokens. Citing OpenRouter data, Nvidia notes that agent AI workloads can consume up to 15 times the Tokens of simple chat requests. This is because an agent might query databases, search news and files, invoke sub-agents to analyze, run code, and repeatedly validate results, all while context continuously accumulates.
This gives rise to a new way of measuring AI infrastructure’s economics: how many effective agent tasks can be completed with the same megawatt of power?
As power and data center resources increasingly become bottlenecks for AI expansion, increasing the number of Tokens produced per unit of electricity may be more important than simply maximizing peak compute capacity.
Vera Rubin Transforms From “GPU Platform” to a Complete AI Factory
From an architectural standpoint, Groq 3 LPX is not an isolated product, but part of Nvidia Vera Rubin’s “ultimate synergy design” strategy.

Nvidia had previously positioned Vera Rubin as an AI factory platform for the Agentic AI era. The system is not simply a cluster of GPUs but is synergistically designed around compute, networking, storage, and software, covering Rubin GPU, Vera CPU, Groq 3 LPX, BlueField-4 DPU, Spectrum-6 SPX, and other components.
Among them, the Rubin GPU handles large-scale training, inference, and context processing tasks; Groq 3 LPX is responsible for high-speed Token generation; Vera CPU handles the vast number of CPU-intensive agent tasks that occur between model invocations, including tool orchestration, code execution, data processing, and simulation.
Nvidia’s previous architecture disclosures show Vera Rubin employs seven chips and five dedicated racks with collaborative design, the intent being that different hardware performs tasks for which each is best suited rather than having one processor type tackle the entire AI workload.
This means Nvidia is upgrading AI infrastructure from a mere “GPU cluster” into a heterogeneous computing system.
For AI cloud providers, this architecture allows resource allocation for training, long-context inference, Token generation, tool invocation, and other stages, boosting overall data center utilization and returns.
SpaceXAI Doubles Down on Vera: Agent AI Moves from Data Centers to Space
As Groq 3 LPX enters large-scale production, SpaceXAI also announced further adoption of Nvidia's Vera Rubin ecosystem.
On August 24, Nvidia announced that SpaceXAI will deploy the Vera CPU for next-generation agentic AI applications. Vera is designed specifically for CPU workloads beyond model inference, including tool invocation, code execution, data processing, task orchestration, and simulation.
Vera utilizes 88 Nvidia-designed Olympus cores and is equipped with high-bandwidth LPDDR5X memory, with bandwidth as high as 1.2TB/s. Nvidia claims that for agentic AI, reinforcement learning, and data processing workloads, Vera’s task completion speed can be up to 1.8 times that of x86 CPUs.
For SpaceXAI, Vera’s value lies not just in enhanced CPU performance.
During agent execution, many tasks occur between GPU inference rounds, such as code execution, data processing, tool invocations, and coordination among different agents. If CPU processing falls short, GPUs are left waiting, reducing the whole AI factory’s utilization.
Therefore, SpaceXAI aims to use Vera for these “non-model” tasks, enabling GPUs to focus more on model training and inference, thereby improving overall system efficiency.
SpaceXAI also plans to expand Grok AI infrastructure based on Vera Rubin as computational capacity grows to multi-GW levels. Further, the company intends to use the optimized Vera Rubin NVL72 in the first generation of Starmind AI satellites.
This signifies Nvidia’s accelerated computing ecosystem is extending from terrestrial data centers to orbital computation.
From “Chatbots” to “Autonomous Execution”, AI Compute Demand is Shifting Gears
From the full-scale production of Groq 3 LPX, to Vera Rubin’s 30x per-megawatt throughput, and SpaceXAI expanding Vera Rubin to Starmind satellites, Nvidia’s latest streak of announcements all point to one paradigm shift: AI applications are moving from generating content to autonomously completing tasks.
In chat and summarization scenarios, users usually need the model to quickly generate a short text; but in agent scenarios, the model must continuously reason, invoke tools, execute code, access external data, and repeatedly verify results.
This means the core metrics for future AI infrastructure may also evolve — moving from “training a bigger model” to focusing on how many effective Tokens are produced per unit of power, how many tasks are completed per unit cost, and ultimately how long it takes for an agent to finish a complex task.
Nvidia’s integration of Groq 3 LPX, Vera CPU, and Rubin GPU into one AI factory system is a redesign in response to this change.
Meanwhile, SpaceXAI offers a more imaginative application direction: If this ecosystem can eventually extend from ground-based data centers to orbit, the boundaries of AI computing will no longer be confined to traditional data center facilities.
Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.
You may also like
XRP Japan Expands Through SBI-Ripple Rails

XRP jumps from $1 to $1.69 in sharp rally as Trump urges Congress on CLARITY Act
Solana DApps generate $35M in revenue, a 29-week high
Daydreams unveils TaskMarket to standardize outsourcing in the agent economy
