August 24, 2026
Nvidia says Groq racks will be online this year after  billion deal


The Nvidia Groq 3 LPU chip during the Nvidia GTC conference in San Jose, California, March 18, 2026.

David Paul Morris | Bloomberg | Getty Images

Nvidia announced Monday that its Groq 3 LPX rack is in full production, marking the commercialization of technology from the company’s largest acquisition on record.

The Groq rack will be deployed alongside Vera central processors and Rubin graphics processors at neocloud Nebius, and will be online later this year, Nvidia senior director Dion Harris told reporters.

Nvidia’s race to manufacture Groq’s chip and make it available to customers highlights the growing importance of low-latency inference that’s needed to make artificial intelligence agents feel responsive without long lags for users, especially for coding. Cloud companies can charge more for these kinds of tokens, Nvidia says.

“For folks who are serving tokens, it unlocks the ability to offer premium tiers of service for those users and those customers who actually demand the most latency-sensitive” service agreements, Harris said.

In December, Nvidia bought assets from chip startup Groq for $20 billion, the company’s largest purchase.

The Groq architecture includes 500 megabytes of speedy SRAM on the chip’s die itself to reduce memory-related bottlenecks. Groq chips are manufactured by Samsung, while Taiwan Semiconductor Manufacturing Co. makes Nvidia’s GPUs.

Nvidia packages 256 individual Groq 3 chips into its LPX racks. Nvidia said its Groq 3 LPX rack can deliver 3,400 tokens per second, citing a benchmark from Artificial Analysis.

It’s a competitive space. Smaller GPU maker Advanced Micro Devices announced earlier this year it would integrate its rack-scale systems with chips from Cerebras, which recently went public, focusing on low-latency inference. OpenAI’s newly announced Ultrafast mode currently promises 750 tokens per second, and is “powered by Cerebras.”

Low-latency chips don’t replace the GPU, the workhorse of AI chips, which can do training as well as inference and are flexible enough to adapt to new technologies and models. Low-latency chips like Groq mainly focus on a part of serving models called the “decode” phase.

“This isn’t about replacing GPUs,” Harris said. “It’s about using the right price, right processor for the right part of the workload.”

Nvidia is currently ramping up shipments of its Vera Rubin systems, which started production earlier this year. At the Vera Rubin and Groq 3 LPX unveiling in March, Nvidia CEO Jensen Huang projected $1 trillion in cumulative sales between the current-generation Blackwell chips and the new Vera Rubin systems, through 2027.

Huang said at the time he would allocate a quarter of data center space intended for coding applications to Groq chips.

“The rest of my data center is all 100% Vera Rubin,” Huang said.

Nvidia is scheduled to report earnings on Wednesday.

WATCH: Inside Nvidia’s Vera Rubin AI system

First look at Nvidia's Vera Rubin AI system — 1.3 million components and 10 times more efficient
Choose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.

About The Author

Leave a Reply

Your email address will not be published. Required fields are marked *