Facing competition from Google’s TPU, NVIDIA CEO Jensen Huang decisively invested $20 billion to partner with hot chip startup Groq, aiming to strengthen AI inference capabilities. Technology investor Gavin Baker highlighted that GPUs face latency bottlenecks in inference: prefill stages suit GPU parallelism, but decode stages are sequential, and GPU performance is limited by HBM memory, leaving FLOPs underutilized.
Groq’s LPU chips leverage on-chip SRAM, running 100x faster than GPUs, handling 300–500 tokens per second at full load. However, each LPU only has 230MB of memory; large models require hundreds of LPUs, taking much more data center space than GPUs and incurring high overall hardware costs.
This strategic move allows NVIDIA not only to upgrade technology but also to defend its moat: integrating Groq’s low-latency inference chips addresses emerging market needs. The rise of TPU revealed GPU limitations in inference, and Groq provides NVIDIA a critical boost to maintain AI leadership.



