The Latency Bottleneck in Large Language Models
While GPUs excel at massively parallel training computations, autoregressive token generation is fundamentally a memory-bandwidth-bound sequential operation. Groq designed a custom architecture from scratch: the Language Processing Unit (LPU).
Eliminating HBM in Favor of On-Die SRAM
Unlike GPUs that rely on external High Bandwidth Memory (HBM), Groq's chip embeds 230 megabytes of ultra-fast static RAM (SRAM) directly on the die, delivering an astonishing 80 terabytes per second of memory bandwidth and generating 500+ tokens per second for real-time voice and agentic applications.