INVESTMENT NOTE
Groq
Making real-time AI fast and economical.
- Validation
NVIDIA licensed Groq technology for its Vera Rubin platform.
- Funding
$650M raised in June 2026 to scale its inference cloud.
- Traction
5M+ developers and 13 data centers (company-reported).
THE PROBLEM
As AI moves from experimentation into everyday products, inference—the work required to serve every answer—has become a critical bottleneck. Developers need responses that are fast, predictable and economical enough to run at enormous scale.
THE APPROACH
Groq’s Language Processing Unit is engineered specifically for inference. Software-scheduled execution makes performance predictable, while on-chip memory keeps data close to the processor and avoids much of the scheduling and memory overhead found in general-purpose GPU systems. The result is high token throughput, very low response latency and more efficient use of hardware. GroqCloud exposes that speed through a familiar API, turning architectural efficiency into a lower cost to serve real-time AI at scale.
WHY THIS TEAM
- Founder Jonathan Ross designed core elements of Google’s first TPU.
- After Groq’s December 2025 technology-licensing agreement with NVIDIA, Ross, president Sunny Madra and several team members joined NVIDIA while Groq remained independent.
- CEO Adam Winter and CFO Matt Eng now lead a team with experience across Meta data centers, xAI, Microsoft and enterprise software — deep systems knowledge paired with commercial execution.
MOMENTUM
- NVIDIA integrated Groq technology into its Vera Rubin platform as the Groq 3 LPX inference accelerator.
- Raised $650 million in June 2026 to expand its inference cloud.
- Reports 13 data centers, more than five million developers and trillions of tokens processed each week; operating figures are company-reported.
