← all investmentsHC / investment note

INVESTMENT NOTE

Groq

Making real-time AI fast and economical.

Founded
2016
Stage
Acquihired
  • Validation

    NVIDIA licensed Groq technology for its Vera Rubin platform.

  • Funding

    $650M raised in June 2026 to scale its inference cloud.

  • Traction

    5M+ developers and 13 data centers (company-reported).

As AI moves from experimentation into everyday products, inference—the work required to serve every answer—has become a critical bottleneck. Developers need responses that are fast, predictable and economical enough to run at enormous scale.

Groq’s Language Processing Unit is engineered specifically for inference. Software-scheduled execution makes performance predictable, while on-chip memory keeps data close to the processor and avoids much of the scheduling and memory overhead found in general-purpose GPU systems. The result is high token throughput, very low response latency and more efficient use of hardware. GroqCloud exposes that speed through a familiar API, turning architectural efficiency into a lower cost to serve real-time AI at scale.

  • Founder Jonathan Ross designed core elements of Google’s first TPU.
  • After Groq’s December 2025 technology-licensing agreement with NVIDIA, Ross, president Sunny Madra and several team members joined NVIDIA while Groq remained independent.
  • CEO Adam Winter and CFO Matt Eng now lead a team with experience across Meta data centers, xAI, Microsoft and enterprise software — deep systems knowledge paired with commercial execution.
  • NVIDIA integrated Groq technology into its Vera Rubin platform as the Groq 3 LPX inference accelerator.
  • Raised $650 million in June 2026 to expand its inference cloud.
  • Reports 13 data centers, more than five million developers and trillions of tokens processed each week; operating figures are company-reported.
Groq LPU processor mounted on an orange accelerator board
Built specifically for inference, Groq’s LPU delivers fast, predictable token generation while reducing the hardware overhead—and cost—of serving AI models.Image: Groq

Selected reading

4 sources