Groq

company
inference-hardwareai-acceleratorgroqrackprefill-decodenon-gpu

Company dossier

Groq — the primary-source profile

Groq

Type: Company (AI inference accelerator vendor)

Groq enters this KB for one specific and useful reason: GroqRack is the non-GPU inference platform in the first independent, phase-separated benchmark this topic has acquired. Everything recorded here comes from that third-party evaluation — not from Groq.

What the independent benchmark found

Usami, Vishwanath & Bethel, "Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators", arXiv:2606.17104, HPAI4S'26 @ IEEE IPDPS 2026. Common model: Llama2-7B.

  • GroqRack achieves significantly lower TPOT (time per output token) during the decode phase — the memory-bandwidth-bound half of inference.
  • Batching is not currently supported on the platform, per the paper — so the decode result is a small-batch / single-stream comparison.
  • GPUs consistently win the compute-intensive prefill phase, and regain the decode throughput advantage as batch size increases.

The honest summary is the authors': each platform has distinct phase-dependent strengths. Groq's demonstrated seat is low-batch, latency-sensitive decode. That is a real and commercially meaningful niche — interactive agents, single-user long generation — but it is not the high-batch throughput regime where hyperscaler serving economics live.

Provenance. The paper's abstract was verified against arXiv; full text was not ingested, so no TTFT/TPOT magnitudes are recorded anywhere in this KB. Directions only. Do not quote a Groq speedup number from this KB — none has been read.

What This KB Does NOT Know About Groq

Effectively everything a dossier would need: architecture (LPU design, SRAM-vs-DRAM memory model), capacity, funding and valuation, customers, pricing, manufacturing partner, roadmap, and whether the lack of batching support is an architectural property or a software gap. No Groq primary source exists here. Confidence is low by construction — this is a stub anchored to one benchmark, not a company profile.

Key Contributions

  • GroqRack wins decode TPOT vs GPUs at small batch on Llama2-7B; loses prefill; loses decode throughput as batch grows. Evidence: moderate — independent (non-vendor) peer-reviewed workshop benchmark, abstract-level only, directional; batching unsupported at time of measurement (Argonne)

Mentioned In

Related Entities

  • NVIDIA — the GPU side of the comparison

Changelog

  • 2026-07-22 — Stub created from the Argonne prefill/decode evaluation. Deliberately minimal: one sourced claim, explicit "what we don't know", and a standing warning that no magnitude figures were read.