Groq
companyCompany dossier
Groq — the primary-source profile
Groq
Type: Company (AI inference accelerator vendor)
Groq enters this KB for one specific and useful reason: GroqRack is the non-GPU inference platform in the first independent, phase-separated benchmark this topic has acquired. Everything recorded here comes from that third-party evaluation — not from Groq.
What the independent benchmark found
Usami, Vishwanath & Bethel, "Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators", arXiv:2606.17104, HPAI4S'26 @ IEEE IPDPS 2026. Common model: Llama2-7B.
- GroqRack achieves significantly lower TPOT (time per output token) during the decode phase — the memory-bandwidth-bound half of inference.
- Batching is not currently supported on the platform, per the paper — so the decode result is a small-batch / single-stream comparison.
- GPUs consistently win the compute-intensive prefill phase, and regain the decode throughput advantage as batch size increases.
The honest summary is the authors': each platform has distinct phase-dependent strengths. Groq's demonstrated seat is low-batch, latency-sensitive decode. That is a real and commercially meaningful niche — interactive agents, single-user long generation — but it is not the high-batch throughput regime where hyperscaler serving economics live.
Provenance. The paper's abstract was verified against arXiv; full text was not ingested, so no TTFT/TPOT magnitudes are recorded anywhere in this KB. Directions only. Do not quote a Groq speedup number from this KB — none has been read.
What This KB Does NOT Know About Groq
Effectively everything a dossier would need: architecture (LPU design, SRAM-vs-DRAM memory model), capacity, funding and valuation, customers, pricing, manufacturing partner, roadmap, and whether the lack of batching support is an architectural property or a software gap. No Groq primary source exists here. Confidence is low by construction — this is a stub anchored to one benchmark, not a company profile.
Key Contributions
- GroqRack wins decode TPOT vs GPUs at small batch on Llama2-7B; loses prefill; loses decode throughput as batch grows. Evidence: moderate — independent (non-vendor) peer-reviewed workshop benchmark, abstract-level only, directional; batching unsupported at time of measurement (Argonne)
Mentioned In
- Prefill/Decode Disaggregation — the benchmark that anchors this page
- Custom Silicon vs GPU — non-GPU silicon's demonstrated seat
Related Entities
- NVIDIA — the GPU side of the comparison
Changelog
- 2026-07-22 — Stub created from the Argonne prefill/decode evaluation. Deliberately minimal: one sourced claim, explicit "what we don't know", and a standing warning that no magnitude figures were read.