Custom Silicon vs GPU
Active FrontierCustom Silicon vs GPU
The AI compute market is experiencing a structural inflection. Custom ASICs built by hyperscalers — Google TPU v7 (Ironwood), Amazon Trainium 3, Microsoft Maia 200, Meta MTIA, OpenAI Titan — are growing at 44.6% CAGR versus 16.1% for GPU-based solutions. Yet this isn't eroding NVIDIA's dominance as simply as headlines suggest. The competitive dynamic is more nuanced: NVIDIA is moving from chip-level to system-level lock-in, while ASICs capture the inference long-tail.
2026 is the year the inflection became measurable in shipments, not just CAGRs. TrendForce's January 2026 forecast puts ASIC-based systems at 27.8% of AI server shipments in 2026 — the highest share since 2023 — with GPUs holding 69.7%. The two numbers together capture the nuance the headlines miss: ASIC growth now outpaces merchant-GPU growth (the first year that has happened), yet GPUs still dominate the installed mix ~7-to-2. The driver is capex: the top five North American CSPs (Google, AWS, Meta, Microsoft, Oracle) are raising capital expenditure ~40% YoY in 2026, and an increasing slice of that budget diverts toward internally-designed chips that generate no revenue for NVIDIA.
The merchant ASIC layer is a duopoly. Hyperscalers don't design these chips alone — they co-design with Broadcom (~60% of the custom-AI-ASIC design market: TPU, MTIA, Maia, OpenAI Titan) and Marvell (~35%: Trainium, Maia). Marvell projects up to ~$11B of AI ASIC revenue in 2026. The investment read-through is that Broadcom and Marvell capture custom-silicon value regardless of which hyperscaler's chip wins — they are the "arms dealers" of the ASIC war.
The demand side is now explicit. The clearest single signal that the inflection is real (not just supply-side cost engineering) is Anthropic's commitment to ~1 million Google TPUs (Ironwood) and well over a gigawatt of capacity online in 2026 — a deal worth tens of billions, with a further multi-gigawatt Google/Broadcom expansion (~5 GW) starting 2027. A frontier lab is voting gigawatt-scale capital for a hyperscaler ASIC over merchant GPUs. But even Anthropic keeps a three-platform strategy (TPUs + AWS Trainium + NVIDIA GPUs), which is exactly why GPU absolute volume keeps rising even as share erodes.
SemiAnalysis argues that NVIDIA's real moat isn't the GPU die — it's the co-designed rack. The Vera Rubin NVL72 integrates six chip types (GPU + CPU + four networking/security chips), cooling, power delivery, and software into a single product. No ASIC vendor matches this full-stack integration. Custom ASICs win on cost efficiency for known, stable workloads and power efficiency for narrow use cases, but they operate in a fundamentally different competitive space than NVIDIA's system-level offering.
NVIDIA's inference market share is projected to fall from 90%+ to 20-30% by 2028. But total inference market size is expanding so rapidly that NVIDIA's absolute revenue grows even as share drops. The premium training and cutting-edge inference market — where models change quarterly and generality matters — remains NVIDIA's stronghold through system integration and annual architecture cadence.
A key wildcard is the open-source software ecosystem. Triton and vLLM could reduce CUDA lock-in over time, making it easier for custom ASICs to capture workloads that currently default to NVIDIA due to software compatibility.
Key Claims
- Custom ASIC CAGR 44.6% vs GPU 16.1% — Hyperscaler chip programs growing nearly 3x faster than GPU solutions; corroborated by TrendForce for 2026. Evidence: strong (multi-source) (SemiAnalysis, TrendForce)
- ASICs = 27.8% of AI server shipments in 2026; GPUs = 69.7% — Highest ASIC share since 2023; first year ASIC growth outpaces merchant GPUs, yet GPUs still dominate the mix. Evidence: moderate (TrendForce forecast) (TrendForce)
- Broadcom ~60% / Marvell ~35% of the ASIC co-design market — Merchant duopoly captures custom-silicon value regardless of hyperscaler winner; Marvell ~$11B AI ASIC revenue 2026. Evidence: moderate (summary-derived) (Tom's Hardware)
- Anthropic commits ~1M Google TPUs + >1 GW capacity in 2026 — Demand-side proof of the inflection; tens of billions, with a ~5 GW Google/Broadcom expansion from 2027. Evidence: strong (first-party) (Anthropic)
- NVIDIA inference share: 90% to 20-30% by 2028 — But absolute revenue grows as total market expands. Evidence: moderate (projection) (SemiAnalysis)
- System-level lock-in replaces chip-level lock-in — Rack co-design (6 chips + cooling + power + software) is the new moat. Evidence: strong (SemiAnalysis)
- Open-source ecosystem could reduce CUDA lock-in — Triton, vLLM lower switching costs for custom ASICs. Evidence: moderate (emerging trend) (SemiAnalysis)
- Custom silicon removes merchant margin but not the memory premium — HBM is 45–50% of a custom accelerator's BOM, priced by the same three suppliers regardless of logic-die owner. Evidence: moderate (single-source model) (Matsuoka)
- Best-case custom build ≈ $0.082/PB vs a $0.022–0.037/PB depreciated-incumbent floor — parity with merchant GPUs, still 2.2–3.3× above an amortized fleet. Evidence: moderate (model output under stipulated parameters) (close read)
- 25% success / 34% mediocre / 41% loss for a custom program. Evidence: weak — elicited subjective probability revised across five rounds; explicitly NOT model output (close read)
- Accelerator advantage is phase-dependent — GPUs win prefill; a non-GPU inference rack wins decode TPOT at small batch; GPUs regain decode throughput at large batch. Evidence: moderate — independent workshop benchmark, abstract-level, directional only (Argonne)
The Two Things Custom Silicon Cannot Design Around
The shipment-share and co-design numbers above describe who is designing chips. Two constraints determine who can actually deploy them, and neither is addressable by a better logic die:
- The memory premium. HBM is 45–50% of a custom accelerator's bill of materials and is "priced by the same three suppliers under the same shortage" regardless of who owns the logic design. Removing NVIDIA's merchant margin (~20–35% of cost) therefore gets a well-executed custom build to roughly $0.082/PB — approximate merchant-GPU parity — against a depreciated-incumbent floor of $0.022–0.037/PB, i.e. still 2.2–3.3× more expensive than an amortized fleet. The competitor a new custom program has to beat is not NVIDIA's price list; it is someone else's sunk capital (Memory Scarcity & Inference Economics).
- The interposer queue. NVIDIA has reserved the majority (~60%) of TSMC's advanced-packaging capacity; hyperscaler ASIC programs queue behind it for the same CoWoS slots at the same foundry (Advanced Packaging & CoWoS Capacity).
A frequently-quoted 25% success / 34% mediocre / 41% loss outcome distribution for a custom accelerator program comes from the same source — but it is elicited subjective probability, revised across five rounds, and explicitly labelled by its author as judgmental assessment rather than model output. Do not present it as a derived result. Stated mitigants that roughly double success probability: an anchor demand customer, a secured HBM long-term agreement, or a non-HBM architecture; staged go/no-go capital gates cut deployed-capital loss probability to ~24%.
Settling the Argument: Benchmark by Phase, Not by Chip
Most GPU-vs-ASIC comparisons compare the wrong thing. Inference splits into a compute-bound prefill phase and a memory-bandwidth-bound decode phase, and accelerator advantage flips between them. The first independent (non-vendor) phase-separated benchmark in this KB — Argonne, Llama2-7B, GPUs vs GroqRack — finds GPUs consistently win prefill, GroqRack wins decode TPOT at small batch, and GPUs regain the decode throughput advantage as batch size grows (Prefill/Decode Disaggregation, Groq).
The read-through: specialized silicon has a real, demonstrable seat in low-batch latency-sensitive decode. High-batch throughput serving — where hyperscaler unit economics actually live — is where GPUs recover. "Which is faster" is not a well-formed question; "which phase, which metric, which batch regime" is.
The Hyperscaler Chip Programs
| Chip | Company | Gen | Co-design partner | Focus / specs (2026) |
|---|---|---|---|---|
| TPU v7 (Ironwood) | 7th | Broadcom | Rack-scale inference; 4,614 FP8 TFLOPS, 192 GB HBM3E, 7.37 TB/s; sold externally (Anthropic) | |
| Trainium 3 | Amazon | 3rd | Marvell | Training accelerator; first AWS 3nm part, ramps from Q2 2026 |
| Maia 200 | Microsoft | 2nd | Broadcom / Marvell | Custom AI accelerator; arrived early 2026 (TSMC 3nm class) |
| MTIA | Meta | 1st+ | Broadcom | Inference accelerator |
| Titan | OpenAI | 1st | Broadcom | New custom program (first appearance in KB) |
Specs for Ironwood / Trainium3 / Maia 200 are summary-derived (Tom's Hardware) — corroborate against vendor materials before flagship use.
The NVIDIA Counter
- System-level integration — No ASIC vendor matches full-stack co-design
- Software ecosystem — CUDA, Triton, framework support create switching costs
- Annual cadence — Architecture updates outpace multi-year ASIC cycles
- Rack-as-product — Vera Rubin NVL72 sells the infrastructure, not just the chip
Open Questions
- Will the compounding effect of multiple hyperscalers each investing billions annually in custom silicon eventually erode NVIDIA's system-level advantage?
- Can open-source inference stacks (Triton, vLLM) meaningfully reduce CUDA lock-in within 2-3 years?
- Do custom ASICs need to replicate rack-scale co-design, or can they win on TCO for specific workload classes?
- How does the training vs inference market split evolve as foundation model training concentrates in fewer labs?
- If HBM is half the BOM and priced identically for everyone, what is the actual defensible margin in a custom program — is it TCO, or is it just supply access?
- Do any hyperscaler ASIC programs publish phase-separated (prefill vs decode) benchmarks, or does the comparison stay apples-to-oranges?
- Does an ASIC program with a secured HBM long-term agreement measurably outperform one without?
Related Concepts
- Rack-Scale AI Compute — NVIDIA's system-level response to ASIC competition
- HBM4 Memory Architecture — Memory bandwidth advantage that custom ASICs must match
- Memory Scarcity & Inference Economics — why removing merchant margin isn't enough
- Prefill/Decode Disaggregation — the right way to benchmark this argument
- Advanced Packaging & CoWoS Capacity — the queue every ASIC program stands in
Backlinks
Pages that reference this concept:
- Custom Silicon Inflection 2026
- NVIDIA Vera Rubin Platform
- TrendForce: ASIC Share of AI Server Shipments 2026
- Custom AI ASIC State of Play (May 2026)
- Anthropic: ~1M Google TPU / 1 GW Deal
Changelog
- 2026-07-22 — Added the two constraints custom silicon cannot design around (HBM at 45–50% of BOM with a $0.082/PB best-case build vs a $0.022–0.037/PB incumbent floor; the CoWoS interposer queue), the 25/34/41 outcome split labelled as elicited subjective probability rather than a derived result, and the phase-aware benchmarking frame from the independent Argonne GPU-vs-GroqRack study. +3 sources.
- 2026-06-24 — Added the 2026 shipment-share quantification (TrendForce: ASIC 27.8% / GPU 69.7%, +28% YoY server growth, CSP capex +40%), the Broadcom ~60% / Marvell ~35% co-design duopoly + Marvell ~$11B 2026 ASIC revenue, Ironwood chip specs and Trainium3/Maia 200 timing, the new OpenAI Titan program, and the Anthropic ~1M-TPU / >1 GW demand anchor. +3 sources.
- 2026-04-09 — Initial compile from SemiAnalysis co-design analysis + Vera Rubin platform.
Related Concepts
Advanced Packaging & CoWoS Capacity
Active FrontierHBM4 Memory Architecture
Memory Scarcity & Inference Economics ($/PB)
Active FrontierPrefill/Decode Disaggregation
Active FrontierRack-Scale AI Compute
Active FrontierTest Your Understanding
Hardware Concepts: The 2026 Substrate
Rack-scale compute, HBM4, custom silicon, quantum error correction, photonics, and memory-centric computing
Hardware Companies & Labs
Match the compute, silicon, quantum, and photonic players to their 2026 signatures
Hardware Timeline: The Substrate Shifts
Order the milestones from Google Willow through Vera Rubin's full production
The Hardware Frontier: Quantum, Foundry & ASICs
Three quantum architectures, foundry economics, the custom-silicon inflection, and the Rubin roadmap
Hardware Speed Round
Quick-fire recall on the numbers and names of the 2026 compute stack
Sourcing Frontier AI Compute in 2026
Make the chip, packaging, memory, and power calls the evidence supports
NVIDIA Dossier: recent revenue
Estimate NVIDIA's revenue across the last quarter.