Custom Silicon vs GPU

Active Frontier
Sign in to track mastery·Sign in·Practice anyway
custom-siliconasicgputputrainiumsemiconductor-economics

Custom Silicon vs GPU

The AI compute market is experiencing a structural inflection. Custom ASICs built by hyperscalers — Google TPU v7 (Ironwood), Amazon Trainium 3, Microsoft Maia 200, Meta MTIA, OpenAI Titan — are growing at 44.6% CAGR versus 16.1% for GPU-based solutions. Yet this isn't eroding NVIDIA's dominance as simply as headlines suggest. The competitive dynamic is more nuanced: NVIDIA is moving from chip-level to system-level lock-in, while ASICs capture the inference long-tail.

2026 is the year the inflection became measurable in shipments, not just CAGRs. TrendForce's January 2026 forecast puts ASIC-based systems at 27.8% of AI server shipments in 2026 — the highest share since 2023 — with GPUs holding 69.7%. The two numbers together capture the nuance the headlines miss: ASIC growth now outpaces merchant-GPU growth (the first year that has happened), yet GPUs still dominate the installed mix ~7-to-2. The driver is capex: the top five North American CSPs (Google, AWS, Meta, Microsoft, Oracle) are raising capital expenditure ~40% YoY in 2026, and an increasing slice of that budget diverts toward internally-designed chips that generate no revenue for NVIDIA.

The merchant ASIC layer is a duopoly. Hyperscalers don't design these chips alone — they co-design with Broadcom (~60% of the custom-AI-ASIC design market: TPU, MTIA, Maia, OpenAI Titan) and Marvell (~35%: Trainium, Maia). Marvell projects up to ~$11B of AI ASIC revenue in 2026. The investment read-through is that Broadcom and Marvell capture custom-silicon value regardless of which hyperscaler's chip wins — they are the "arms dealers" of the ASIC war.

The demand side is now explicit. The clearest single signal that the inflection is real (not just supply-side cost engineering) is Anthropic's commitment to ~1 million Google TPUs (Ironwood) and well over a gigawatt of capacity online in 2026 — a deal worth tens of billions, with a further multi-gigawatt Google/Broadcom expansion (~5 GW) starting 2027. A frontier lab is voting gigawatt-scale capital for a hyperscaler ASIC over merchant GPUs. But even Anthropic keeps a three-platform strategy (TPUs + AWS Trainium + NVIDIA GPUs), which is exactly why GPU absolute volume keeps rising even as share erodes.

SemiAnalysis argues that NVIDIA's real moat isn't the GPU die — it's the co-designed rack. The Vera Rubin NVL72 integrates six chip types (GPU + CPU + four networking/security chips), cooling, power delivery, and software into a single product. No ASIC vendor matches this full-stack integration. Custom ASICs win on cost efficiency for known, stable workloads and power efficiency for narrow use cases, but they operate in a fundamentally different competitive space than NVIDIA's system-level offering.

NVIDIA's inference market share is projected to fall from 90%+ to 20-30% by 2028. But total inference market size is expanding so rapidly that NVIDIA's absolute revenue grows even as share drops. The premium training and cutting-edge inference market — where models change quarterly and generality matters — remains NVIDIA's stronghold through system integration and annual architecture cadence.

A key wildcard is the open-source software ecosystem. Triton and vLLM could reduce CUDA lock-in over time, making it easier for custom ASICs to capture workloads that currently default to NVIDIA due to software compatibility.

Key Claims

  • Custom ASIC CAGR 44.6% vs GPU 16.1% — Hyperscaler chip programs growing nearly 3x faster than GPU solutions; corroborated by TrendForce for 2026. Evidence: strong (multi-source) (SemiAnalysis, TrendForce)
  • ASICs = 27.8% of AI server shipments in 2026; GPUs = 69.7% — Highest ASIC share since 2023; first year ASIC growth outpaces merchant GPUs, yet GPUs still dominate the mix. Evidence: moderate (TrendForce forecast) (TrendForce)
  • Broadcom ~60% / Marvell ~35% of the ASIC co-design market — Merchant duopoly captures custom-silicon value regardless of hyperscaler winner; Marvell ~$11B AI ASIC revenue 2026. Evidence: moderate (summary-derived) (Tom's Hardware)
  • Anthropic commits ~1M Google TPUs + >1 GW capacity in 2026 — Demand-side proof of the inflection; tens of billions, with a ~5 GW Google/Broadcom expansion from 2027. Evidence: strong (first-party) (Anthropic)
  • NVIDIA inference share: 90% to 20-30% by 2028 — But absolute revenue grows as total market expands. Evidence: moderate (projection) (SemiAnalysis)
  • System-level lock-in replaces chip-level lock-in — Rack co-design (6 chips + cooling + power + software) is the new moat. Evidence: strong (SemiAnalysis)
  • Open-source ecosystem could reduce CUDA lock-in — Triton, vLLM lower switching costs for custom ASICs. Evidence: moderate (emerging trend) (SemiAnalysis)
  • Custom silicon removes merchant margin but not the memory premium — HBM is 45–50% of a custom accelerator's BOM, priced by the same three suppliers regardless of logic-die owner. Evidence: moderate (single-source model) (Matsuoka)
  • Best-case custom build ≈ $0.082/PB vs a $0.022–0.037/PB depreciated-incumbent floor — parity with merchant GPUs, still 2.2–3.3× above an amortized fleet. Evidence: moderate (model output under stipulated parameters) (close read)
  • 25% success / 34% mediocre / 41% loss for a custom program. Evidence: weak — elicited subjective probability revised across five rounds; explicitly NOT model output (close read)
  • Accelerator advantage is phase-dependent — GPUs win prefill; a non-GPU inference rack wins decode TPOT at small batch; GPUs regain decode throughput at large batch. Evidence: moderate — independent workshop benchmark, abstract-level, directional only (Argonne)

The Two Things Custom Silicon Cannot Design Around

The shipment-share and co-design numbers above describe who is designing chips. Two constraints determine who can actually deploy them, and neither is addressable by a better logic die:

  1. The memory premium. HBM is 45–50% of a custom accelerator's bill of materials and is "priced by the same three suppliers under the same shortage" regardless of who owns the logic design. Removing NVIDIA's merchant margin (~20–35% of cost) therefore gets a well-executed custom build to roughly $0.082/PB — approximate merchant-GPU parity — against a depreciated-incumbent floor of $0.022–0.037/PB, i.e. still 2.2–3.3× more expensive than an amortized fleet. The competitor a new custom program has to beat is not NVIDIA's price list; it is someone else's sunk capital (Memory Scarcity & Inference Economics).
  2. The interposer queue. NVIDIA has reserved the majority (~60%) of TSMC's advanced-packaging capacity; hyperscaler ASIC programs queue behind it for the same CoWoS slots at the same foundry (Advanced Packaging & CoWoS Capacity).

A frequently-quoted 25% success / 34% mediocre / 41% loss outcome distribution for a custom accelerator program comes from the same source — but it is elicited subjective probability, revised across five rounds, and explicitly labelled by its author as judgmental assessment rather than model output. Do not present it as a derived result. Stated mitigants that roughly double success probability: an anchor demand customer, a secured HBM long-term agreement, or a non-HBM architecture; staged go/no-go capital gates cut deployed-capital loss probability to ~24%.

Settling the Argument: Benchmark by Phase, Not by Chip

Most GPU-vs-ASIC comparisons compare the wrong thing. Inference splits into a compute-bound prefill phase and a memory-bandwidth-bound decode phase, and accelerator advantage flips between them. The first independent (non-vendor) phase-separated benchmark in this KB — Argonne, Llama2-7B, GPUs vs GroqRack — finds GPUs consistently win prefill, GroqRack wins decode TPOT at small batch, and GPUs regain the decode throughput advantage as batch size grows (Prefill/Decode Disaggregation, Groq).

The read-through: specialized silicon has a real, demonstrable seat in low-batch latency-sensitive decode. High-batch throughput serving — where hyperscaler unit economics actually live — is where GPUs recover. "Which is faster" is not a well-formed question; "which phase, which metric, which batch regime" is.

The Hyperscaler Chip Programs

ChipCompanyGenCo-design partnerFocus / specs (2026)
TPU v7 (Ironwood)Google7thBroadcomRack-scale inference; 4,614 FP8 TFLOPS, 192 GB HBM3E, 7.37 TB/s; sold externally (Anthropic)
Trainium 3Amazon3rdMarvellTraining accelerator; first AWS 3nm part, ramps from Q2 2026
Maia 200Microsoft2ndBroadcom / MarvellCustom AI accelerator; arrived early 2026 (TSMC 3nm class)
MTIAMeta1st+BroadcomInference accelerator
TitanOpenAI1stBroadcomNew custom program (first appearance in KB)

Specs for Ironwood / Trainium3 / Maia 200 are summary-derived (Tom's Hardware) — corroborate against vendor materials before flagship use.

The NVIDIA Counter

  1. System-level integration — No ASIC vendor matches full-stack co-design
  2. Software ecosystem — CUDA, Triton, framework support create switching costs
  3. Annual cadence — Architecture updates outpace multi-year ASIC cycles
  4. Rack-as-product — Vera Rubin NVL72 sells the infrastructure, not just the chip

Open Questions

  • Will the compounding effect of multiple hyperscalers each investing billions annually in custom silicon eventually erode NVIDIA's system-level advantage?
  • Can open-source inference stacks (Triton, vLLM) meaningfully reduce CUDA lock-in within 2-3 years?
  • Do custom ASICs need to replicate rack-scale co-design, or can they win on TCO for specific workload classes?
  • How does the training vs inference market split evolve as foundation model training concentrates in fewer labs?
  • If HBM is half the BOM and priced identically for everyone, what is the actual defensible margin in a custom program — is it TCO, or is it just supply access?
  • Do any hyperscaler ASIC programs publish phase-separated (prefill vs decode) benchmarks, or does the comparison stay apples-to-oranges?
  • Does an ASIC program with a secured HBM long-term agreement measurably outperform one without?

Related Concepts

Backlinks

Pages that reference this concept:

Changelog

  • 2026-07-22 — Added the two constraints custom silicon cannot design around (HBM at 45–50% of BOM with a $0.082/PB best-case build vs a $0.022–0.037/PB incumbent floor; the CoWoS interposer queue), the 25/34/41 outcome split labelled as elicited subjective probability rather than a derived result, and the phase-aware benchmarking frame from the independent Argonne GPU-vs-GroqRack study. +3 sources.
  • 2026-06-24 — Added the 2026 shipment-share quantification (TrendForce: ASIC 27.8% / GPU 69.7%, +28% YoY server growth, CSP capex +40%), the Broadcom ~60% / Marvell ~35% co-design duopoly + Marvell ~$11B 2026 ASIC revenue, Ironwood chip specs and Trainium3/Maia 200 timing, the new OpenAI Titan program, and the Anthropic ~1M-TPU / >1 GW demand anchor. +3 sources.
  • 2026-04-09 — Initial compile from SemiAnalysis co-design analysis + Vera Rubin platform.
Custom Silicon vs GPU | KB | MenFem