PAPER2026-04-01·Unspecified (arXiv preprint)·arXiv 2604.26966

Towards Topology-Aware Very Large-Scale Photonic AI Accelerators

(arXiv preprint authors — see paper)
COMPILED NOTES

Identifies the 'Utilization Wall' — topology-dominated scaling bottleneck in photonic accelerators; symmetric-grid topologies recover up to 6x utilization, cut memory access >40%

Towards Topology-Aware Very Large-Scale Photonic AI Accelerators

Abstract

Addresses scaling limits in photonic AI accelerators by proposing a modular, scale-out architecture built from 4×4 photonic tensor core units, incorporating hardware-realistic constraints (insertion loss, fanout penalties, laser power limits) that conventional monolithic-scaling approaches ignore.

Key Contributions

  1. Modular 4×4 photonic tensor core scale-out architecture — an alternative to monolithic photonic accelerator scaling.
  2. Hardware-realistic constraint modeling — insertion loss, fanout penalties, and laser-power limits are incorporated directly, rather than assumed away.
  3. The "Utilization Wall" — a novel, photonics-specific bottleneck: at scale, performance becomes governed by grid topology rather than raw hardware size (number of processing elements). This is a distinct failure mode from the electronic "memory wall."
  4. Symmetric Grid Rule — symmetric topologies achieve up to 6x utilization improvement and reduce memory access by >40% compared to linear/asymmetric configurations of the same nominal hardware size.

Methodology / Results

Evaluated on representative DNN workloads (GoogleNet, ResNet-18, MobileNet, AlphaGo Zero), scaling up to 1,024 processing elements to characterize where and why utilization collapses as photonic accelerators scale out.

Why It Matters (lens: optical-computing — photonic compute, long-term / what to watch)

Directly extends the KB's existing SimPhony finding (harnessing-photonics-for-machine-intelligence.md) — where SimPhony showed peripheral (DAC/ADC) overheads dominate photonic AI energy budgets at the device level, this paper shows an analogous topology-driven bottleneck at the system/scale-out level. Together they argue that photonic AI accelerators face two independent, compounding scaling walls (device-level peripheral overhead + system-level topology), not one — a materially more cautious picture than vendor marketing (Q.ANT, Lightmatter) typically presents, and directly relevant to the LENS's forbidden pattern: "hype about replacing GPUs without a benchmark."

Limitations

  • Full author list and institutional affiliation were not resolved from the fetched abstract page — flag for a follow-up fetch of the arXiv HTML/PDF to complete proper attribution.
  • Evaluated via simulation on representative workloads, not a fabricated/measured hardware demonstration.

Source: Towards Topology-Aware Very Large-Scale Photonic AI Accelerators — arXiv preprint 2604.26966, April 2026.

RELATED · IN THE BASE
Towards Topology-Aware Very Large-Scale Photonic AI Accelerators | Knowledge Base | MenFem