Towards Topology-Aware Very Large-Scale Photonic AI Accelerators
Identifies the 'Utilization Wall' — topology-dominated scaling bottleneck in photonic accelerators; symmetric-grid topologies recover up to 6x utilization, cut memory access >40%
Towards Topology-Aware Very Large-Scale Photonic AI Accelerators
Abstract
Addresses scaling limits in photonic AI accelerators by proposing a modular, scale-out architecture built from 4×4 photonic tensor core units, incorporating hardware-realistic constraints (insertion loss, fanout penalties, laser power limits) that conventional monolithic-scaling approaches ignore.
Key Contributions
- Modular 4×4 photonic tensor core scale-out architecture — an alternative to monolithic photonic accelerator scaling.
- Hardware-realistic constraint modeling — insertion loss, fanout penalties, and laser-power limits are incorporated directly, rather than assumed away.
- The "Utilization Wall" — a novel, photonics-specific bottleneck: at scale, performance becomes governed by grid topology rather than raw hardware size (number of processing elements). This is a distinct failure mode from the electronic "memory wall."
- Symmetric Grid Rule — symmetric topologies achieve up to 6x utilization improvement and reduce memory access by >40% compared to linear/asymmetric configurations of the same nominal hardware size.
Methodology / Results
Evaluated on representative DNN workloads (GoogleNet, ResNet-18, MobileNet, AlphaGo Zero), scaling up to 1,024 processing elements to characterize where and why utilization collapses as photonic accelerators scale out.
Why It Matters (lens: optical-computing — photonic compute, long-term / what to watch)
Directly extends the KB's existing SimPhony finding (harnessing-photonics-for-machine-intelligence.md) — where SimPhony showed peripheral (DAC/ADC) overheads dominate photonic AI energy budgets at the device level, this paper shows an analogous topology-driven bottleneck at the system/scale-out level. Together they argue that photonic AI accelerators face two independent, compounding scaling walls (device-level peripheral overhead + system-level topology), not one — a materially more cautious picture than vendor marketing (Q.ANT, Lightmatter) typically presents, and directly relevant to the LENS's forbidden pattern: "hype about replacing GPUs without a benchmark."
Limitations
- Full author list and institutional affiliation were not resolved from the fetched abstract page — flag for a follow-up fetch of the arXiv HTML/PDF to complete proper attribution.
- Evaluated via simulation on representative workloads, not a fabricated/measured hardware demonstration.
Source: Towards Topology-Aware Very Large-Scale Photonic AI Accelerators — arXiv preprint 2604.26966, April 2026.