Skip to content
Rung 05 DataSwitch rung
REPORT2024-08-20 · Epoch AI

Can AI scaling continue through 2030? (Epoch AI)

Jaime Sevilla, Tamay Besiroglu, Ben Cottier, Josh You, Edu Roldán, Pablo Villalobos, Ege Erdil
Compiled notes
What it moved

Indexed web ≈ 500T tokens (100T–3,000T) vs ~15T in the largest known training sets; text 'data wall' in about five years at 4x/yr compute. The exhaustion window it summarises (Epoch's June 2024 data paper, read alongside) MOVED: 2022 said high-quality text gone before 2026 (median ~2024); 2024 says the ~300T-token effective stock is fully used 2026–2032 (80% CI), central 2028 compute-optimal, 2027 at 5x overtraining. Ranks data behind power and chips as a constraint.

How it was read: three Epoch AI pages were fetched on 2026-09-24 and read in full as text:

  1. The named source — "Can AI scaling continue through 2030?" (report, 20 August 2024). Its Data scarcity section, the summary and the conclusion were read closely; the power, chip and latency sections were read for the ranking of constraints only.
  2. The exhaustion window it summarises — "Will we run out of data? Limits of LLM scaling based on human-generated data" (Villalobos et al., paper page, 6 June 2024, https://epoch.ai/publications/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data). The 2030 report says in its data section that it "summarize[s] our previous work on data scarcity"; the dated window is stated on this page, not on the 2030 page.
  3. The earlier estimate it replaced — "Will we run out of ML data? Evidence from projecting dataset size trends" (paper page, 10 November 2022, https://epoch.ai/blog/will-we-run-out-of-ml-data-evidence-from-projecting-dataset).

Only the Epoch web pages were read, not the underlying arXiv PDFs. None of the three pages shows a revision date or changelog as fetched; "the window moved" is established by the 2024 paper's own comparison with the 2022 one, quoted below.

The window, and how it moved

EstimatePublishedWhat runs outWhen
Epoch 20222022-11-10High-quality language data"before 2026" — median 2024.5 (90% CI 2023.5–2025.7) on the historical projection, 2024.1 (2023.2–2025.3) on the compute projection
Epoch 20222022-11-10Low-quality language data2030 to 2050
Epoch 20242024-06-06The ~300 trillion-token effective stock of public human-written text (quality- and repetition-adjusted)Between 2026 and 2032 (80% confidence interval)
Epoch 2024, compute-optimal training2024-06-06SameEnough for a ~5e28 FLOP model, "a level we expect to be reached in 2028"
Epoch 2024, overtrained 5×2024-06-06Same2027
Epoch 2024, overtrained 100×2024-06-06Same2025 (Llama 3-70B, for scale, was overtrained 10×)

The move, in Epoch's own words (2024 page): "Our 2022 paper predicted that high-quality text data would be fully used by 2024, whereas our new results indicate that might not happen until 2028." The stated reasons are a change of method and two findings: carefully filtered web data trains as well as curated sources, and models can be trained on the same data for several passes, which "further expanded our estimate of the effective stock by a factor of 2-5x".

A correction to the proposal. DATA-3 was proposed on a search summary that the date had "moved later, to around 2028 at the earliest, and to around 2027 under heavy overtraining". On the page, 2028 is the central compute-optimal date, not the earliest; the window is 2026–2032; 2027 is modest (5×) overtraining, and heavy (100×) overtraining gives 2025.

Where that leaves the date today (2026-09-24): the current window opened this year. The headline "we run out in 2026", still widely repeated, is a misreading of the 2022 figure (which said before 2026, for high-quality text only) and of the 2024 range (which starts in 2026).

What the 2030 report adds

  • The stock. The largest known training sets are "on the order of 15 trillion tokens" of public text and code. The indexed web holds about 500 trillion tokens after deduplication — 30× more — with a range from 100T (compiled corpora such as Common Crawl only) to 3,000T (counting private data).
  • The wall. Using the whole indexed web would allow 30× more data and 30× more parameters, about 900× the compute, "up to 8e28 FLOP". At the recent 4× a year growth in training compute, "we would run into this 'data wall' for text data in about five years" — about 2029, counting from the August 2024 publication (the year is this KB's arithmetic).
  • Other modalities push it out, not away. Adding image, video and audio gives an effective stock of 400 trillion to 20 quadrillion tokens (the summary's range; the data section itself says 450 trillion to 23 quadrillion — the page disagrees with itself slightly), enough for training runs of 6e28 to 2e32 FLOP by 2030. The main assumptions (read from the page's parameter table): the text stock grows 3–10% a year (median 7%); only 10–40% of it survives quality filtering (median 20%); it is reused for 3–15 passes (median 5).
  • Synthetic data is excluded from those numbers, deliberately: "we conservatively rely on estimates from multimodal data, excluding all types of synthetic data." Epoch's rough estimate is that generating a model's training set synthetically would about double the compute needed to train it; it flags model collapse as the main risk.
  • Copyright is treated as a quality problem, not a quantity problem: restrictions "may disproportionately affect high-quality sources such as books and reputable news outlets."
  • Data is not the first limit. "The constraint likely to bind first is power, followed by the capacity to manufacture enough chips." The report's bottom line is that ~2e29 FLOP training runs are likely feasible by 2030.

What number this moves

The data rung had no source on supply. This gives it two numbers and a date:

  • ~500T tokens of indexed web text against ~15T used — a 30× headroom, as of August 2024.
  • 2026–2032 for the ~300T-token effective stock to be fully used, central 2028 — Epoch's current window, moved four years later than its 2022 estimate.

How it reaches the token price: a data wall raises the cost of making a better model (licensed, private or synthetic data instead of free web text — synthetic data at roughly twice the training compute), which is paid once and spread across every token the model serves. It does not change the cost of serving an existing model. This rung still cannot convert either into cents per token.

Limitations

  • Forecasts, not measurements, and Epoch's own forecast has already moved once by four years. The 2026–2032 range is an 80% interval; 2028 depends on how models are trained.
  • Aged. The 2030 report dates from August 2024 and the window from June 2024. Training sets larger than "15T tokens" have been reported since; this KB has not checked a newer Epoch update.
  • Public text only. Private data (the 3,000T upper bound), licensed corpora and synthetic data sit outside the headline stock.
  • Epoch AI is an independent research organisation; the 2030 report and the data paper share authors (Sevilla, Besiroglu, Villalobos), so they are not independent confirmations of each other.

Standing

I only read about it. A piece built on this is a coverage piece. The teachable content is the moving date itself — why a forecast of "running out" slid from 2024 to 2028, and what "runs out" means (the stock being fully used by a frontier training run, not the internet emptying).

No live market call rests on this rung (gate 3 unmet on data).


Sources: Can AI scaling continue through 2030? (Epoch AI, 2024-08-20) · Will we run out of data? (2024-06-06) · Will we run out of ML data? (2022-11-10).

Related in the base