Skip to content
Rung 05 DataSwitch rung
Rung 05 / What it learns from

Data

What the model is trained and grounded on, and what that constrains.

3
Sources
3
Concepts
4
Entities
A stack of polished steel drive platters on a copper spindle, on paper.

In scope: training-data supply and exhaustion, synthetic data, licensing and provenance, retrieval corpora, data quality effects on capability per token, contamination. Out: retrieval as a harness mechanism (harnesses), model architecture (models).

A banded drill core in a tray, the thin copper band marking the clean licensed seam.
SyntheticLicensingProvenanceExhaustion
Sources compiled for this topic
TypeSourcePublished
REPORTReddit, Inc. — Form S-1 Registration Statement (data licensing arrangements, January 2024)
Reddit, Inc. · Reddit, Inc. (SEC filing, CIK 0001713445)

A negotiated, primary-source price on licensed training data: January 2024 data licensing arrangements with an aggregate contract value of $203.0M over 2–3 year terms, at least $66.4M to be recognised in 2024, substantially all from one (unnamed) partner — against $15.2M of ALL 'other revenue' in 2023. Ingested as the fallback after the Q2 2026 10-Q (DATA-1) failed its check: its $43.3M is 'other revenue', not a licensing line. 31 months old — never cite as current.

2024-02-22
REPORTCan AI scaling continue through 2030? (Epoch AI)
Jaime Sevilla, Tamay Besiroglu, Ben Cottier, Josh You, Edu Roldán, Pablo Villalobos, Ege Erdil · Epoch AI

Indexed web ≈ 500T tokens (100T–3,000T) vs ~15T in the largest known training sets; text 'data wall' in about five years at 4x/yr compute. The exhaustion window it summarises (Epoch's June 2024 data paper, read alongside) MOVED: 2022 said high-quality text gone before 2026 (median ~2024); 2024 says the ~300T-token effective stock is fully used 2026–2032 (80% CI), central 2028 compute-optimal, 2027 at 5x overtraining. Ranks data behind power and chips as a constraint.

2024-08-20
ANALYSISWhat Authors Need to Know About the Anthropic Settlement (Bartz v. Anthropic, $1.5bn class settlement)
The Authors Guild · The Authors Guild

A court-set per-unit price for pirated training data: ~$3,000 per title across ~500,000 qualifying titles from a $1.5bn fund, against 7 million copies from LibGen and PiLiMi. The fair-use ruling covered TRAINING; the $1.5bn covered ACQUISITION — the distinction most coverage collapses. Final-approval date unresolved (hearing 2026-05-14 per this source).

2025-09-25