← The reading
The bibliography
§10 · Model Evaluation and Benchmarks
CAPTURED
13 On the list1 Starred0 In the Atlas13 To read
Rungs
Where this section moves a numberNo study unit draws on this section yet.
The pick
The source author's must-read for this sectionThe papers
13 papers- Benchmark^2: Systematic Evaluation of LLM Benchmarks ↗
- OdysseyArena: Benchmarking Large Language Models for Long-Horizon, Active and Inductive Interactions ↗
- Large Language Model Reasoning Failures ↗
- Emergent Misalignment Is Easy, Narrow Misalignment Is Hard ↗
- Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips ↗
- LLMStructBench: Benchmarking Large Language Model Structured Data Extraction ↗
- Quantifying Construct Validity in Large Language Model Evaluations ↗
- Lost in Stories: Consistency Bugs in Long Story Generation by LLMs ↗
- ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence ↗
- Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs ↗
- ClawArena: Benchmarking AI Agents in Evolving Information Environments ↗
- AlphaEval: Evaluating Agents in Production ↗★
- Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use ↗
What counts as read
A paper is on this list because someone worth reading put it there. That is a pointer, not a claim: it counts as read only once it has a close-read file in kb/<topic>/raw/, which is what a close-read link on a row means. There is deliberately nothing to tick off here — the study desk is the only writer of study state, and a second way to mark something done is a second version of the truth.