Rung 04 Evaluation & MeasurementSwitch rungClose
Holistic Agent Leaderboard (HAL)
organizationType: Public leaderboard and research project (Princeton; Kapoor et al., ICLR 2026)
HAL runs standard agent set-ups across nine benchmarks — coding, web navigation, science tasks and customer service — with shared scaffolds and central cost tracking, and plots accuracy against cost rather than accuracy alone. That makes it the one public source on this rung that publishes what its numbers cost to produce: about $40,000 for 21,730 runs across 9 models and 9 benchmarks, single seed, grown to 26,597 runs by April 2026 (AI evals are becoming the new compute bottleneck).
Its own audits are also where several of this rung's validity findings come from: a "do-nothing" agent passing 38% of τ-bench airline tasks as originally built, and data leakage found in one scaffold's logs that led to that scaffold's removal in December 2025. HAL has paused new model evaluations to work on reliability (same source).
HAL is known here only through the EvalEval article; its own paper (arXiv:2510.11977) has not been ingested.
Mentioned In
- The Cost of Evaluation — the $40,000 single-seed total and the ~$320,000 eight-repeat figure.
- Benchmark Validity — the do-nothing agent result.
- AI evals are becoming the new compute bottleneck
Related Entities
- UK AI Security Institute (AISI) — the other evaluator on this rung that owns no model.