Skip to content
Rung 04 Evaluation & MeasurementSwitch rung

Holistic Agent Leaderboard (HAL)

organization
agent-benchmarksleaderboardevaluation-cost

Type: Public leaderboard and research project (Princeton; Kapoor et al., ICLR 2026)

HAL runs standard agent set-ups across nine benchmarks — coding, web navigation, science tasks and customer service — with shared scaffolds and central cost tracking, and plots accuracy against cost rather than accuracy alone. That makes it the one public source on this rung that publishes what its numbers cost to produce: about $40,000 for 21,730 runs across 9 models and 9 benchmarks, single seed, grown to 26,597 runs by April 2026 (AI evals are becoming the new compute bottleneck).

Its own audits are also where several of this rung's validity findings come from: a "do-nothing" agent passing 38% of τ-bench airline tasks as originally built, and data leakage found in one scaffold's logs that led to that scaffold's removal in December 2025. HAL has paused new model evaluations to work on reliability (same source).

HAL is known here only through the EvalEval article; its own paper (arXiv:2510.11977) has not been ingested.

Mentioned In

Related Entities

Related concepts