STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios
State-based personal-agent benchmark (40 tasks, 11 frontier models) scored by final SYSTEM-STATE correctness rather than text; reports a measured per-task API cost ~$0.17 (cheapest, MiniMax-tier) to ~$6.55 (Claude-Opus-4.7-tier) — a ~38x spread — the KB's FIRST dated absolute cost-per-task levels, partly closing the topic's sharpest cost-per-task gap. The $0.17-$6.55 figures are search-surfaced from the paper's cost table, NOT in the fetched abstract; full-text confirmation owed. API list-price-based, personal-computing scenarios only.
STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios
Abstract
"Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existing benchmarks still rely on sandboxed artifacts, static task design, and coarse scoring, which hinder scalability and limit progress toward reliable personal-agent evaluation. This paper introduces STAGE-Claw, an automated framework for building and evaluating realistic personal-agent scenarios in state-based personal-computing environments. Given a task hint, STAGE-Claw automatically creates and validates a realistic benchmark task with its environment, task prompts, ground truth, and related verification programs. Agents are then evaluated in realistic operating environments, where performance is measured by the correctness of the final system state rather than only the textual response. Using STAGE-Claw, this paper creates a benchmark with 40 challenging real scenario agent tasks, evaluates 11 frontier models, and analyzes their task scores, costs, tool-call reliability, and common failure patterns."
Key Contributions
- Automated generation + validation of realistic agent tasks (environment, prompts, ground truth, verification programs) from a task hint — scalable benchmark authoring.
- State-based scoring: correctness of the final system state, not the textual response — the harder, more faithful signal for personal-computing agents.
- A 40-task benchmark spanning real scenarios, run across 11 frontier models.
- Per-run analysis of task score, dollar cost, tool-call reliability, and failure patterns — the cost axis is what this KB needs.
Results — the cost-per-task numbers (why this closes the gap)
- Measured API cost per task: ≈ $0.17 (cheapest, MiniMax-tier) to ≈ $6.55 (most expensive, Claude-Opus-4.7-tier) — a ~38× spread across the 11 models on the same 40 tasks. [paper cost table, surfaced via search of the paper; the fetched abstract references "costs" without tabulating them — full-text read owed to confirm the exact per-model figures]
- Time cost per task: ≈ 3–15 minutes per model. [same provenance caveat]
- This is the first source in this KB to give a dated, absolute, per-task dollar level for real agentic workloads — the frontier's #1 sharp gap ("cost-per-task: framework in, absolute dated levels still missing") is now partly closed with a concrete 2026 range.
Methodology
State-based environments in personal-computing settings; each task carries a verification program that checks the final system state. Agents operate the environment with tool calls; scoring is pass/fail on end-state correctness plus cost/reliability instrumentation.
Limitations
- Author affiliations and the exact per-model cost table are not in the fetched abstract — the $0.17–$6.55 range is search-surfaced from the paper and needs a full-text confirmation pass.
- 40 tasks is a focused, not exhaustive, benchmark; personal-computing scenarios may not generalize to coding/research agents.
- Costs are API list-price-based (not realised enterprise prices).
Why it matters (inference-economics lens)
Pairs with Cost-of-Pass (2504.13359) (the cost÷accuracy methodology) and AI Tokenomics (2606.24616) (value≠cost theory) by supplying the missing measured absolute levels: a real 2026 agentic workload costs single-digit dollars per task at the frontier and ~$0.17 at the cheap end. Sits directly on [cost-per-task] and the [price-decline-distribution] concepts.
Source: STAGE-Claw (arXiv:2606.10394), submitted 2026-06-09. Ingested 2026-07-24 at abstract grade + search-surfaced cost table (flagged). Provenance: [preprint] for the framework/design; the specific per-task $ figures are [preprint — cost table, search-surfaced, full-text confirmation owed].