Rung 5 · Gated
Agentic AI
Agentic AI trackLINK PENDING
Last, because the harness only becomes interesting once the model underneath it is not a mystery. This is the rung the flagship argument is aimed at — that almost all of the useful behaviour lives in the system around the model.
Unscoped
Progress
Allocated
Quota
Read it
Seat
Gated
State
Syllabus length not fixed — no published lecture count to cite.
Takes the technical slot when rung 4 closes.
Gate
Follows CS336. The deep-RL family only after the CS336 alignment on-ramp is logged.
Artifacts
Consume → do → output- The harness flagship — the earned version of the argumentNot yet — nothing has landed for this slot.
- Design notes into the machine that runs this siteEarmarked for the-machine — building on the workshop; not counted until it ships there.
Feeds
What closing this sharpensThe reading
Sections of the catalogue this unit draws on- §7Agent Systems and Tool Use33 papersHarnesses
- §8Coding Agents and Software Engineering16 papersHarnesses
The papers
The rows behind the sections above- Dr. Zero: Self-Evolving Search Agents Without Training Data ↗
- Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline ↗
- Large Language Model Agents Are Not Always Faithful Self-Evolvers ↗
- Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems ↗
- AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent ↗
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments ↗
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks ↗
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook ↗
- AgentLAB: Benchmarking LLM Agents Against Long-Horizon Attacks ↗
- Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use ↗
- Evaluating Theory of Mind and Internal Beliefs in LLM-Based Multi-Agent Systems ↗
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents ↗
- Hyperagents ↗
- Multi-User Large Language Model Agents ↗
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents ↗
- Natural-Language Agent Harnesses ↗
- Drop the Hierarchy and Roles: How Self-Organizing LLM Agents Outperform Designed Structures ↗
- The Art of Building Verifiers for Computer Use Agents ↗
- Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering ↗
- Autogenesis: A Self-Evolving Agent Protocol ↗
- From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company ↗
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond ↗
- From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills ↗
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate ↗
- Recursive Multi-Agent Systems ↗
- ClawGym: A Scalable Framework for Building Effective Claw Agents ↗
- Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes ↗
- Is Grep All You Need? How Agent Harnesses Reshape Agentic Search ↗★
- AI for Auto-Research: Roadmap & User Guide ↗
- A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents ↗
- Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost ↗
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills ↗
- Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems ↗
- AI IDEs or Autonomous Agents? Measuring the Impact of Coding Agents on Software Development ↗
- CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding ↗
- AutoHarness: Improving LLM Agents by Automatically Synthesizing a Code Harness ↗
- Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? ↗★
- LongCLI-Bench: A Preliminary Benchmark and Study for Long-Horizon Agentic Programming in Command-Line Interfaces ↗
- On Data Engineering for Scaling LLM Terminal Capabilities ↗
- SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale ↗
- Qwen3-Coder-Next Technical Report ↗
- BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? ↗
- Coding Agents Are Effective Long-Context Processors ↗
- Effective Strategies for Asynchronous Software Engineering Agents ↗
- Meta-Harness: End-to-End Optimization of Model Harnesses ↗
- Scaling Coding Agents via Atomic Skills ↗
- Frontier Coding Agents Can Now Implement an AlphaZero Self-Play ML Pipeline for Connect Four That Performs Comparably to an External Solver ↗
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses ↗
- Code as Agent Harness ↗