← The reading
The bibliography
§7 · Agent Systems and Tool Use
the-machine
CAPTURED
33 On the list1 Starred0 In the Atlas33 To read
Rungs
Where this section moves a number- Agentic AIRung 5
The pick
The source author's must-read for this sectionThe papers
33 papers- Dr. Zero: Self-Evolving Search Agents Without Training Data ↗
- Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline ↗
- Large Language Model Agents Are Not Always Faithful Self-Evolvers ↗
- Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems ↗
- AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent ↗
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments ↗
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks ↗
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook ↗
- AgentLAB: Benchmarking LLM Agents Against Long-Horizon Attacks ↗
- Learning to Rewrite Tool Descriptions for Reliable LLM-Agent Tool Use ↗
- Evaluating Theory of Mind and Internal Beliefs in LLM-Based Multi-Agent Systems ↗
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents ↗
- Hyperagents ↗
- Multi-User Large Language Model Agents ↗
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents ↗
- Natural-Language Agent Harnesses ↗
- Drop the Hierarchy and Roles: How Self-Organizing LLM Agents Outperform Designed Structures ↗
- The Art of Building Verifiers for Computer Use Agents ↗
- Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering ↗
- Autogenesis: A Self-Evolving Agent Protocol ↗
- From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company ↗
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond ↗
- From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills ↗
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate ↗
- Recursive Multi-Agent Systems ↗
- ClawGym: A Scalable Framework for Building Effective Claw Agents ↗
- Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes ↗
- Is Grep All You Need? How Agent Harnesses Reshape Agentic Search ↗★
- AI for Auto-Research: Roadmap & User Guide ↗
- A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents ↗
- Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost ↗
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills ↗
- Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems ↗
What counts as read
A paper is on this list because someone worth reading put it there. That is a pointer, not a claim: it counts as read only once it has a close-read file in kb/<topic>/raw/, which is what a close-read link on a row means. There is deliberately nothing to tick off here — the study desk is the only writer of study state, and a second way to mark something done is a second version of the truth.