
In scope: scaffolding and orchestration, tool use, context and memory management, retries and recovery, agent evaluation and benchmarks, failure modes at step three, agent security and the exploitation surface, and cost per completed task. The load-bearing finding this rung already holds: for long-horizon agentic work the harness is often a stronger performance determinant than the model.
01Harness-Induced Belief Divergenceactive02Agent Harness Evolutionactive03Agent Memory Architecturesactive04Agentic Reasoningactive05The Binding Constraint Thesisactive06Evolutionary Algorithm Discoveryactive07Measuring Harness Effectsactive08LLM Tool Useactive09Multi-Agent Systemsactive10Reinforcement Learning for Agentsactive11Tool-Chain Navigationactive