01THE SYSTEM AROUND THE MODEL· SURGING

Harnesses & Agent Systems

The system around the model — the rung that turns a token price into a TASK price. In scope: scaffolding and orchestration, tool use, context and memory management, retries and recovery, agent evaluation and benchmarks, failure modes at step three, agent security and the exploitation surface, and cost-per-completed-task. The load-bearing finding this topic already holds: for long-horizon agentic work the harness is often a stronger performance determinant than the model. Connor's seat here is `ran it` — he runs desks, subagents and workflows daily, which is the standing the teaching doctrine requires for a teaching-level claim.

15SOURCES
15CONCEPTS
0ENTITIES
SOURCE MIX
14 P1 R0 A0 N
ACTIVITY · 20W
OrchestrationTool useAgent memoryFailure modes
PAPER
2026-01-18

Agentic Reasoning for Large Language Models

Tianxin Wei et al. · Multiple institutions

Three-layer framework for agentic reasoning: foundational, self-evolving, multi-agent

PAPER
2026-03-06

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review

Mohamed Amine Ferrag et al. · Multiple institutions

Unified taxonomy of ~60 benchmarks, agent framework comparison, collaboration protocols survey

PAPER
2026-04-01

Agentic Tool Use in Large Language Models

Hu Jinchao et al. · Harbin Institute of Technology Shenzhen, TikTok Inc

Unified evolutionary framework for LLM tool use: prompting, supervised, RL paradigms

PAPER
2026-02-07

Agentic AI Security & Autonomous Red-Teaming

Ashok Kumar Kanagala · Independent Researcher, Boston, MA

Red-teaming framework for agentic AI: permission escalation, hallucination, orchestration flaws, memory manipulation, supply chain

PAPER
2026-03-08

Memory for Autonomous LLM Agents: Mechanisms, Evaluation, and Emerging Frontiers

Pengfei Du · Not specified

Write-manage-read taxonomy, 5 mechanism families, three-dimensional taxonomy for agent memory

PAPER
2025-02-17

A-MEM: Agentic Memory for LLM Agents

Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, Yongfeng Zhang · Multiple institutions

Agentic memory with Zettelkasten-inspired note construction, dynamic linking, memory evolution

PAPER
2026-04-06

Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities

Charafeddine Mouzouni · OPIT – Open Institute of Technology; Cohorte AI, Paris, France

Large-scale systematic taxonomy of LLM agent exploitation triggers across 12 attack dimensions, identifying goal reframing as the sole reliable trigger while ruling out nine others, with GPT-4.1 achieving complete immunity across 1,850 trials.

PAPER
2026-04-06

Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw

Zijun Wang, Haoqin Tu, Letian Zhang, Hardy Chen, Juncheng Wu, Xiangyan Liu, Zhenlong Yuan, Tianyu Pang, Michael Qizhe Shieh, Fengze Liu, Zeyu Zheng, Huaxiu Yao, Yuyin Zhou, Cihang Xie · UC Santa Cruz, National University of Singapore, Tencent, ByteDance, UC Berkeley, UNC-Chapel Hill

First real-world safety evaluation of a deployed personal AI agent (OpenClaw), introducing the CIK taxonomy and showing that poisoning any single dimension raises attack success rate from 24.6% to 64–74%.

PAPER
2026-04-11

The Amazing Agent Race: Strong Tool Users, Weak Navigators

Zae Myung Kim, Dongseok Lee, Jaehyung Kim, Vipul Raheja, Dongyeop Kang · University of Minnesota Twin Cities, Yonsei University, Google DeepMind

DAG-structured benchmark of 1,400 Wikipedia navigation tasks revealing that current best agents achieve only 37.2% accuracy with navigation errors dominating (27–52% of failures), exposing compositional reasoning as the primary frontier bottleneck.

PAPER
2025-12-16

Model-First Reasoning LLM Agents: Reducing Hallucinations through Explicit Problem Modeling

Annu Rana, Gaurav Kumar

Two-phase reasoning: LLMs construct explicit problem models before generating solutions. Reduces constraint violations vs CoT and ReAct across five planning domains.

PAPER
2026-07-14

Rethinking the Evaluation of Harness Evolution for Agents

Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, Teng Xiao · Multiple institutions (inferred UW/AI2 affiliations, not stated verbatim in abstract)

Automatic harness evolution for LLM agents does not consistently outperform simple test-time-scaling baselines under matched feedback/inference budget, and generalizes poorly to held-out tasks (Terminal-Bench 2.1, GPT-5.4 + Claude Opus 4.6) — a methodological check on 'scaffold gains' claims in the agentic-harness literature.

PAPER
2026-05-07

Stop Comparing LLM Agents Without Disclosing the Harness

Yunbei Zhang, Janet Wang, Yingqiang Ge, Weijie Xu, Jihun Hamm, Chandan K. Reddy · Not stated verbatim on abstract page

Formalizes the Binding Constraint Thesis — for long-horizon agentic tasks across comparable-capability models the execution harness is often a stronger performance determinant than the model — and measures harness-induced variance at 7.80x model-induced variance (18.48 vs 2.37 pp^2, 6/9 ranking reversals) in a 3-model x 3-harness SWE-bench Verified experiment; proposes an ETCSOVG Harness Card disclosure standard + variance-decomposition protocol. Directly on the MenFem 'edge is the harness' thesis.

PAPER
2026-05-27

Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Yilun Yao, Xinyu Tan, Chao-Hsuan Liu, Yaoming Li, Zhengyang Wang, Wenhan Yu, Zhewen Tan, Yuxuan Tian, Guangxiang Zhao, Lin Sun, Xiangzheng Zhang, Tong Yang · Not stated verbatim on abstract page

Diagnostic benchmark (106 sandboxed offline tasks, 5,194 execution trajectories) that isolates configuration-level harness effects from model capability by fixing task/budget/eval and varying only the harness across model backends; finds substantial variation in completion, quality, efficiency, and failure behavior, and names execution-alignment decoupling as the dominant failure class. Empirical complement to Stop-Comparing.

PAPER
2026-07-05

Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents

Haiwen Yi, Xinyuan Song · Not stated verbatim on abstract page

Shows the harness changes an agent's multi-step beliefs (progress, risk, recoverability, constraints, failure mode, uncertainty, future success, repair cost, next action) even with task/env/model fixed; introduces a belief-rollout diagnostic, a cross-harness belief-divergence metric split into arrival (interface) + growth (horizon) terms, and BIWM (no-training trajectory alignment). Terminal success often preserved while decision-driving beliefs diverge. Freshest of the July harness cluster.

REPORT
2025-05-14

AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms

Google DeepMind · Google DeepMind

Gemini-powered evolutionary coding agent; 0.7% Google compute recovery; math breakthroughs

Harnesses & Agent Systems | Knowledge Base | MenFem