PAPER2026-06-05·arXiv 2606.07017

The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective

Xiaoou Liu, Tiejin Chen, Weibo Li, Xiyang Hu, Hua Wei
COMPILED NOTES

Reframes foundation-model-agent robustness as a classical sim-to-real problem over the four MDP elements (Observation/Action/Transition/Reward); argues the LLM-agent field should import robotics' domain randomization rather than reinvent it; proposes unified vocabulary + standardized stress-test benchmarks

The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective

Abstract (verbatim)

"Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical control have mature frameworks to address this gap, the foundation model community is treating agent robustness as an entirely novel phenomenon. Our paper proposes formalizing the foundation model agent evaluation and training gap as a classical sim-to-real problem structured entirely around the four elements of a Markov Decision Process, including Observation, Action, Transition, and Reward. In this paper, we set a comprehensive research agenda that translates classical discrepancies into the foundation model domain and advocates for adopting established solutions like domain randomization. We provide concrete examples, such as a multilingual tool calling to demonstrate how severe observation space gaps lead to operationally invalid actions despite correct semantic intent. Ultimately, this agenda aims to drive a paradigm shift, yielding a unified vocabulary and standardized stress test benchmarks to foster a new generation of highly trustworthy agents for reliable real-world applications."

Key Contributions

  • MDP framing of agent robustness — restructures the sim-to-real problem using four MDP components: Observation gap, Action gap, Transition gap, Reward gap. Each classical robotics discrepancy is mapped onto its foundation-model-agent analogue.
  • "It's a solved problem" thesis — robotics and classical control already have mature frameworks; the LLM-agent community is treating robustness as novel and should import, not reinvent. Advocates adopting domain randomization.
  • Concrete worked examplemultilingual tool calling: a severe observation-space gap can yield "operationally invalid actions despite correct semantic intent" — the model understands the request but emits an unexecutable call.
  • Research agenda — a unified vocabulary plus standardized stress-test benchmarks for trustworthy agents.

Significance

This is a bridge paper, and a notable one for a robotics KB: it argues the conceptual machinery the robotics field built for sim-to-real (domain randomization, system identification, the MDP gap decomposition) is the right lens for the much larger LLM-agent field. It inverts the usual flow of ideas (robotics borrowing from ML) and validates sim-to-real as a general theory of deployment robustness, not a robotics-specific trick. It is a position/agenda paper (KDD Blue Sky), so evidence is conceptual, not empirical.

Limitations

  • Position paper — sets an agenda rather than reporting experiments; no benchmark results yet.
  • The proposed standardized stress-test benchmarks are advocated, not delivered.

Source: arXiv 2606.07017, submitted June 5, 2026. Accepted to the KDD 2026 Blue Sky Ideas Track.

RELATED · IN THE BASE
The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective | Knowledge Base | MenFem