ArXiv

A game theory for foundation models shows new paths to rational cooperation through similarity inference

Authors
Alexander Meulemans, Maciej Wołczyk, Marissa A. Weis...
Categories
cs.AI
arXiv
https://arxiv.org/abs/2608.03958v1
PDF
https://arxiv.org/pdf/2608.03958v1

Brief

The paper studies interactions of foundation-model agents in social dilemmas and finds that optimal planning leads to robust cooperation, opposing Nash-style expectations of defection. To explain this, the authors propose the embedded Bayesian agent and the embedded equilibrium: agents model themselves as part of the environment, infer behavioral similarity, and use their own deliberation as evidence of partners' actions, yielding new paths to rational cooperation.

Why it matters

On 2026-08-04 (arXiv:2608.03958v1), Meulemans et al. report that foundation-model agents using optimal planning in stylized social dilemmas consistently converge to stable cooperation, contradicting classical game-theory predictions of mutual defection (paper: 75 pages, 11 figures).

Key details

  • The authors introduce the 'embedded Bayesian agent' and a new solution concept, the 'embedded equilibrium,' where agents maintain epistemic uncertainty about their own algorithms and perform similarity inference—treating their own deliberation as evidence that similar partners will act similarly.
Source evidence

Abstract

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce theembedded Bayesian agent,' a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,' a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.

Comment: 75 pages, 11 figures