Mechanism Design for Alignment and Control

2026-09-01Artificial Intelligence

Artificial IntelligenceComputer Science and Game Theory
AI summary

The authors create a way to design rules (mechanisms) for AI agents when we don't know what they want or what they can do. Their approach makes sure the agents tell the truth and follow instructions, even when agents can hide their abilities but can't fake them. They explain when and how we can predict agent behaviors using math tools like nested cyclical monotonicity. Then, they show examples like an agent pretending to be less skilled, balancing between being understandable and aligned with goals, using peer reviews to keep agents honest, encouraging competition, and improving oversight as more agents get involved.

mechanism designAI agentsalignmentcapabilitiesrevelation principlenested cyclical monotonicitysandbaggingpeer scoringreward shapingoversight
Authors
Dirk Bergemann, Andrew Koh, Stephen Morris
Abstract
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.