Calibrating the Instrument: Controllability of an LLM-Driven Synthetic Population
2026-07-01 • Multiagent Systems
Multiagent Systems
AI summaryⓘ
The authors study a synthetic population tool that simulates how groups respond to different messages from institutions. They test if the tool behaves in a consistent and predictable way when given messages ranging from very positive to very negative. Their experiment shows that the tool reliably reflects the hidden group structures they built in, meaning it controls responses well before comparing to real people. They also found that one message thought to be slightly positive was actually seen as negative by the tool, which helped them improve the message design. The authors share their detailed data and analyses to help understand how this synthetic instrument works internally.
Generative Synthetic Populationspopulation synthesisagent-based modellingLLM agentscontrollabilityinternal validityinstitutional communicationslatent structureinstrument calibrationnoise floor
Authors
Mirko Degli Esposti
Abstract
Generative Synthetic Populations (GSP) -- the convergence of population synthesis, agent-based modelling, and LLM agents -- are attracting growing interest for urban simulation and institutional communication research. Before any GSP instrument is used on a real population, a more basic question must be answered: does it respond to stimuli of known valence in an ordered, replicable, group-structured way? We call this controllability. We ask not whether a synthetic population tracks humans, but whether it tracks itself: whether the latent structure we impose on it is recovered in its own responses. This internal-validity question is logically prior to any claim about external validity, just as characterising an instrument's response function must precede using it to test a theory. We report SIVE (Synthetic Instrument Validation Experiment): a fictional municipality (Montelago) with 120 synthetic personas of known latent structure, exposed to seven conditions spanning strongly positive to strongly negative institutional communications about a water network. Seven pre-registered criteria, evaluated across a temperature sweep, jointly assess fidelity, stability, noise floor, specificity, sensitivity, and ordering. All seven pass at every temperature. A central finding turns a calibration failure into a diagnostic success: a message designed as "weakly positive" was identified by the instrument as functionally negative, traced to unresolved problems, uncertainty, and institutional passivity in its text; a redesigned version restored the expected ordering and interacts with agents' latent trust in unanticipated ways. A noise sub-experiment shows the instrument's intrinsic noise is roughly half the cross-agent estimate and stable across temperatures. Individual trajectories reveal coherent micro-dynamics that summary statistics obscure. Full data are available via an interactive explorer.