Silicon Sampling via Cross-Survey Transfer
2026-07-03 • Artificial Intelligence
Artificial IntelligenceComputation and LanguageComputers and SocietyMultiagent Systems
AI summaryⓘ
The authors studied how large language models (LLMs) can imitate human answers in surveys by predicting responses to new questions based on previous answers. They tested this method using data from a Taiwanese election survey and found that LLMs can predict unseen answers fairly well, close to traditional machine learning methods. They also discovered that some types of questions, like those about political party attitudes, are easier to predict than others, such as views on sovereignty. Additionally, they examined limitations of LLMs and found these issues are more complex and occur in other models too. Overall, their work helps show both the potential and limits of using LLMs for survey research.
Large Language ModelsSilicon SamplingCross-Survey TransferSurvey ResearchZero-Shot LearningRandom ForestVariance CollapseSafety AlignmentTaiwan Election and Democratization Study
Authors
Chan-Tung Ku, Chan Hsu, Pei-Cing Huang, Frank Cheng-shan Liu, I-Ling Cheng, Yihuang Kang
Abstract
Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.