Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

2026-07-01Computation and Language

Computation and Language
AI summary

The authors studied how large language models (LLMs) that act like different 'personas' perform when answering multiple-choice questions (MCQs). They noticed that sometimes these personas are unstable, meaning their answers can change a lot depending on things like the type of question or how the task is set up. They created three ways to measure this instability and found it varies by model type, size, and question type, with math and commonsense questions causing more change. They also discovered that how the question is asked (prompt format) affects stability more than some other settings. Their work shows it's important to check how stable these personas are when using LLMs for such tasks.

Large Language ModelsPersona-driven GenerationMultiple-choice Question AnsweringInstability MetricsPrompt EngineeringModel SizeModel FamiliesHyperparametersTask AccuracyCommonsense Reasoning
Authors
César Guerra-Solano, Xiang Lorraine Li
Abstract
Persona-driven generations (PDGs) have seen prolific use in research and industry applications, where a large language model (LLM) takes on a 'persona' while completing some task. While persona expressed through free-form text (like dialogue) has substantial work investigating stability or consistency, relatively, persona expressed in non-text-heavy outputs (like in multiple-choice question answering, or MCQA) is often overlooked. We work to address this gap, seeking to understand the instability of LLM PDGs in MCQA tasks. We develop three metrics investigating the performance, outcome, and question correctness stability, evaluating three distinct dimensions. Using these metrics, we find that instability varies consistently between model families and model size, and across question domains, with math/commonsense questions leading to greater instability. We also find task prompt format introduces more prediction instability than other hyperparameters, like temperature. Finally, we find that instability is related to task accuracy, and using our instability metrics, find different experimental settings that result in different best and worst personas for tasks, despite their similarity. This reveals the importance of checking hyperparameter instability in PDGs.