Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers
2026-08-14 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors studied how developers gather people's moral opinions to train AI systems that make ethical decisions. They found that three key choices—what features to ask about, who gets to vote, and how questions are worded—can all influence the moral preferences collected. These choices vary by context, political beliefs, and question framing, which means that aggregated votes alone might not produce fair or clear AI ethics. The authors suggest that each step should be carefully reviewed and openly shared to improve transparency.
moral preference elicitationfeature scopingvoter samplingquestion framingAI ethicspolitical ideologymoral foundations theoryAI alignmentaggregation bias
Authors
Taenyun Kim, Edyta Bogucka, Daniele Quercia
Abstract
As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices in the moral AI elicitation pipeline: feature scoping, voter sampling, and question framing. In other words, they decide which features go to a vote, which voters to include, and how to present the question. These choices are often opaque, undocumented, and treated as technical details rather than normative ones. We examine each of these choices within a common empirical study and show that each can shape the preferences produced by moral AI elicitation. Across two phases (N = 809) in three deployment contexts (i.e., AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of the deceased), we examine the three main stages of the moral AI elicitation pipeline. First, morally relevant features shift across contexts. This suggests that feature schemas should not be assumed to transfer across deployment domains. Second, preferences differ by political ideology for roughly one-third of features, with some differences reversing direction. The ideological composition of the voter pool can therefore affect the resulting aggregated preference profile. Third, the wording of the elicitation question can narrow or widen ideological gaps by up to a full scale point. The framing conditions also change how moral foundations are associated with participants' judgments. Taken together, these findings suggest that voting-based alignment cannot deliver fair or transparent AI by aggregation alone; at minimum, each stage of the moral AI elicitation pipeline should be audited and disclosed.