Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation
2026-08-03 • Computers and Society
Computers and SocietyArtificial IntelligenceComputation and Language
AI summaryⓘ
The authors discuss how evaluating fairness in AI models is not just about looking at what the model outputs but also about deciding what those outputs should be compared to. They point out that current fairness checks assume we know the sensitive attributes on the input, but for generative models, it's unclear what the target demographic distribution should be. They create a framework to decide these target distributions by considering factors like the context and why a certain demographic mix is chosen. Their findings show that different ways of setting these targets lead to significant differences in measuring fairness, meaning choosing the target is an important part of the fairness evaluation itself.
fairness evaluationgenerative modelssensitive attributesdemographic distributiontarget distributiongroup fairnessAP-BenchJensen-Shannon divergenceworkforce-composition fidelitygeographic prior
Authors
Zeshen Zheng, Yujia He, Qianmian Lin, Xiangyue Huang, Wenqing Chen
Abstract
Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. When a model generates "a CEO in the United States," the prompt leaves demographic realization to the model. Existing group fairness definitions assume that sensitive attributes are given on the input side. Generative audits instead examine output-side demographic composition, yet the targets they compare it against are typically supplied rather than justified. The upstream question is what the target distribution should be. We formalize this missing-target problem for demographic-value-unspecified generation and decompose target construction into four commitments: the evaluative object, prior admissibility, allocation, and operationalization. In this framework, we admit the geographic prior under a geographic-membership interpretation for the declared public-world use. The occupational prior, under an incumbency interpretation, requires an independently defended objective such as workforce-composition fidelity. Instantiating this construction in AP-Bench, we find substantial distribution divergence from geography-derived targets, ranging from 0.508 to 0.606 on a 0-to-1 scale. Replacing each geography-derived target with an equal-category comparator, while holding generations and measurement fixed, produces model-specific mean absolute cell-level $\mathrm{JSD}_2$ changes ranging from 0.279 to 0.355. Target construction is therefore not a preliminary to fairness evaluation but a component of it. What we supply is not a universal target, but a framework that makes explicit the justification required before a distribution can serve as a fairness standard.