Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse

2026-07-12Computation and Language

Computation and Language
AI summary

The authors studied how different automated methods of processing interview videos can affect scientific results. They looked at two parts: the way speakers are identified and the way language features are measured. They found that differences in speaker identification mostly affect cases with few video samples, while measurement methods can cause bigger and more systematic disagreements. Overall, general emotion proportions stay stable, but subtle relationship patterns can be inconsistent depending on the method used. The authors suggest researchers carefully check that their findings hold up across different processing and measurement methods.

speaker diarizationASR transcript cleaningsentence segmentationaffective valenceepistemic modalityLLM annotationkeyword lexiconpipeline sensitivitymeasurement instrumentYouTube interviews
Authors
Bo Chen
Abstract
Computational social science increasingly relies on automated preprocessing pipelines -- speaker diarization, ASR transcript cleaning, sentence segmentation -- to convert raw media into analyzable text. When these pipelines produce different outputs from the same input, two distinct sources of instability can arise: the preprocessing pipeline itself (diarization method, segmentation rules) and the downstream measurement instrument (LLM annotation vs.\ keyword lexicon). Using 256 YouTube interviews across 41 public figures from five domains, we compare two speaker-diarization pipelines and two measurement methods, all targeting the coupling between affective valence and epistemic modality. We find that (1) preprocessing pipeline sensitivity is concentrated in speakers with limited video samples (N $\leq 5$); for the four best-sampled speakers (N $\geq 16$), the mean absolute pipeline-induced change in $r(\text{neg}, \text{emph})$ is only $0.13$; (2) cross-method disagreement is larger and more systematic -- the LLM and keyword-lexicon methods assign opposite coupling directions to several well-sampled speakers, even within the same preprocessing pipeline; and (3) aggregate valence proportions are highly stable ($|Δp(\text{neg})| < 6$pp) regardless of pipeline or method, masking both sources of instability. The contribution is a diagnostic framework that separates pipeline effects from measurement effects: researchers studying cross-dimensional relationships in interview data should verify that their conclusions are robust to both sources of variation, with particular attention to measurement method choice.