Surprisal Theory is Tautological (without Rational Grounding)

2026-07-23Computation and Language

Computation and Language
AI summary

The authors explain that surprisal theory, which links how hard it is to process a word to how surprising it is under a language model, is actually a tautology without extra rules. This means you can always find some language model that fits any pattern of difficulty, so the theory can't be proven wrong unless the model is restricted. Previous work assumed the language model matched the data humans experience, but newer studies show better data models don't always predict human difficulty better. The authors suggest that to make meaningful predictions, the language model must come from an understanding of how people process language, not just from data patterns.

surprisal theorylanguage modelpsycholinguisticsprocessing difficultytautologyaffine functioncorpusmodel evaluationrationalist approachcognitive constraints
Authors
Ryan Cotterell
Abstract
Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose surprisal is an affine function of it under mild technical conditions. Therefore, because any pattern of difficulty is consistent with some language model, without an additional constraint on the language model, surprisal theory makes no falsifiable predictions. The tautology was long obscured by an assumption implicit in two decades of psycholinguistic work---that the relevant language model is the distribution that generated the training corpus, so that improving corpus fit improves predictions of human behavior. Recent empirical work has undermined this assumption, demonstrating that better corpus models can be worse predictors of processing difficulty. I conclude that breaking the tautology requires a rationalist intervention, i.e., the relevant language model must be derived from a non-empirically motivated model of the comprehender, which could be based on, for instance, memory constraints or processing goals, and that, thus, does not depend on the behavioral data surprisal theory is meant to explain.