LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior
2026-07-21 • Artificial Intelligence
Artificial IntelligenceComputer Science and Game Theory
AI summaryⓘ
The authors studied how tools that detect AI-written text (LLM detectors) affect the way people use AI in their work. They found that imperfect detectors can make people use AI more, not less, and sometimes even lower the quality of the final output. Their model shows users try to avoid being detected by changing how they use AI, which can backfire. They also observed real examples where detection caused a rise and then a fall in certain word usage patterns. Overall, the authors demonstrate that detection tools can unintentionally change user behavior and output quality in surprising ways.
Large Language ModelsLLM detectionUser behaviorOutput qualityHeuristicsContent post-processingModeling incentivesArXiv abstractsWord frequency analysisDownstream metrics
Authors
Meena Jagadeesan, Tatsunori Hashimoto, Jon Kleinberg
Abstract
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns. Detectors operate as an intervention that steers not only the detected attribute itself, but also downstream metrics such as LLM usage and output quality. In this work, we demonstrate how imperfect LLM detectors lead to counterintuitive impacts on these downstream metrics, by distorting how users are incentivized to use LLMs in their workflow. We develop a stylized model which captures how users strategically choose how much to use the LLM and how to post-process content to reduce the detected attribute. Using this model, we show that LLM detection can counterintuitively lead humans to increase their LLM usage. Moreover, even when reducing the detected attribute improves output quality, we find that introducing an LLM detector can lead users to produce lower quality outputs. In contrast, we show that detectors result in a clean "rise-then-fall" pattern for the detected attribute, which we empirically reproduce for word frequencies on arXiv abstracts. Altogether, our work illustrates how LLM detection can distort LLM usage and output quality, uncovering failure modes when LLM detectors operate as an intervention on these downstream metrics.