Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

2026-07-28Computation and Language

Computation and Language
AI summary

The authors studied whether large language models (LLMs) imitate the grammar patterns of people they are talking to, similar to how humans do in conversations. They tested various Llama and Gemma models by swapping in model-generated replies within real human dialogues and checked how much these models reused certain grammar rules right after a human turn. They found that all models did show this kind of grammatical imitation more than random chance, especially for rarer grammar patterns. Instruction-tuned models imitated previous turns more than before tuning but also showed some increase in overlap with unrelated turns, suggesting tuning changes mimicry in complex ways. Additionally, models generally matched the previous turn better in word choice and meaning than humans did, with tuning boosting semantic similarity consistently.

Syntactic convergenceLarge language modelsContext-free grammarInstruction tuningLexical similaritySemantic similarityPretrained modelsHuman dialogueOpen-weight modelsTurn-adjacent reuse
Authors
Zandi Eberstadt
Abstract
Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented feature of human dialogue widely considered to operate below conscious awareness. Whether large language models exhibit analogous syntactic convergence toward human users relative to human baselines and across a broad range of syntactic constructions remains an open question. Using substitution-paradigm data in which model generations replace one speaker's turns in pre-existing human dialogues, this study measures turn-adjacent reuse of context-free grammar (CFG) rules across sixteen open-weight Llama and Gemma models (1B-70B, pretrained and instruction-tuned) at 1,901 matched positions per model. Every model showed greater CFG-rule overlap with the preceding human turn than with a sampled unrelated human prime, and in every model this actual-versus-random difference was larger for lower-frequency rules. Each instruction-tuned model also showed greater natural-output overlap with the actual prime than the human response it replaced, and all eight matched architecture pairs exhibited greater actual-prime overlap after instruction tuning. However, relative to pretrained variants, instruction-tuned outputs overlapped more with unrelated primes, showed a smaller actual-versus-random increment, and had lower conditional rule-reuse odds once target rule-set size was held constant. In exploratory analyses, each model exhibited greater mean lexical and semantic similarity to the preceding turn than the matched human responses did. Instruction-tuned models additionally produced responses with greater mean semantic similarity than their pretrained counterparts in all eight architecture pairs, whereas the lexical similarity results were more heterogeneous.