Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations
2026-08-04 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors studied how different ways people input text, like typing on a keyboard or speaking, affect language model accuracy. They created a tool called HIVE to simulate mistakes from typing and voice transcription. They found that errors from voice transcriptions cause more problems than typing mistakes, mainly because voice errors change important parts of the input. Typing errors have less impact unless many words are completely destroyed. Also, improving models by simple retraining or extra thinking helps mostly with typing errors, not voice errors.
language modelsvoice transcriptionkeyboard inputinput perturbationsorthographic noisedisfluencyinstruction tuningtokenizationmodel robustnesslightweight adaptation
Authors
Zizhao Hu, Nathan Elijah Segura, Mohammad Rostami, Jesse Thomason
Abstract
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools. How do they impact an LLM's performance? In this paper we present HIVE (Human Input-Variation Engine), a suite of voice transcription perturbations and QWERTY keyboard perturbations. We use HIVE to evaluate how robust models are to these perturbations. We present seven findings. (i) Voice transcription perturbations lower accuracy across every instruction-tuned model we test, and it is the structure of the transcription rather than its fillers that carries the cost. (ii) QWERTY keyboard perturbations cost less, and a model absorbs a lot of them before accuracy falls away. (iii) Both trace back to one cause, how many of the question's tokens survive the perturbation: destroying a token is what hurts, while adding new ones alongside it costs little. (iv) The gap between the two channels appears only where the answer must be constructed or deduced; on multiple choice there is none. (v) The harm does not solely come from test-set contamination. (vi) It cannot be trained away with lightweight adaptation. (vii) A thinking budget recovers the keyboard channel almost entirely but leaves the spoken registers untouched, and compressed speech is worse with it.