Evaluating AI Models' Capability to Automate Voice Phishing Attacks
2026-07-10 • Cryptography and Security
Cryptography and SecurityComputers and Society
AI summaryⓘ
The authors studied how likely people in the U.S. are to fall for phone scams that use AI-generated voices instead of real humans. They found that many people might comply with these AI scams, especially in emotional situations, with about 16.5% overall compliance across different scam types. Some AI voice models were almost as convincing as real voices, making these scams easier and cheaper to run than before. The authors point out that the main danger is how cheaply AI can automate scams, not necessarily because the AI is more persuasive than humans. This has important implications for protecting consumers and controlling AI technology.
voice phishingvishingAI voice synthesislarge language modelsscam compliancecaller persuasivenesseconomic viabilityautomationconsumer protectionmodel release policies
Authors
Fred Heiding, Claudio Mayrink Verdun, Simon Lermen, Andrew Kao, Vitor Albiero, Lauren Deason, Irina-Elena Veliche, Christine Lehane
Abstract
Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults' susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play.AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the "relative-in-distress" category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or "superhuman" persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies.