DETECT-3B-Omni is Agnostic of Content and Demographics
2026-07-03 • Sound
SoundArtificial Intelligence
AI summaryⓘ
The authors studied a deepfake audio detector called DETECT-3B-Omni to see if it unfairly focuses on what is being said or who is speaking. They tested over 10,000 audio clips from many speakers and different AI voice systems. Their results show the detector works equally well no matter the content, speaker's gender, age, or region. This means the detector bases its decisions on audio features, not on the speech meaning or speaker identity.
deepfake audioAI voice cloningaudio detectionsemantic independenceResemble AIGDPR complianceequivalence testingspeaker characteristicsdetection accuracy
Authors
Nicolas M. Müller, Aditya Tirumala Bukkapatnam, Dominik Schnieders, Zohaib Ahmed
Abstract
A trustworthy and GDPR-compliant deepfake audio detector must base its decisions on acoustic artifacts, not on what is being said or who is speaking. We present a large-scale study of semantic independence for Resemble AI's detector, DETECT-3B-Omni. Using 10,240 audio samples from diverse US English speakers across 30 states, generated through 8 different AI voice-cloning systems, we test whether detection accuracy depends on spoken content (benign versus malicious), speaker gender, speaker age, or speaker region. Using equivalence testing, our results show that the accuracy difference between any two of these groups is at most 2 percentage points, at 99% confidence. The detector therefore identifies AI-generated audio with equivalent accuracy regardless of what the audio says or who the speaker is.