Detecting AI-Generated Video: A Vision-Language Dual-View Survey
2026-07-12 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionComputation and Language
AI summaryⓘ
The authors explain that as AI-generated videos become more realistic, old methods that look for simple flaws no longer work well. They suggest focusing on checking if the video's content matches real-world facts, calling this Factual Fidelity Verification. The authors organize existing detection techniques into a four-level system that uses both visual and language information, moving from simple checks to advanced reasoning about the video’s meaning. They review many studies, discuss current challenges, and suggest ways to improve reliable and understandable detection methods.
AI-generated videosDeepfake detectionFactual Fidelity VerificationVision-language modelsSpatiotemporal consistencyCross-modal reasoningArtifact analysisSemantic verificationEvaluation benchmarks
Authors
Dylan Xinming Hou, Juntian Zhang, Xu Gu, Yichen Wu, Nils Lukas, Gus Xia, Xiuying Chen, Yuhan Liu
Abstract
The evolving realism of AI-generated Videos (AIGC-V) is rapidly rendering traditional artifact-centric detection insufficient, necessitating a paradigm shift from low-level inspection to high-level semantic verification. This paper presents a comprehensive survey of AIGC-V detection, reframing the task as Factual Fidelity Verification, which asks whether the events, entities, and physical processes depicted in a video are consistent with real-world facts. To systematize this rapidly evolving field, we propose a Vision-Language Dual-View taxonomy that organizes existing methods into a hierarchical, four-layer landscape, spanning intrinsic cue analysis, spatiotemporal consistency modeling, cross-modal consistency reasoning, and language-guided world-level reasoning. This dual-view framing highlights a fundamental transition from artifact matching in traditional deepfake detection to evidence-based semantic verification enabled by vision-language models and agentic reasoning pipelines. Based on a systematic review of 221 works, we synthesize AIGC-V generation paradigms, survey the landscape of detection methods, and review evaluation metrics and benchmarks in line with proposed views. Finally, we discuss current challenges and identify promising directions toward robust, explainable, and trustworthy detection.