One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
2026-09-03 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionArtificial Intelligence
AI summaryⓘ
The authors present EditVid, a new tool for editing videos that works without needing extra training. It uses special techniques to keep the video looking consistent, preserve important details, and apply changes only where needed. EditVid can handle many types of edits guided by instructions or example videos, like changing styles or replacing subjects. Tests show it performs better than other methods that also require no training, and people generally prefer its results.
video editinginstruction-guided editingreference-guided editingsparse causal memorytoken injectionlatent blendingstyle transfersubject replacementFiVE benchmarkIVEBench
Authors
Adheesh Sunil Juvekar, Onkar Kishor Susladkar, Kiet A. Nguyen, Muntasir Wahed, Nabeel Bashir, Xiaona Zhou, Tianjiao Yu, Vedant Shah, Ismini Lourentzou
Abstract
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token injection for long-range identity preservation, and soft latent blending for edit locality. The same framework supports instruction-guided and reference-guided edits, including style transfer, attribute modification, object insertion, part-level editing, and subject replacement. On FiVE, EditVid achieves 78.16 FiVE-Acc, compared with 58.95 for the strongest evaluated training-free baseline, while obtaining competitive results on IVEBench. A user study further shows a 51.8\% overall preference for EditVid over 7 competing methods.