DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding

2026-07-01Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors tackle zero-shot video temporal grounding, where a system finds events in videos based on language queries without special training. They note that simple matching methods struggle with complicated queries needing an understanding of event order and cause. Their method, DART, uses a smart system to decide if a query is easy or hard, then applies either a quick prediction or a detailed step-by-step reasoning process. This approach leads to better accuracy while analyzing fewer video frames. They tested it on datasets and showed improvements over previous methods.

zero-shot learningvideo temporal groundingvision-language modelsdeterminantal point processspectral entropytemporal event localizationmulti-stage reasoningCharades-STAActivityNet Captions
Authors
Zhengbo Zhang, Mark He Huang, Zhigang Tu, Ming-Hsuan Yang
Abstract
Zero-shot video temporal grounding (VTG) localizes events in untrimmed videos from natural language queries without task-specific training. Existing methods rely on frame-query feature matching, which suffices for simple events but struggles with complex multi-stage queries that require understanding temporal ordering and causal structure -- a disparity we call the reasoning gap. We propose DART (Difficulty-Adaptive Routing for Temporal Grounding), which bridges this gap by coupling difficulty-aware routing with structured reasoning in large vision-language models. A query-conditioned Determinantal Point Process (DPP) serves a dual role: selecting diverse, query-relevant keyframes as temporal evidence, and providing spectral entropy as a difficulty indicator. Simple queries are routed to a Fast path for direct prediction, while complex queries follow a Slow path with Temporal Markup Prompting, which decomposes localization into global event analysis, per-frame temporal role annotation, and boundary extraction. On Charades-STA and ActivityNet Captions, DART achieves state-of-the-art zero-shot performance across both identically distributed and multiple out-of-distribution settings, improving mIoU by up to 3.5 points over the strongest baseline while using over 7 times fewer frames. The project homepage is available at https://dart-vtg.github.io/.