Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models

2026-07-12Computer Vision and Pattern Recognition

Computer Vision and Pattern RecognitionRobotics
AI summary

The authors explore whether controlling AI-generated driving videos requires retraining the underlying model. They show that by using a special method called rectified-flow, they can guide the planned driving path during video generation without changing the original model. This is done by adding rules about driving behavior at the time the video is created. However, they also find that the generated video does not fully follow the intended path yet because the parts of the model that combine video and trajectory information need better connection for full control.

driving world modelsvideo diffusionrectified-flowego trajectoryenergy guidanceOpen-Sora 2.0MM-DiT backboneself-attentioncross-stream couplingsampling time control
Authors
Xiyan Su, Frank Diermeyer, Markus Lienkamp
Abstract
Driving world models built on large video-diffusion backbones generate realistic scenes but are hard to control: enforcing a traffic norm typically means retraining the backbone or conditioning it on hand-built layouts. We ask whether controllability requires training at all. Our experiment shows that a rectified-flow driving world model, which jointly generates future video and a planned ego trajectory, can have its planned trajectory steered entirely at sampling time by differentiable energy functions that encode driving norms, without knowledge-specific retraining of the diffusion backbone. Concretely, we demonstrate that a world model built on Open-Sora 2.0 MM-DiT backbone can be steered to brake at a counterfactual target by injecting energy guidance at sampling time. However, we find that the generated video does not yet follow the steered trajectory through the backbone's joint self-attention and identify the cross-stream coupling as a crucial requirement for end-to-end-controllable rollouts.