Principia: Relational Physics Tests for Video Models
2026-09-03 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors created a new way to test if video models understand physical rules by checking if pairs of objects follow the same physical laws, no matter the camera setup. They made a benchmark called Principia that looks at eight different physics phenomena using real videos and measures how well models predict consistent behavior. They found that current video generators and vision-language models struggle to accurately reflect these physical relationships. Even the best models performed much worse on this new test than on previous benchmarks.
physical reasoningvideo modelsNewtonian physicscalibration-independentrelational consistencybenchmarkmotion dynamicsphysics phenomenavision-language modelsphysical violation
Authors
Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan, Anand Bhattad
Abstract
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.