Training Object Permanence in World Models
Object permanence and solidity are hallmarks of human cognitive priors. Video generation models are a paradigmatic class of current world models and have begun to show emergent reasoning abilities, making them candidates for building human-like physical intelligence. The paper asks whether video models have emergent object permanence and, if not, whether they can be trained with a core-cognition-inspired dataset.
The authors introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive-science-inspired tasks divided into six cognitive categories. Blender generators randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. They release a 1.5M-sample training corpus and a 300-question exam.
On the exam, they evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation models, including PWM-WROP, their 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. They release the data, exam, model answers, scores, weights, and PWM, their native-PyTorch training stack on AWS Trainium2.