World Models
Learned simulators and action-conditioned video prediction, with the JEPA counterargument.
This domain overview lists 8 articles, ordered as what the term denotes, then the latent-dynamics and generative-video families that each claim it, then the argument that generation is the wrong objective entirely.
An editorial comparison of six world-model example groups: what they predict, in what representation, and for what purpose. The survey-defined functional criterion is decision-relevant prediction, not visual plausibility alone.
Dreamer, TD-MPC2, and DayDreamer: compact learned dynamics for imagination-based control.
Cosmos, Genie, and GR-2: action-conditioned video prediction and the conditioning-strength problem.
V-JEPA 2 and LeCun's case that prediction in representation space beats pixel generation.
Generated content inside real physics engines beats generated dynamics: RoboGen, Holodeck, RoboCasa.
Learn dynamics, plan through them, and improve from imagined rollouts: the common structure behind Dreamer, TD-MPC2 and robotic world models.
Visual fidelity is not enough: evaluate action sensitivity, rollout consistency, task progress, policy ranking and real-world agreement.
Learned dynamics, explicit physics and hybrid simulation compared by controllability, coverage, speed, debugging and downstream policy value.