Robot Wiki

World Models

Learned simulators and action-conditioned video prediction, with the JEPA counterargument.

This domain overview lists 8 articles, ordered as what the term denotes, then the latent-dynamics and generative-video families that each claim it, then the argument that generation is the wrong objective entirely.

  1. An editorial comparison of six world-model example groups: what they predict, in what representation, and for what purpose. The survey-defined functional criterion is decision-relevant prediction, not visual plausibility alone.

  2. Dreamer, TD-MPC2, and DayDreamer: compact learned dynamics for imagination-based control.

  3. Cosmos, Genie, and GR-2: action-conditioned video prediction and the conditioning-strength problem.

  4. V-JEPA 2 and LeCun's case that prediction in representation space beats pixel generation.

  5. Generated content inside real physics engines beats generated dynamics: RoboGen, Holodeck, RoboCasa.

  6. Learn dynamics, plan through them, and improve from imagined rollouts: the common structure behind Dreamer, TD-MPC2 and robotic world models.

  7. Visual fidelity is not enough: evaluate action sensitivity, rollout consistency, task progress, policy ranking and real-world agreement.

  8. Learned dynamics, explicit physics and hybrid simulation compared by controllability, coverage, speed, debugging and downstream policy value.