Explore how world models can be evaluated across prediction, simulation, planning, control, robotics, embodied intelligence, agents and long-horizon consistency.
A world model can be good at one capability and weak at another. A visually convincing model may not support planning. A short-horizon predictor may drift over long rollouts. A control model may succeed in one environment but fail to generalize.
| Benchmark family | Capability | Useful metrics | Best for | Key evaluation question |
|---|
Always report the task, dataset or environment, observation space, action space, horizon, model size, compute, training data and evaluation protocol together with the metric.