WORLD MODELS · EVALUATION NAVIGATOR

World Model Benchmark

Explore how world models can be evaluated across prediction, simulation, planning, control, robotics, embodied intelligence, agents and long-horizon consistency.

Capability→ Task→ Metric→ Protocol→ Validation

Why World Models need multidimensional evaluation

A world model can be good at one capability and weak at another. A visually convincing model may not support planning. A short-horizon predictor may drift over long rollouts. A control model may succeed in one environment but fail to generalize.

No single score captures world-model quality. Evaluate the capability that matters for the intended system.

Benchmark Navigator

—Benchmark families
—Current matches
—Unique metrics
12Evaluation dimensions
Benchmark family Capability Useful metrics Best for Key evaluation question

Core evaluation dimensions

PredictionCan the model predict future states accurately?
Temporal consistencyDoes the environment remain coherent through time?
Spatial consistencyAre geometry and object relationships preserved?
Action conditioningDo actions produce the intended state changes?
ControllabilityCan actions reliably steer generated futures?
Planning utilityDoes the model improve decisions or policies?
Control performanceCan an agent complete tasks using the world model?
Physical plausibilityAre learned dynamics useful and physically consistent?
Long-horizon stabilityHow quickly do errors accumulate during rollouts?
GeneralizationDoes the model work outside familiar conditions?
UncertaintyCan the model represent multiple plausible futures?
EfficiencyWhat compute, latency and memory are required?

Recommended benchmark workflow

Define capability→ Choose environment→ Select metrics→ Short horizon→ Long horizon→ Planning / control→ Efficiency→ Generalization

Always report the task, dataset or environment, observation space, action space, horizon, model size, compute, training data and evaluation protocol together with the metric.