ArXiv

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Authors
Yuxue Yang, Shuyao Shang, Jiahe Wang...
Categories
cs.CV
arXiv
https://arxiv.org/abs/2608.02603v1
PDF
https://arxiv.org/pdf/2608.02603v1

Brief

WorldExam introduces a hierarchical benchmark (1,474 cases, eight tasks) that extends evaluation of controllable video/world models beyond visual fidelity to inherent reactivity—how scenes should plausibly react to state and goals. The paper (Yang et al., arXiv 2026-08-03) evaluates 20 camera-, action-, and language-driven models and reveals distinct failure modes and that high visual quality or instruction following does not guarantee scene reactivity.

Why it matters

WorldExam is a hierarchical diagnostic benchmark with 1,474 cases across eight tasks and four evaluation levels—Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity—designed to evaluate camera-, action-, and language-driven controllable video models.

Key details

  • The authors evaluated 20 representative models (paper published 2026-08-03) and found a capability split: camera-driven models excel at camera control but lack dynamic interaction support; action-driven models control subjects precisely yet often leave the world unresponsive; language-driven models handle interactions better but follow complex controls less faithfully; no model combined broad task coverage with consistently strong performance.
Source evidence

Abstract

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.

Comment: Project Website: https://WorldExam.github.io