Twitter/X

On 2026-06-03 Sihyun Yu introduced VSTAT, a new benchmark for visual state…

Brief

Sihyun Yu released VSTAT (2026-06-03), a benchmark for visual state tracking that tests MLLMs on simple video tasks — counting cups, reading typed words, and counting page flips — which humans handle easily but models do not. The project positions visual state tracking as a major near-term challenge and ties it to how brains infer and maintain internal world states from noisy inputs.

Why it matters

On 2026-06-03 Sihyun Yu introduced VSTAT, a new benchmark for visual state tracking (vision-x-nyu.github.io/vstat).

Key details

  • VSTAT poses simple video tasks — count cups, read typed words, count page flips — that humans solve easily but current multimodal LLMs (MLLMs) fail to track reliably.
  • Researchers frame visual state tracking as a forthcoming 'grand challenge' for vision and link it to the question of how the brain builds and maintains an internal world state from noisy, incomplete visual observations.
Cleaned source text

how does the brain build and track an internal state of the world from (possibly incomplete and noisy) visual observations?

Sihyun Yu (@sihyun_yu)

Can MLLMs actually track what's happening in a video?

The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't.

vision-x-nyu.github.io/vstat…

🧵 [1/11]

Video

— https://nitter.net/sihyun_yu/status/2062000569938756054#m