Twitter/X

VSTAT (introduced 2026-06-03) reveals a substantial perceptual gap

Brief

VSTAT is a new benchmark showing a large human–MLLM perceptual gap: simple video tasks (counting cups, reading typed words, counting page flips) are easy for humans but often fail for MLLMs. The benchmark targets latent-space tracking of evolving world states, treats text as one probe, and aims to expand to pixels, actions, and other modalities.

Why it matters

VSTAT (introduced 2026-06-03) reveals a substantial perceptual gap: humans solve simple video tasks easily while MLLMs fail on examples like counting cups, reading typed words, and counting page flips.

Key details

  • VSTAT is designed to evaluate evolving world-state tracking in videos in latent space (not just pixel tracking); the authors plan to probe beyond text into modalities such as pixels and actions (vision-x-nyu.github.io/vstat).
Cleaned source text

Working on this benchmark has been a lot of fun along the way—huge shout-out to my amazing collaborators!

Sihyun Yu (@sihyun_yu)

Can MLLMs actually track what's happening in a video?

The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't.

vision-x-nyu.github.io/vstat…

🧵 [1/11]

Video

— https://nitter.net/sihyun_yu/status/2062000569938756054#m