Twitter/X

Pinzhi Huang and Sihyun Yu (vision-x-nyu) released VSTAT on 2026-06-03, a…

Brief

VSTAT, introduced by Pinzhi Huang and Sihyun Yu (vision-x-nyu) on 2026-06-03, is a new benchmark targeting visual state tracking in videos. It tests time-sensitive skills—counting cups, reading typed text, and tracking page flips—demonstrating that humans handle these tasks easily while leading multimodal language models struggle, highlighting gaps in entity/event identification and temporal state mapping.

Why it matters

Pinzhi Huang and Sihyun Yu (vision-x-nyu) released VSTAT on 2026-06-03, a benchmark measuring visual state tracking in videos.

Key details

  • VSTAT focuses on simple, time-dependent tasks—counting cups, reading typed words, and counting page flips—that humans solve easily but current frontier multimodal/MLLMs perform poorly on.
  • Authors argue state tracking requires identifying entities/events and mapping their state evolution over time, and existing models are surprisingly weak at this core video understanding skill.
Cleaned source text

Frontier multimodal models are surprisingly bad at it, so we built a benchmark to measure it.

Meet VSTAT!

Sihyun Yu (@sihyun_yu)

Can MLLMs actually track what's happening in a video?

The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't.

vision-x-nyu.github.io/vstat…

🧵 [1/11]

Video

— https://nitter.net/sihyun_yu/status/2062000569938756054#m