Frontier multimodal models are surprisingly bad at it, so we built a benchmark to measure it.
Meet VSTAT!
Sihyun Yu (@sihyun_yu)
Can MLLMs actually track what's happening in a video?
The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't.
vision-x-nyu.github.io/vstat…
🧵 [1/11]
Video
— https://nitter.net/sihyun_yu/status/2062000569938756054#m