Working on this benchmark has been a lot of fun along the way—huge shout-out to my amazing collaborators!
Sihyun Yu (@sihyun_yu)
Can MLLMs actually track what's happening in a video?
The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't.
vision-x-nyu.github.io/vstat…
🧵 [1/11]
Video
— https://nitter.net/sihyun_yu/status/2062000569938756054#m