We are releasing PerceptionBench, a benchmark that isolates visual perception and evaluates it as a set of atomic capabilities - discovered from how today's models fail, rather than defined in advance.
From frontier-model failures across 42 benchmarks, we derive 10 atomic perceptual capabilities and construct 3,000 verified questions, each isolating a single capability and answerable by looking, with no reasoning or external knowledge required.
Blog: kimi.com/blog/perception-ben…
GitHub: github.com/MoonshotAI/Percep…
Hugingface: huggingface.co/datasets/moon…