Twitter/X

PerceptionBench is a new visual-perception benchmark released by Moonshot that…

Brief

PerceptionBench, released by Moonshot, breaks visual perception into 10 empirically discovered atomic capabilities after analyzing frontier-model failures across 42 benchmarks. It provides 3,000 verified, single-capability questions that require only visual inspection (no external knowledge or reasoning). Public artifacts include a blog post, GitHub repo, and a Hugging Face dataset.

Why it matters

PerceptionBench is a new visual-perception benchmark released by Moonshot that isolates perception into 10 atomic capabilities derived from failures of frontier models across 42 existing benchmarks.

Key details

  • The dataset contains 3,000 verified questions, each designed to isolate a single perceptual capability and be answerable purely by looking (no reasoning or external knowledge required).
  • Resources and code are published publicly: blog announcement, GitHub repository (github.com/MoonshotAI/Percep...), and a Hugging Face dataset (huggingface.co/datasets/moon...).
Source evidence

We are releasing PerceptionBench, a benchmark that isolates visual perception and evaluates it as a set of atomic capabilities - discovered from how today's models fail, rather than defined in advance.

From frontier-model failures across 42 benchmarks, we derive 10 atomic perceptual capabilities and construct 3,000 verified questions, each isolating a single capability and answerable by looking, with no reasoning or external knowledge required.

Blog: kimi.com/blog/perception-ben…
GitHub: github.com/MoonshotAI/Percep…
Hugingface: huggingface.co/datasets/moon…