Twitter/X

Simon Willison announced smevals, a small evaluation suite for testing models…

Brief

Simon Willison describes smevals, a compact evals framework he built with Jesse Vincent's Prime Radiant applied AI research lab to measure the capabilities of various models, prompts, and harnesses. The post (published 2026-07-31) links to a detailed write-up at simonwillison.net/2026/Jul/3 outlining the framework and its intended use cases.

Why it matters

Simon Willison announced smevals, a small evaluation suite for testing models, prompts, and harnesses, with a blog post linked at simonwillison.net/2026/Jul/3…

Key details

  • smevals was developed in collaboration with Jesse Vincent's Prime Radiant applied AI research lab to answer capability questions across different models and prompt/harness configurations (post published 2026-07-31).
Source evidence

More on my own blog as well: simonwillison.net/2026/Jul/3…

Link

smevals—a small eval suite for evaluating models, prompts, and harnesses

I've been working with Jesse Vincent's Prime Radiant applied AI research lab building out this evals framework to help answer questions about the capabilities of different models. The result is …
simonwillison.net