Twitter/X

@simonw (Simon Willison) published an example report on 2026-07-31 showing…

Brief

Simon Willison shared an example report (2026-07-31) that shows the results of running and grading an evaluation suite across different AI models. The linked static.simonwillison.net/sta… report contains the comparative grading output and demonstrates the format and detail the evaluation suite can produce when assessing multiple models.

Why it matters

@simonw (Simon Willison) published an example report on 2026-07-31 showing results from running and grading an evaluation suite against multiple models

Key details

  • The report is hosted at static.simonwillison.net/sta… and demonstrates comparative model grading output produced by the evaluation suite
Source evidence

Here's an example of the kind of report it can produce having run and graded an evaluation suite against different models static.simonwillison.net/sta…