Twitter/X

Supabase launched “Supabase Evals,” a benchmark that runs AI coding agents…

Brief

Supabase introduced Supabase Evals, a benchmarking suite that runs AI coding agents such as Claude Code, Codex, and Open Code against real tasks and scores their performance. @paulg argues this will become universal: services will adopt agent-compatible benchmarks or risk obsolescence, since agent-incompatible services will fail.

Why it matters

Supabase launched “Supabase Evals,” a benchmark that runs AI coding agents (Claude Code, Codex, Open Code) on real Supabase tasks and scores their output.

Key details

  • @paulg predicts every service will adopt similar agent-focused benchmarks, arguing that services unusable by agents will eventually go out of business (“limit case = all services”).
Source evidence

This seems like a great idea. I bet one day all services used by agents will do this. Which in the limit case = all services, since those that can't be used by agents will go out of business.

Supabase (@supabase)

Introducing Supabase Evals.

Our benchmark for how well AI coding agents build with Supabase. We run agents like Claude Code, Codex, and Open Code against real tasks and score what they do.

— https://nitter.net/supabase/status/2083282155170340898#m