Twitter/X

Supabase open-sourced “Supabase Evals” (announced by @kiwicopple on 2026-08-03)…

Brief

Supabase Evals is an open‑source benchmark released by @kiwicopple on 2026-08-03 that evaluates how well AI coding agents build with Supabase. It executes agents on real Supabase tasks, applies automated scoring, and provides comparative results — including runs for Claude Code, Codex, and OpenCode — so users can inspect agent performance.

Why it matters

Supabase open-sourced “Supabase Evals” (announced by @kiwicopple on 2026-08-03) as a public benchmark for AI coding agents.

Key details

  • The benchmark runs agents on real Supabase development tasks and automatically scores their outputs against those tasks.
  • Initial comparisons highlighted agent performance for Claude Code, Codex, and OpenCode (results available from the project).
Source evidence

we open-sourced Supabase Evals

it's our benchmark for how well AI coding agents build with @supabase. we run agents against real tasks and score them

see how Claude Code, Codex, and OpenCode perform: