Please @cursor_ai please god fix.
Matthew Lam (@mattlam_)
introducing OpenBench v1, an open framework for measuring AI performance and efficiency for your codebase and use case.
Companies are realizing that you can't simply tokenmaxx, and are frantically looking for better ways to use and measure AI use and efficiency. One approach is evidenced by the rise in model routers: OpenRouter, Cloudflare, Databricks, Vercel, and now Ramp to name a few.
But they'll also need the ability to evaluate how agents are performing in their actual use cases and codebase. A good example is @DoorDash's recent evals on their codebase and prs. This will become even more important as companies explore different model + harness combinations.
OpenBench will be built to this direction, focusing on the cross between correctness vs efficiency in both token use and latency. Starting off, OpenBench makes it easy for anyone to add to the task set, and run a variety of harness + models. The framework also makes it easy to add any custom harness variant, for example I've been testing codex variants with ablations against the stock codex harness. OpenBench will also have tooling to help with discovering/adding reliable tasks from your repo.
Today I have integrated harnesses: codex, claude, cursor, devin, grok build, pi, and eval'ed them for different models like gpt 5.6-sol, GLM 5.2, Kimi K3, Grok 4.5, etc. measuring correctness %, token use (in/out/cache), and latency.
Video
— https://nitter.net/mattlam_/status/2079605387121049605#m