Twitter/X

Daniel Higgins (@Daniel00Higgins) posted on 2026-03-11 pleading “Please…

Brief

Daniel Higgins pleads with @cursor_ai to "please god fix" (post created 2026-03-11), attaching Matthew Lam's 2026-07-21 thread introducing OpenBench v1 — an open framework for measuring AI correctness versus efficiency across token use and latency, with harnesses (codex, claude, cursor, devin, grok build, pi) and evaluations of gpt 5.6-sol, GLM 5.2, Kimi K3, Grok 4.5.

Why it matters

Daniel Higgins (@Daniel00Higgins) posted on 2026-03-11 pleading “Please @cursor_ai please god fix.”

Key details

  • Matthew Lam's 2026-07-21 thread introduces OpenBench v1: an open framework measuring correctness vs efficiency (token use and latency), supports custom harnesses and task discovery, includes harnesses codex, claude, cursor, devin, grok build, pi, and reports correctness %, token in/out/cache, and latency across models like gpt 5.6-sol, GLM 5.2, Kimi K3, Grok 4.5.
Source evidence

Please @cursor_ai please god fix.

Matthew Lam (@mattlam_)

introducing OpenBench v1, an open framework for measuring AI performance and efficiency for your codebase and use case.

Companies are realizing that you can't simply tokenmaxx, and are frantically looking for better ways to use and measure AI use and efficiency. One approach is evidenced by the rise in model routers: OpenRouter, Cloudflare, Databricks, Vercel, and now Ramp to name a few.

But they'll also need the ability to evaluate how agents are performing in their actual use cases and codebase. A good example is @DoorDash's recent evals on their codebase and prs. This will become even more important as companies explore different model + harness combinations.

OpenBench will be built to this direction, focusing on the cross between correctness vs efficiency in both token use and latency. Starting off, OpenBench makes it easy for anyone to add to the task set, and run a variety of harness + models. The framework also makes it easy to add any custom harness variant, for example I've been testing codex variants with ablations against the stock codex harness. OpenBench will also have tooling to help with discovering/adding reliable tasks from your repo.

Today I have integrated harnesses: codex, claude, cursor, devin, grok build, pi, and eval'ed them for different models like gpt 5.6-sol, GLM 5.2, Kimi K3, Grok 4.5, etc. measuring correctness %, token use (in/out/cache), and latency.

Video

— https://nitter.net/mattlam_/status/2079605387121049605#m