Twitter/X

At Databricks, AI coding token spend is growing exponentially; GLM 5.2, Opus 4.8…

Brief

Databricks reports exponential growth in AI coding token spend and identifies GLM 5.2, Opus 4.8, and GPT 5.6‑Sol as the best quality-per-dollar models on their Coding Bench. They note Opus 5.0 regresses on cost versus 4.8, argue hard budgets misalign incentives, and recommend investing in routing, harnesses, evals, and an open/proprietary model mix to optimize economics.

Why it matters

At Databricks, AI coding token spend is growing exponentially; GLM 5.2, Opus 4.8, and GPT 5.6‑Sol sit on the “efficiency frontier” as the best quality-per-dollar models.

Key details

  • Opus 5.0 shows cost regressions versus Opus 4.8, demonstrating that newer model versions are not always more cost-efficient.
  • Hard budgets are the wrong primitive—Databricks warns biggest AI spenders can be the most AI-leveraged engineers, and says there is no single best model: routing, harnesses, evals, and the right mix of open vs proprietary models materially change economics, and Databricks is investing heavily in those four areas.
Source evidence

ok DBRX gets it

Yuchen Jin (@Yuchenj_UW)

At Databricks, AI coding token spend is growing exponentially.

  1. The “efficiency frontier” matters. On Databricks Coding Bench, GLM 5.2, Opus 4.8, and GPT 5.6-Sol sit on that frontier: the best quality per dollar.

  2. We see cost regressions when comparing Opus 5.0 to 4.8. Newer model doesn’t always mean more efficient.

  3. Hard budgets are the wrong primitive. Your biggest AI spenders may also be your most AI-leveraged engineers.

  4. There is no best model for every task. Routing, harnesses, evals, and the right mix of open and proprietary models can radically change the economics.

We’re investing heavily in all 4, both for ourselves and for our customers.

— https://nitter.net/Yuchenj_UW/status/2085779009913430237#m