Twitter/X

April 2025 was a low point

Brief

Meta pivoted from April 2025's Llama 4 failure — an 18 score on the Artificial Analysis Intelligence Index and a shelved Behemoth — to an aggressive rebuild: $14.3B for 49% of Scale AI, Alexandr Wang as chief AI officer, and massive hiring. Muse Spark 1.2 now scores 82.9 on Terminal-Bench 2.1 (behind only Opus 5), powers Muse Code, and was produced via a 27‑day 1.1→1.2 data flywheel.

Why it matters

April 2025 was a low point: Llama 4 underperformed on independent benchmarks, leaked internal messages showed tweaks to inflate evals, the Behemoth project was shelved indefinitely, and Meta's model scored 18 on the Artificial Analysis Intelligence Index.

Key details

  • Mark Zuckerberg responded with heavy moves: $14.3B to acquire 49% of Scale AI, Alexandr Wang installed as chief AI officer at age 28, hundreds-of-millions-of-dollars pay packages to recruit researchers from OpenAI and Google, and a nine-month teardown of Meta's stack, architecture, infrastructure, and data pipelines.
  • Muse Spark's turnaround: Muse Spark arrived in April, then Muse Spark 1.2 (four months later) scored 82.9 on Terminal-Bench 2.1 — ahead of GPT-5.6, Grok 4.5 and Gemini 3.6 and four points behind Opus 5; the 1.1→1.2 jump took 27 days because 1.1 generated and graded training data. Meta is funding this with a $130–145B capex guide (~$375M/day) despite a 91% drop in free cash flow, and has launched Muse Code beta powered by Muse Spark 1.2.
Source evidence

Sixteen months ago, Reddit's top review of Meta's flagship AI model called it a pathetic release from one of the richest corporations on the planet. Today Meta shipped a coding model that trails exactly one lab on earth. The story of those sixteen months is Mark refusing to lose.

April 2025 was the bottom. Llama 4 underperformed on nearly every independent benchmark, leaked internal messages described tweaks to inflate eval scores, and Behemoth got shelved indefinitely. Meta's model scored 18 on the Artificial Analysis Intelligence Index. Eighteen.

Mark's response was to buy his way into the room. $14.3B for 49% of Scale AI. Alexandr Wang installed as chief AI officer at 28. Pay packages worth hundreds of millions to pull researchers from OpenAI and Google, with the new team seated near his own desk. Then nine months tearing down the entire stack, architecture, infrastructure, and data pipelines included.

Muse Spark arrived this April with a rare move. Meta named its own weakness in the launch post, admitting performance gaps in coding and long-horizon agents. Its model trailed Claude Sonnet 4.6, a mid-tier model, on Terminal-Bench Hard.

Four months later, Muse Spark 1.2 posts 82.9 on Terminal-Bench 2.1. Ahead of GPT-5.6. Ahead of Grok 4.5. Ahead of Gemini 3.6. Behind only Opus 5, and by four points. The version jump from 1.1 took 27 days, because 1.1 generated and graded the training data for 1.2. The flywheel now turns in weeks.

Underneath it sits money no one else will spend. Meta guided to $130-145B of capex this year, roughly $375M every single day, more than 2024 and 2025 combined. Free cash flow fell 91% last quarter. He raised the spending floor anyway.

Even the launch chart carries the message. Meta published a graphic where it finishes second on its own internal benchmark. You only ship that chart when you want people watching the slope, not the standings.

Mark Zuckerberg (@finkd)

Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.

— https://nitter.net/finkd/status/2085080750034940201#m