Twitter/X

Elon Musk posted on 2026-07-21 that Grok Build is now powered by Grok 4.5, adding…

Brief

Grok Build is now powered by Grok 4.5 (announced by Elon Musk on 2026-07-21), introducing a native subagent view, Plan Mode integration, mouse support and a fullscreen terminal UI; installation is shown via curl -fsSL https://x.ai/cli/inst.... XFreeze claims Grok 4.5 ranked #1 on the Long-Horizon Terminal-Bench by binary pass rate—beating Claude Fable 5, Claude Opus 4.8 and GPT-5.6-sol—and argues this demonstrates superior ability to maintain context, recover from mistakes and complete multi-hundred-step workflows in real-world coding and automation tasks.

Why it matters

Elon Musk posted on 2026-07-21 that Grok Build is now powered by Grok 4.5, adding a native subagent view, Plan Mode integration, mouse support and a fullscreen terminal UI; install shown as curl -fsSL https://x.ai/cli/inst...

Key details

  • XFreeze reports Grok 4.5 ranked #1 on the Long-Horizon Terminal-Bench by binary pass rate, outperforming Claude Fable 5, Claude Opus 4.8 and GPT-5.6-sol
  • Under the benchmark’s strict metric—counting only tasks with a perfect reward and zero errors—Grok 4.5 “finished clearly ahead,” a claim framed as evidence of stronger long-horizon agentic performance for multi-step coding, automation and engineering workflows
Source evidence

Try Grok Build, it’s awesome!

X.ai/cli

Link

Grok Build

Grok Build is now powered by Grok 4.5, our new model — with a native subagent view, Plan Mode integration, mouse support, and a fullscreen terminal UI. Install with curl -fsSL https://x.ai/cli/inst...
x.ai

X Freeze (@XFreeze)

Grok 4.5 now also ranked #1 on the Long-Horizon Terminal-Bench by binary pass rate, outperforming Claude Fable 5, Claude Opus 4.8 and GPT-5.6-sol

Under the strictest scoring metric - where a task counts only if it is fully solved with a perfect reward and zero errors......Grok 4.5 finished clearly ahead of every other tested model

This matters for real-world coding, automation and difficult engineering work because long-horizon terminal tasks require much more than making partial progress

The model has to maintain context, recover from mistakes and successfully complete an entire workflow across hundreds of steps

Grok 4.5 is showing serious strength on complex agentic tasks over extended periods of time

— https://nitter.net/XFreeze/status/2079264002316927312#m