Twitter/X

On Terminal-Bench 2.1, an OpenSpace v2 agent scored 65.2% on a cold run and 78.7%…

Brief

OpenSpace v2 claims to convert agent skills from static files into reusable, evidence-backed capabilities: validated by execution traces, ranked by real-task quality signals, and evolved over time to produce compounding capability rather than one-off runs. The release emphasizes a Skill Hub, controlled evolution, local-first design, and compatibility with multiple agent backbones.

Why it matters

On Terminal-Bench 2.1, an OpenSpace v2 agent scored 65.2% on a cold run and 78.7% on a warm run after using OpenSpace’s skill reuse and evolution mechanism — an absolute +13.5 percentage-point improvement without changing the underlying LLM backbone.

Key details

  • OpenSpace v2 provides a 'Quality-First Skill Hub' where skills are structured packages carrying real-task signals (selected, applied, completed, failed, evolved), attach execution traces as evidence, support controlled evolution, are local-first, and integrate with agents like Claude Code, Codex, OpenClaw, Hermès, and nanobot.
Source evidence

OpenSpace v2 shows strong performance gains on Terminal-Bench 2.1 under the same LLM backbone.

With a cold run, the agent achieves 65.2%. After using OpenSpace’s skill reuse and evolution mechanism, the warm run reaches 78.7%, delivering a +13.5 absolute point improvement without changing the underlying model.

This highlights OpenSpace’s core value: agents do not get better only by switching to larger models. They can also improve by accumulating validated skills, reusing what worked before, and turning real task experience into compounding capability.

With OpenSpace, skills are validated through real task traces, reused across runs, and evolved over time. The agent does not start from scratch every time. It brings back what actually worked.

This is the shift we care about:
From one-off agent execution
to compounding agent capability.

Chao Huang (@huang_chao4969)

OpenSpace v2 is now released 🚀

The Quality-First Skill Hub for AI Agents.

Today’s agents can run tasks, call tools, and write code. But they still lack one critical layer: knowing which skills actually work in the real world.

OpenSpace v2 turns agent skills from “files in a folder” into reusable, reviewable, evidence-backed capabilities.

Every skill can now carry real task signals: whether it was selected, applied, completed, failed, or evolved.

This means agents do not just accumulate skills.

They learn which skills to trust, when to reuse them, and how to improve them over time.

What’s new in OpenSpace v2:
• A Skill Hub for agents
Browse, organize, and reuse agent skills as structured packages instead of scattered prompt files.

• Quality you can actually inspect
Each skill comes with real-task quality signals, so agents know what has worked before.

• Task traces as evidence
Upload execution traces to show how a skill performed in real tasks, not just how it looks in a README.

• Skills that improve over time
OpenSpace supports controlled skill evolution, so useful skills can be refined without losing history or trust.

• Local-first by design
Agents can import and run skills locally while still benefiting from a shared skill ecosystem.

• Built for real agent workflows
Use it through a refreshed dashboard or TUI, across agents like Claude Code, Codex, OpenClaw, Hermès, nanobot, and more.

Agents should not get better by guessing.

They should get better from real work, real outcomes, and real evidence.

👉 Open-Source: github.com/HKUDS/OpenSpace

— https://nitter.net/huang_chao4969/status/2078879870428586490#m