Twitter/X

Levie (created 2026-03-11, updated 2026-04-04) claims the "harness" — the…

Brief

Levie argues the harness — orchestration that decomposes tasks and routes to the right model — will become as critical as model capability as workloads grow to tens/hundreds of millions of tokens. Composio's benchmark illustrates large per-task cost gaps (Hermes/Pi ≈ $0.39 vs Claude Code $1.47) using Kimi K3 pricing assumptions.

Why it matters

Levie (created 2026-03-11, updated 2026-04-04) claims the "harness" — the orchestration that breaks down work and routes it to the right model — will become the most important variable next to model capability as tasks scale to tens or hundreds of millions of tokens, driving accuracy and cost.

Key details

  • Composio's benchmark reports average per-task costs: Hermes Agent $0.39, Pi Agent $0.40, Codex $0.47, OpenCode $0.51, Kimi Code $0.54, and Claude Code $1.47 (Claude ≈3.7× Pi). Median costs follow the same gap (Pi/Hermes $0.29, Claude $0.72). Costs were computed using Kimi K3 list prices: $3/1M input, $0.30/1M cached input, $15/1M output tokens.
Source evidence

No idea if these specific numbers generalize across tasks, but directionally it’s clear that the harness is going to become the most important variable -right next to model capability- in the AI stack.

The ability for harnesses to break down work in the most efficient way and route to the right model at the right time is going to be a huge variable for maximizing accuracy and reducing costs.

We’re actually still incredibly early in this journey. The harness didn’t matter that much when tasks only took hundreds of thousands or millions of tokens. But as we have tasks that take tens of millions and hundreds of millions of tokens, this becomes a major variable. Huge opportunity ahead.

Composio (@composio)

Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi:

  • $0.39 Hermes Agent
  • $0.40 Pi Agent
  • $0.47 Codex
  • $0.51 OpenCode
  • $0.54 Kimi Code
  • $1.47 Claude Code

The median cost tells the same story: $0.29 in Pi Agent and Hermes, $0.35 in OpenCode, $0.38 in Kimi Code, $0.39 in Codex and $0.72 in Claude Code, so the cost gap holds for a typical task and is not driven by a few expensive runs.

We calculated these costs using Kimi K3’s list prices: $3/1M input tokens, $0.30/1M cached input tokens, and $15/1M output tokens.

— https://nitter.net/composio/status/2083161879489220722#m