No idea if these specific numbers generalize across tasks, but directionally it’s clear that the harness is going to become the most important variable -right next to model capability- in the AI stack.
The ability for harnesses to break down work in the most efficient way and route to the right model at the right time is going to be a huge variable for maximizing accuracy and reducing costs.
We’re actually still incredibly early in this journey. The harness didn’t matter that much when tasks only took hundreds of thousands or millions of tokens. But as we have tasks that take tens of millions and hundreds of millions of tokens, this becomes a major variable. Huge opportunity ahead.
Composio (@composio)
Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi:
- $0.39 Hermes Agent
- $0.40 Pi Agent
- $0.47 Codex
- $0.51 OpenCode
- $0.54 Kimi Code
- $1.47 Claude Code
The median cost tells the same story: $0.29 in Pi Agent and Hermes, $0.35 in OpenCode, $0.38 in Kimi Code, $0.39 in Codex and $0.72 in Claude Code, so the cost gap holds for a typical task and is not driven by a few expensive runs.
We calculated these costs using Kimi K3’s list prices: $3/1M input tokens, $0.30/1M cached input tokens, and $15/1M output tokens.
— https://nitter.net/composio/status/2083161879489220722#m