Twitter/X

Mixture of Agents (MoA) runs several LLMs in parallel and an aggregator…

Brief

Mixture of Agents (MoA) ensembles parallel LLM drafts and an aggregator synthesis; Together AI's MoA hit 65.1% on AlpacaEval 2.0 vs GPT-4o 57.5% using open models, but the exam wasn't released and was written by the lab. MoA recovers collective knowledge and favours judge-driven benchmarks that reward long answers, while adding sequential-layer latency for only single-digit accuracy gains.

Why it matters

Mixture of Agents (MoA) runs several LLMs in parallel and an aggregator synthesizes their answers; Together AI's MoA scored 65.1% on AlpacaEval 2.0 vs GPT-4o's 57.5% using only open models, and the exam used to claim superiority was not released and was authored by the lab.

Key details

  • Ensembles cannot create new capabilities: MoA only recovers more of what member models collectively know and reduces errors from the strongest model, so gains reflect coordination across the public frontier rather than new model knowledge.
  • MoA disproportionately helps benchmarks that use LLM judges and reward longer, fuller answers (AlpacaEval 2.0, MT-Bench); Nous/Hermes advertises '8% higher than Opus 4.8 and 11% higher than GPT 5.5' on an unshipped benchmark, while MoA increases latency (response time floors at the slowest model per sequential layer) and yields single-digit accuracy gains at much higher time cost.
Source evidence

An AI lab says it beat the two best models on the planet. The exam it used as proof has not been released, and the same lab wrote it.

They used a technique called Mixture of Agents. Several LLMs answer the same prompt in parallel, then an aggregator model reads every answer and writes a synthesized one on top. Together AI's original version scored 65.1% on AlpacaEval 2.0 against GPT-4o's 57.5%, using only open models. The method is real and the lift is real.

The ceiling is the part that gets skipped. An ensemble cannot answer a question that none of its members can answer. It recovers more of what those models already collectively know, and it stops the strongest one from being wrong as often. New capability never enters the system. So "beyond the frontier" really means the public frontier, coordinated, missing less.

Where MoA wins matters too. AlpacaEval 2.0 and MT-Bench both hand grading to an LLM judge, and both reward longer, fuller answers. An aggregator fusing five drafts produces exactly that. Part of the margin comes from the judge liking the shape of the output before correctness even enters the picture.

So "8% higher than Opus 4.8 and 11% higher than GPT 5.5" lands differently once you see where it was measured. The lab's own benchmark, still unshipped, scored in the format that flatters an aggregator most.

Latency is the bill. Proposers inside a layer run in parallel, but the layers run in sequence, so response time floors at the slowest model in each layer, stacked. One answer feels fine. An agent chaining forty calls feels every layer, every time.

MoA earns its place as a way to squeeze more out of models you already pay for. Price it honestly: single-digit accuracy gains, a multiple on latency, measured on a test nobody outside the lab has seen.

Nous Research (@NousResearch)

The strongest models are gated and access is granted only to a select few.

Hermes Agent now exposes MoA presets as virtual models, giving you capabilities beyond the publicly available frontier: 8% higher than Opus 4.8 and 11% higher than GPT 5.5 on our upcoming benchmark.

Video

— https://nitter.net/NousResearch/status/2070610321278988385#m