An AI lab says it beat the two best models on the planet. The exam it used as proof has not been released, and the same lab wrote it.
They used a technique called Mixture of Agents. Several LLMs answer the same prompt in parallel, then an aggregator model reads every answer and writes a synthesized one on top. Together AI's original version scored 65.1% on AlpacaEval 2.0 against GPT-4o's 57.5%, using only open models. The method is real and the lift is real.
The ceiling is the part that gets skipped. An ensemble cannot answer a question that none of its members can answer. It recovers more of what those models already collectively know, and it stops the strongest one from being wrong as often. New capability never enters the system. So "beyond the frontier" really means the public frontier, coordinated, missing less.
Where MoA wins matters too. AlpacaEval 2.0 and MT-Bench both hand grading to an LLM judge, and both reward longer, fuller answers. An aggregator fusing five drafts produces exactly that. Part of the margin comes from the judge liking the shape of the output before correctness even enters the picture.
So "8% higher than Opus 4.8 and 11% higher than GPT 5.5" lands differently once you see where it was measured. The lab's own benchmark, still unshipped, scored in the format that flatters an aggregator most.
Latency is the bill. Proposers inside a layer run in parallel, but the layers run in sequence, so response time floors at the slowest model in each layer, stacked. One answer feels fine. An agent chaining forty calls feels every layer, every time.
MoA earns its place as a way to squeeze more out of models you already pay for. Price it honestly: single-digit accuracy gains, a multiple on latency, measured on a test nobody outside the lab has seen.
Nous Research (@NousResearch)
The strongest models are gated and access is granted only to a select few.
Hermes Agent now exposes MoA presets as virtual models, giving you capabilities beyond the publicly available frontier: 8% higher than Opus 4.8 and 11% higher than GPT 5.5 on our upcoming benchmark.
Video
— https://nitter.net/NousResearch/status/2070610321278988385#m