Twitter/X

On 2026-06-13 @hamptonism predicts that within a year, production agents will…

Brief

The post asserts that agent self-improvement—agents that analyze their production failures, create better evals from them, and iterate autonomously—will soon feel obvious. It highlights Adaline 2.0 (Arsh Shah Dilbagi) as a functioning early layer that turns traces into behaviors and auto-generates evals, tests candidate agents, and surfaces winners for human review.

Why it matters

On 2026-06-13 @hamptonism predicts that within a year, production agents will autonomously read their own failures, convert those failures into improved evaluation tasks, and iteratively sharpen without manual tracing and babysitting.

Key details

  • Arsh Shah Dilbagi’s Adaline 2.0 is presented as an early working example: it converts Traces → Behaviors → Issues → auto-generated Evals/Data, then generates and tests new agent candidates so engineers only review and ship the winners.
Source evidence

in a year this is gonna feel obvious.

agents that quietly read their own production failures → turn them into better evals → and keep getting sharper without you manually babysitting every trace.

early version of that is already here and it’s kinda wild to see working.

Video

Arsh Shah Dilbagi (@arshdilbagi)

Introducing Adaline 2.0 - The Agent Self-Improvement Layer

Adaline turns Traces into Behaviors,
Behaviors surface Issues,
Issues become auto-generated Evals + Data,
Adaline then generates new agent candidates and tests them.

You review the winners and ship!

Video

— https://nitter.net/arshdilbagi/status/2065826083224834345#m