Twitter/X

After deployment agents produce thousands of traces, drift over time, and fail in…

Brief

Adxtyahq argues the real challenge for agents is post-deployment: thousands of traces, model drift, and rare-edge failures that are impractical to hunt manually, so eval generation is becoming its own infrastructure layer. Arsh Shah Dilbagi’s Adaline 2.0 formalizes this: converting traces into behaviors and issues, auto-generating evaluations and data, creating and testing candidate agents, and surfacing winners for human review and deployment.

Why it matters

After deployment agents produce thousands of traces, drift over time, and fail in rare edge cases; this creates an operational bottleneck because teams lack time to manually find and triage failures (post dated 2026-06-13, author @adxtyahq).

Key details

  • Arsh Shah Dilbagi announced Adaline 2.0, an "Agent Self-Improvement Layer" that pipelines Traces → Behaviors → Issues → auto-generated Evals + Data → new agent candidates + tests, so humans only review and ship the winning candidates.
Source evidence

everyone wants agents that can use more tools, bigger context windows, better models

but the actual problem starts after deployment

your agent generates thousands of traces, drifts over time, fails in weird edge cases and nobody has time to manually find any of it

feels like eval generation is becoming its own infrastructure layer

Arsh Shah Dilbagi (@arshdilbagi)

Introducing Adaline 2.0 - The Agent Self-Improvement Layer

Adaline turns Traces into Behaviors,
Behaviors surface Issues,
Issues become auto-generated Evals + Data,
Adaline then generates new agent candidates and tests them.

You review the winners and ship!

Video

— https://nitter.net/arshdilbagi/status/2065826083224834345#m