Twitter/X

@kyliestew: got to sit down & talk about how we use evals to make sure our dashboard agent actually solves r...

got to sit down & talk about how we use evals to make sure our dashboard agent actually solves real user problems :) thanks for the chat, @braintrust!

Braintrust (@braintrust)

Cloudflare's dashboard agent spans its entire developer platform, from deploying Workers to debugging production instances. Keeping that much surface area performing well requires more than gut checks.

Their team uses Braintrust to run LLM-as-a-judge scorers to measure conversation resolution, gates every skill and prompt change with evals in CI/CD, and benchmarks sub-agents against models of varying capacity before choosing which to use.

Read more → braintrustdata.link/evals-cl…

Video

— https://nitter.net/braintrust/status/2085763574983524381#m