Twitter/X

@ankrgyl (2026-03-20) argues that the product "flywheel" relies on observability…

Brief

@ankrgyl (2026-03-20) argues that the product "flywheel" relies on observability (o11y) and continuous interaction with evaluation tooling, and states his team is actively improving that integration. He links an Awesome Agents roundup — “Best LLM Eval Tools in 2026: 6 Options Tested” — comparing DeepEval, Braintrust, Langfuse, LangSmith, Inspect AI, and RAGAS to support that focus on LLM evaluation and testing.

Source evidence

title: @ankrgyl: Glad this resonated because it's been our vision from the start. The flywheel is strongest when you ...
author: @ankrgyl
contenttype: tweet
publication: Twitter/X
published: 2026-03-20T18:36:16+00:00
source
url: https://x.com/ankrgyl/status/2035062888197759400

word_count: 79

Glad this resonated because it's been our vision from the start. The flywheel is strongest when you have o11y and evals feeding each other continuously.

We are doing a lot of work to make this better and more capable

Awesome Agents (@awagents)

Best LLM Eval Tools in 2026: 6 Options Tested

A data-driven comparison of DeepEval, Braintrust, Langfuse, LangSmith, Inspect AI, and RAGAS - the top LLM evaluation frameworks for teams building AI in production.

LlmEvaluation #AiTesting

— https://nitter.net/awagents/status/2035038918505058363#m