Twitter/X

Twitter user @ankrgyl (2026-04-07) argues that sandboxing evaluation runs both…

Brief

Twitter user @ankrgyl (2026-04-07) argues that sandboxing evaluation runs both improves reproducibility and lets teams test many ideas at scale. They report Braintrust now natively supports sandboxed eval execution via AWS Lambda and @modal, with additional execution options promised soon, positioning Braintrust for scalable, reproducible model evaluation.

Source evidence

title: @ankrgyl: Sandboxing evals is an incredible way to (a) get more reproducibility and (b) test a lot more ideas ...
author: @ankrgyl
contenttype: tweet
publication: Twitter/X
published: 2026-04-07T21:47:34+00:00
source
url: https://x.com/ankrgyl/status/2041634010691203089

word_count: 37

Sandboxing evals is an incredible way to (a) get more reproducibility and (b) test a lot more ideas at scale.

This is now natively supported in Braintrust with support for AWS Lambda, @modal, and more options soon.