Box is a great example of a company using evals both on their own product and to share public research.
Check out the work @sidharths00 and team are doing to eval new models on realistic tasks: blog.box.com/how-gpt-56-hand…
Link
How GPT-5.6 handles real enterprise work
Box benchmarks GPT 5.6's Sol, Terra, and Luna models on real enterprise tasks across twelve industries.
blog.box.com
Braintrust (@braintrust)
Box's early evals were manual, with responses marked in a spreadsheet or a Box Note.
With Braintrust, they built a practice of running evals programmatically.
Now product managers have a custom eval runner that pulls versioned, schema-validated datasets out of Braintrust, hits the agent endpoint, grades the outputs, and loads results back in.
Read more → braintrustdata.link/customer…
Video
— https://nitter.net/braintrust/status/2079995778773274726#m