Twitter/X

On 2026-07-22 @ankrgyl highlighted Box's public benchmarking of GPT-5.6 (models…

Brief

Box benchmarked GPT-5.6’s Sol, Terra, and Luna models across twelve industries (blog by @sidharths00) and publicly shared results. The company transitioned from manual spreadsheet/Box Note evaluations to a programmatic Braintrust-based workflow, giving product managers a custom eval runner that uses versioned, schema-validated datasets to call agents, grade outputs, and ingest results.

Why it matters

On 2026-07-22 @ankrgyl highlighted Box's public benchmarking of GPT-5.6 (models Sol, Terra, Luna) across twelve industries in the blog 'How GPT-5.6 handles real enterprise work' led by @sidharths00 and team.

Key details

  • Box moved from manual evals (spreadsheet/Box Note) to a programmatic pipeline with Braintrust: a custom eval runner that pulls versioned, schema-validated datasets, calls agent endpoints, grades outputs, and loads results back into Braintrust.
Source evidence

Box is a great example of a company using evals both on their own product and to share public research.

Check out the work @sidharths00 and team are doing to eval new models on realistic tasks: blog.box.com/how-gpt-56-hand…

Link

How GPT-5.6 handles real enterprise work

Box benchmarks GPT 5.6's Sol, Terra, and Luna models on real enterprise tasks across twelve industries.
blog.box.com

Braintrust (@braintrust)

Box's early evals were manual, with responses marked in a spreadsheet or a Box Note.

With Braintrust, they built a practice of running evals programmatically.

Now product managers have a custom eval runner that pulls versioned, schema-validated datasets out of Braintrust, hits the agent endpoint, grades the outputs, and loads results back in.

Read more → braintrustdata.link/customer…

Video

— https://nitter.net/braintrust/status/2079995778773274726#m