Twitter/X

François Chollet calls for open, auditable frameworks to evaluate model behavior…

Brief

François Chollet urges development of open, auditable frameworks for evaluating model behavior and points to Cyril Gorlla and CTGT's work. Gorlla raises provenance questions (quoting Reed Albergotti) and reports that at 8k-token production budgets a 120B model scores 83.61% on FinanceReasoning (vs Kimi K3 81.93%, Inkling 65.13%) on one H100 with 62–160× lower per-query cost, while unlimited-budget large models still have higher raw accuracy.

Why it matters

François Chollet calls for open, auditable frameworks to evaluate model behavior and credits Cyril Gorlla and the CTGT team for important work in that space.

Key details

  • Cyril Gorlla highlighted Reed Albergotti's question: "What is the nationality of an American model distilled from a Chinese model that was distilled from an American models?" to underline evaluation and provenance issues.
  • At 8k-token budgets, the authors report a 120B model scoring 83.61% on FinanceReasoning—above Kimi K3 (81.93%) and Inkling (65.13%)—running on a single H100 at 62–160x lower cost per query; at unlimited budget, larger models win on raw accuracy.
Source evidence

We need open frameworks to evaluate model behavior. Discussions need to be grounded in auditable measurements rather than "us vs them" vibes.

@cyrilgorlla and the team at CTGT are doing important work in this space

Cyril Gorlla (@CyrilGorlla)

@ReedAlbergotti broke it at @semafor this morning, and his question is the one that lingers:

"What is the nationality of an American model distilled from a Chinese model that was distilled from an American models?"

At the 8k token budgets production systems actually run, our 120B scores 83.61% on FinanceReasoning. Above Kimi K3 (81.93%) and Inkling (65.13%). At 62 to 160x lower cost per query, on one H100.

At unlimited budget the big models win on raw accuracy.

— https://nitter.net/CyrilGorlla/status/2082506941570707961#m