We need open frameworks to evaluate model behavior. Discussions need to be grounded in auditable measurements rather than "us vs them" vibes.
@cyrilgorlla and the team at CTGT are doing important work in this space
Cyril Gorlla (@CyrilGorlla)
@ReedAlbergotti broke it at @semafor this morning, and his question is the one that lingers:
"What is the nationality of an American model distilled from a Chinese model that was distilled from an American models?"
At the 8k token budgets production systems actually run, our 120B scores 83.61% on FinanceReasoning. Above Kimi K3 (81.93%) and Inkling (65.13%). At 62 to 160x lower cost per query, on one H100.
At unlimited budget the big models win on raw accuracy.
— https://nitter.net/CyrilGorlla/status/2082506941570707961#m