Twitter/X

François Chollet says he was wrong in 2023 and early 2024 about the role LLMs…

Brief

François Chollet admits he was wrong in 2023 and early 2024 about underestimating LLMs, saying he updated his view in December 2024 after OpenAI's o3 test-time compute (TTC) breakthrough. He cites o3's 75.7% on the ARC-AGI public leaderboard and argues TTC and external harnesses—not scaling base LLMs alone—are essential, since base models still fail ARC 1 and simple math.

Why it matters

François Chollet says he was wrong in 2023 and early 2024 about the role LLMs would play and updated his view in December 2024 after OpenAI's o3 test-time compute (TTC) breakthrough.

Key details

  • OpenAI o3 scored 75.7% on the ARC-AGI public leaderboard; Chollet argues this result and the TTC breakthrough show that test-time compute and external 'harnesses' are critical to achieving fluid intelligence.
  • Chollet asserts that simply scaling base LLMs did not solve AGI: current base LLMs, even scaled, still do not perform well on ARC 1 and 'can't even reliably do simple math operations.'
Source evidence

One thing I want to make perfectly clear: back in 2023 and early 2024, I was wrong about the role that LLMs would come to play. I underestimated their long-term importance. I have acknowledged this many times.

This was the moment I changed my mind, in December 2024, following the o3 test-time compute breakthrough: arcprize.org/blog/oai-o3-pub…

I did not initially see that LLMs could work as a base to build systems actually capable of fluid intelligence. Then in late 2024 I updated my views.

And here's what did not happen: the early 2023 narrative that all we needed to solve AGI was scaling up base LLMs did not pan out. To this day, current base LLMs (considerably scaled up compared to the models from that time) still do not perform well on something as easy as ARC 1 -- and can't even reliably do simple math operations. TTC and harnesses are in fact critical, and the TTC breakthrough was not obvious.

Link

OpenAI o3 Breakthrough High Score on ARC-AGI-Pub | ARC Prize

OpenAI o3 scores 75.7% on ARC-AGI public leaderboard.
arcprize.org