Twitter/X

François Chollet wrote he was wrong about the role of LLMs in 2023 and early 2024…

Brief

Gary Marcus cites François Chollet's December 2024 reversal—Chollet says he underestimated LLMs in 2023–early 2024 and flipped after the OpenAI o3 test-time-compute breakthrough—to argue Chollet vindicates Marcus's 2022 position that pure scaling of base LLMs failed; Marcus highlights persistent failures on ARC 1 and simple math and points to TTC/harnesses as the critical advance (o3: 75.7% on ARC-AGI).

Why it matters

François Chollet wrote he was wrong about the role of LLMs in 2023 and early 2024 and says he changed his mind in December 2024 after the OpenAI o3 test-time-compute (TTC) breakthrough.

Key details

  • Gary Marcus asserts Chollet's admission confirms Marcus's 2022 claim that pure scaling of base LLMs didn’t solve AGI; Marcus claims current base LLMs still fail on ARC 1 and basic math and that TTC plus harnesses are critical.
  • OpenAI's o3 system achieved a 75.7% score on the ARC-AGI public leaderboard, cited as the TTC breakthrough that changed Chollet's view.
Source evidence

pure scaling didn’t work. @fchollet confirms exactly what I said in 2022.

François Chollet (@fchollet)

One thing I want to make perfectly clear: back in 2023 and early 2024, I was wrong about the role that LLMs would come to play. I underestimated their long-term importance. I have acknowledged this many times.

This was the moment I changed my mind, in December 2024, following the o3 test-time compute breakthrough: arcprize.org/blog/oai-o3-pub…

I did not initially see that LLMs could work as a base to build systems actually capable of fluid intelligence. Then in late 2024 I updated my views.

And here's what did not happen: the early 2023 narrative that all we needed to solve AGI was scaling up base LLMs did not pan out. To this day, current base LLMs (considerably scaled up compared to the models from that time) still do not perform well on something as easy as ARC 1 -- and can't even reliably do simple math operations. TTC and harnesses are in fact critical, and the TTC breakthrough was not obvious.

Link

OpenAI o3 Breakthrough High Score on ARC-AGI-Pub | ARC Prize

OpenAI o3 scores 75.7% on ARC-AGI public leaderboard.
arcprize.org

— https://nitter.net/fchollet/status/2085416724266983682#m