To note, base LLMs still suffer from the same limitations wrt generalization that were discussed at length, by many, in an extensive body of academic research pre-2024. Scaling them up did not fix those limitations (and they still do not beat ARC 1). New techniques did.