Twitter/X

Adam Hunt (@RealAdamHunt) says he has flipped from bullish to bearish on AI…

Brief

Adam Hunt argues that recent progress has produced spiky, domain-specialized models rather than steadily more general intelligence: coding and math capabilities have surged while simple logic and normal English performance have stagnated or degraded (he even notes worse language output compared with 'o3'). He traces the pattern to RL optimization on lucrative, well-recorded tasks (coding) and credits chain-of-thought and web search post-2024 with a temporary boost in cross-domain ability. He references the 'spiky bubble' AGI image (Tomas Pueyo, Nov 2025) and suggests scaling has expanded some spikes at the expense of others. Hunt gives this view ~40% credence, predicts the spikiness will be apparent within 1–2 years, foresees sector-specific RL efforts (health, law) and model-linking, and flags increased cyber-security danger from powerful code-focused models.

Why it matters

Adam Hunt (@RealAdamHunt) says he has flipped from bullish to bearish on AI progress and places about 40% confidence in his bearish thesis.

Key details

  • He claims recent scaled models are becoming more 'spiky' not more general: the coding/math spike has improved (even superhuman in places) while simple logic and everyday language have stagnated or worsened—he specifically says language outputs are worse compared to models like 'o3'.
  • Hunt attributes the spikiness to reinforcement learning concentrated on high-reward, well-recorded domains (notably coding, where benchmarks and training data are abundant) and argues that chain-of-thought plus web search after 2024 temporarily amplified apparent generality.
  • He predicts this pattern will become widely obvious in the next 1–2 years, expects efforts to RL-specialize models for sectors like health or law (or link such specialist models), and warns code-focused superpowered models pose heightened cyber-security risks.
Source evidence

Okay, hear me out.

We turn creating machine general intelligence into a verifiably correct task...

Then we RL the shit out of the models to make them good at it.

Adam Hunt (@RealAdamHunt)

Recently I've flipped from being bullish to being bearish about AI.

I think I'm updating my bearishness to be more solidly bearish. Early thoughts (which I hope to be disproven in the next year or so, I would prefer progress) and my reasoning:

The whole 'it turns out if you keep training and scaling the models more they develop broad new capabilities in lots of domains' thesis is wrong (sorry Demis). The recent batch of models haven't got more general, they've got less general. This is most obvious in the fact that their language outputs have got much worse in comparison to e.g. o3. If they were gaining generalist capacities we would expect them to be describing their work in ever more graceful and comprehensive prose!

The image that was being shared as the AGI thesis (November 2025, Tomas Pueyo) was the spiky bubble that has a current spike or two out past human capabilities (e.g. on coding or math) but below human on other capabilities on the other spikes - the future prediction was that as the models scale/advance, every spike would grow bit by bit until the whole center encompasses the human capabilities, with super-superhuman on some spikes. I think it seems like what's actually happened in the last few models has been that the coding/math spike has grown, but leaving behind or even at the cost of the other spikes. The models are no better at some simple logic, language (and sometimes worse!).

This makes sense from a simple RL perspective; you can't RL something endlessly on one domain of tasks and expect it to improve on the other tasks. The fact that early LLMs did seem to improve generally was a byproduct of the written language corpus covering everything - that corpus is general, so training it on that gave the appearance of something generally intelligent and becoming more generally intelligent as it got better at replicating that corpus. But the actual logic and underlying ground truths behind the language aren't captured efficiently enough and weren't effectively RLd in - they top out at some point (I guess this happened around the time that there was the 'has scaling hit a wall' discussion in late 2024). Chain of thought was then a genuine breakthrough, along with web search, which plugged into that general LLM global-corpus intelligence to lead to post 2024 gains.

The AI companies have since worked out that coding works (and pays) really well (basically this is because the entire job is nearly perfectly recorded and exists as training data, and you can set up clear benchmarks and rewards). The recent models (and benchmarks) have been maxxing that and we've seen degradation on normal English use for that reason. This could still be transformative, leading to extremely powerful (and potentially dangerous, particularly in cyber security) models but it's not a pathway to AGI.

I'm probably at about 40% confidence about this. It fits my current observations of AI progress and has a basic explanatory model. It doesn't account for potential breakthroughs, which is a major reason for discounting.
To make some predictions, I guess if I'm right this will become broadly apparent and more widely acknowledged in the next year or two, as we see how the spikiness of models that keep getting released develops.
Maybe there will be efforts to concentrate on specific spikes e.g. health or law which require going back to earlier models and RLing on a different data set/with different rewards/benchmarks. Maybe those separate models can be linked together to give a more apparently general model. How capital intensive that is/the potential profitability will be a defining question. But I just don't see general abilities emerging atm, and I don't think we will any time soon. Good news - a whole industry of tackling important specific problems/sectors can open up!

— https://nitter.net/RealAdamHunt/status/2082034423344853168#m