Twitter/X

@0xDevShah (2026-08-07) asserts humans have practiced continual learning for…

Brief

Dev Shah (@0xDevShah) argues humans have exhibited continual learning for 200k+ years, but translating that to ML fails because gradients 'violently fight' in latent space and backprop drives models into local minima, preventing native neuroplasticity. He contends solving meta/empirical learning is necessary to identify Pareto‑optimal retention, and points to Dwarkesh Patel's predictions on continual learning.

Why it matters

@0xDevShah (2026-08-07) asserts humans have practiced continual learning for 200k+ years, yet most people remain 'stuck in local maxima' performing 'npc jobs'.

Key details

  • Current neural architectures resist continual learning: gradients 'violently fight' in latent space, neuroplasticity isn't native, and backprop inherently finds and remains in local minima, forcing awkward continuous real‑time updates.
  • He argues we must first solve meta/empirical learning to determine what is Pareto‑optimal to retain; he links Dwarkesh Patel's '8 Predictions for the Era of Continual Learning' for related perspectives.
Source evidence

humans have had continual learning for 200k+ years and most of us are still stuck in local maxima doing npc jobs. it sounds great but when you try to implement it, you realize the gradients just violently fight each other in the latent space. neuroplasticity just isn't a native property of these architectures. backprop fundamentally wants to find a local minimum and stay there. this forces it into continuous real-time updates, and it fights with the math.

you probably need to solve meta/emprical learning first to figure out what's actually pareto optimal to keep.

Dwarkesh Patel (@dwarkesh_sp)

8 Predictions for the Era of Continual Learning.

Also up on YouTube, pod feed, and Substack.

Video

— https://nitter.net/dwarkesh_sp/status/2085781456375218232#m