humans have had continual learning for 200k+ years and most of us are still stuck in local maxima doing npc jobs. it sounds great but when you try to implement it, you realize the gradients just violently fight each other in the latent space. neuroplasticity just isn't a native property of these architectures. backprop fundamentally wants to find a local minimum and stay there. this forces it into continuous real-time updates, and it fights with the math.
you probably need to solve meta/emprical learning first to figure out what's actually pareto optimal to keep.
Dwarkesh Patel (@dwarkesh_sp)
8 Predictions for the Era of Continual Learning.
Also up on YouTube, pod feed, and Substack.
Video
— https://nitter.net/dwarkesh_sp/status/2085781456375218232#m