Random ramble on why I think robotics will be solved soon and why it’s symbiotic to what is being built by the labs:
Intelligence is prediction which is equivalent to optimal compression.
This was first talked about and proven mathematically in the 60s (Solomonoff) but was under-discussed until Ilya came along.
The concept is pretty simple. By forcing an algo to predict the next word you force it to first compress reality into its simplest possible form in order to then have a robust enough representation off which it can model and predict.
LeCun was skeptical that text-based prediction was enough to model reality. So far he’s been proven wrong under his original goal posts (glorified search) which were far too pessimistic re how far we can take LLMs as new tricks like inference-time search & RL alignment bridged us to where we are now by ameliorating compounding errors.
There is a narrative that maxis are naively assuming that we can simply continue to scale up LLMs without ever addressing the impossibility of using text to represent spatial environments (gravity, object permanence etc). This is false.
We always understood the need for other forms of abstraction but what we appreciated is that once you have built a medium to automate the process of digital abstraction you meaningfully speed up the time it takes to jumpstart new vectors of compression.
Because if you need to compress to predict and the process of compressing the physical world is ultimately a digital endeavor (via sensors) then text-based compression is complimentary to predicting abstract states in latent space….not orthogonal.
I would guess this is why Anthropic hasn’t bothered with video or world models.
I would also guess this is why Musk and Sam are rushing to get sensors out into the real world.
We are going to get a GPT type of moment for robotics sooner than you think because we are getting very good at compression which is directly transferable from LLMs to the VLA models driving robots.
And one layer is on top of the other. The human brain analyzes what the eyes see and the hand feels. One feeds the other.
What is not talked about enough is the tipping point that follows when we have sufficient robotic capabilities to justify the rollout into the economy which in turn drives a massive step up in data collection that can be fed back to LLMs.
Because yes solving the spatial domain is valuable in and of itself. Thats how you get material abundance within our current societal construct. It’s enough to basically just replace what humans do now with a humanoid with a low enough error rate that it basically nukes the marginal cost of everything down to zero.
But most of what is truly valuable in our economy was invented by human minds abstracting away at night in their kitchen as they compressed all the spatial input they collected over the course of their work day or while reading books (compressing compression).
Once we have the loop going what we’ll see is an explosion of new inventions as the model synthesizes what it observes while doing menial tasks and abstracts to a point where it can introduce bounded vol which is what we humans call that imagination.
This becomes an easily automatable process when models can touch the real world.
To be clear I think models will be doing this long before we solve robotics. I’m just saying solving robotics massively increases the space within which models can explore and invent.
Thoughts welcome.
(P.S. Not AI slop but might as well be because this is a topic that is way outside my circle of competence as a pod boi)