I've had a theory about super smart people who aren't always successful, that it boils down to focusing too much on verifiable rewards with short feedback loops.
It feels like there's a lesson that might carry over to LLMs here.
I've had a theory about super smart people who aren't always successful, that it boils down to focusing too much on verifiable rewards with short feedback loops.
It feels like there's a lesson that might carry over to LLMs here.