Running RL asynchronously is the key to faster and cheaper training runs. We've been doing A LOT of research here to make the most performant RL training stack for open weight models.
Check it out!
Applied Compute (@appliedcompute)
Controlling staleness in an async RL stack has not been well understood, so we derived a closed-form formula that predicts staleness in advance. Our predictions match measured staleness from production RL runs within a fraction of a step. Building these simulations has led to key insights on how we maintain high GPU utilization during our RL runs without sacrificing ML performance.
— https://nitter.net/appliedcompute/status/2073119551336919102#m