Incredible how Z. ai literally has their RL infrastructure open source.
The entire OPD post-training of GLM-5.2 took on this slime platform took ~2 days.
github.com/THUDM/slime
Author highlights that Z. ai’s RL stack appears publicly available through the THUDM/slime GitHub repo and reports that running OPD post-training for the GLM-5.2 model on that slime platform took roughly two days, implying accessible, fast RL tooling for model fine-tuning or post-training workflows.
Author claims Z. ai has its reinforcement learning (RL) infrastructure published as open source via the THUDM/slime repository on GitHub.
Incredible how Z. ai literally has their RL infrastructure open source.
The entire OPD post-training of GLM-5.2 took on this slime platform took ~2 days.
github.com/THUDM/slime