Twitter/X

Hamish Ivison (ICML) introduced Tmax, an open RL-trained terminal-agent model…

Brief

Tmax is an open-release set of RL-trained terminal-agent models from Hamish Ivison (presented at ICML) that reportedly outperforms prior open efforts on terminal use under default settings and a 65,000-token budget. The authors publish the full recipe, training data, model weights, and rollout logs; researcher Suhail publicly recommended the work on 2026-06-27.

Why it matters

Hamish Ivison (ICML) introduced Tmax, an open RL-trained terminal-agent model family that outperforms prior open work on terminal tasks under default settings.

Key details

  • Tmax evaluations used shorter context/token budgets (65,000 tokens) and the team is releasing all data, model weights, and rollout logs publicly.
  • Researcher @Suhail endorsed the release (2026-06-27) as a strong, open recipe for post-training LLMs with reinforcement learning, recommending it for study.
Source evidence

This is a very good entry into post training LLMs with RL. The whole recipe and data is open. Highly recommend!

Hamish Ivison @ ICML (@hamishivi)

Trained some terminal agents with friends!

Introducing Tmax, open RL terminal agent models. Under default settings and shorter length (65k) token budgets, tmax outperforms prior open work on terminal use. We are releasing all data+weights+rollouts publically!

— https://nitter.net/hamishivi/status/2069047986920071263#m