Tmax: A simple recipe for terminal agents
Summary
Tmax introduces a simplified RL training recipe for terminal agents, achieving state-of-the-art performance with a 9B parameter model using a novel data generation taxonomy and an expanded open-source dataset.
View Cached Full Text
Cached at: 06/23/26, 05:40 AM
Paper page - Tmax: A simple recipe for terminal agents
Source: https://huggingface.co/papers/2606.23321
Abstract
A novel RL training approach for terminal agents achieves superior performance using a simplified recipe and expanded dataset, enabling effective training with fewer parameters than previous methods.
Terminal-using agents have quickly become the most popular downstream application oflanguage models(LMs). Despite their prevalence, relatively little academic work has examined RL-based training of these models, likely due to difficult benchmarks, a lack of data, and a lack of simple baseline recipes. We present Tmax, the strongest open RL recipe forterminal agentsto date, bringing open data recipes closer to the frontier. While simple, our recipe achieves 27\% onTerminal-Bench 2.0with only 9B parameters, outperforming much larger models from prior work. Concretely, we generate data using a novel taxonomy, combiningdifficulty control,personas, andverifier diversification, which allows us to cheaply generate large amounts of terminal environments for RL andSFT training. We open-source our terminal dataset, which is over 2.5x larger than previously released terminal-agent datasets. We then train open-weight models using RL with our data, using a simple,outcome-only recipe. We release our data, models, and code as a strong baseline for future open academic work onterminal agentsat https://github.com/hamishivi/tmax.
View arXiv pageView PDFProject pageGitHubAdd to collection
Models citing this paper12
#### allenai/tmax-27b 2.65M• Updatedabout 2 hours ago • 4
#### allenai/tmax-9b 9B• Updatedabout 2 hours ago • 3
#### allenai/qwen35-9b-openthoughts 9B• Updatedabout 2 hours ago • 4 • 2
#### allenai/tmax-2b 2B• Updatedabout 2 hours ago • 1
Browse 12 models citing this paper## Datasets citing this paper11
#### allenai/tmax-15k-open-instruct Updatedabout 2 hours ago • 14 • 1 #### allenai/tmax-sft Updatedabout 2 hours ago • 6 #### allenai/TMax-15K Viewer• Updatedabout 2 hours ago • 14.6k • 2 • 2 #### allenai/open-instruct-endless-terminals Updatedabout 2 hours ago • 1 Browse 11 datasets citing this paper### Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.23321 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
TMax: A Simple Recipe for Terminal Agents
TMax presents a straightforward method for building AI agents that operate in terminal environments, combining practical design principles for effective command-line automation.
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
T1 is a 122B Mixture-of-Experts model trained with reinforcement learning for long-horizon terminal tasks, achieving state-of-the-art results on benchmarks like Terminal-Bench 2.1 and surpassing models such as GPT-5.4 and GLM-5.1.
@hamishivi: Trained some terminal agents with friends! Introducing Tmax, open RL terminal agent models. Under default settings and …
Introducing Tmax, open reinforcement learning terminal agent models that outperform prior open work on terminal use. All data, weights, and rollouts are being released publicly.
@Suhail: This is a very good entry into post training LLMs with RL. The whole recipe and data is open. Highly recommend!
Hamish Ivison and team release Tmax, an open-source RL-trained terminal agent model that outperforms prior work under standard settings. All data, weights, and rollouts are publicly released.
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
LiteCoder-Terminal-Gen introduces a zero-dependency synthetic pipeline that generates executable terminal training environments, producing SFT and RL datasets that enable language agents to achieve significant performance gains on Terminal Bench benchmarks.