Tmax: A simple recipe for terminal agents

Hugging Face Daily Papers Papers

Summary

Tmax introduces a simplified RL training recipe for terminal agents, achieving state-of-the-art performance with a 9B parameter model using a novel data generation taxonomy and an expanded open-source dataset.

Terminal-using agents have quickly become the most popular downstream application of language models (LMs). Despite their prevalence, relatively little academic work has examined RL-based training of these models, likely due to difficult benchmarks, a lack of data, and a lack of simple baseline recipes. We present Tmax, the strongest open RL recipe for terminal agents to date, bringing open data recipes closer to the frontier. While simple, our recipe achieves 27\% on Terminal-Bench 2.0 with only 9B parameters, outperforming much larger models from prior work. Concretely, we generate data using a novel taxonomy, combining difficulty control, personas, and verifier diversification, which allows us to cheaply generate large amounts of terminal environments for RL and SFT training. We open-source our terminal dataset, which is over 2.5x larger than previously released terminal-agent datasets. We then train open-weight models using RL with our data, using a simple, outcome-only recipe. We release our data, models, and code as a strong baseline for future open academic work on terminal agents at https://github.com/hamishivi/tmax.
Original Article
View Cached Full Text

Cached at: 06/23/26, 05:40 AM

Paper page - Tmax: A simple recipe for terminal agents

Source: https://huggingface.co/papers/2606.23321

Abstract

A novel RL training approach for terminal agents achieves superior performance using a simplified recipe and expanded dataset, enabling effective training with fewer parameters than previous methods.

Terminal-using agents have quickly become the most popular downstream application oflanguage models(LMs). Despite their prevalence, relatively little academic work has examined RL-based training of these models, likely due to difficult benchmarks, a lack of data, and a lack of simple baseline recipes. We present Tmax, the strongest open RL recipe forterminal agentsto date, bringing open data recipes closer to the frontier. While simple, our recipe achieves 27\% onTerminal-Bench 2.0with only 9B parameters, outperforming much larger models from prior work. Concretely, we generate data using a novel taxonomy, combiningdifficulty control,personas, andverifier diversification, which allows us to cheaply generate large amounts of terminal environments for RL andSFT training. We open-source our terminal dataset, which is over 2.5x larger than previously released terminal-agent datasets. We then train open-weight models using RL with our data, using a simple,outcome-only recipe. We release our data, models, and code as a strong baseline for future open academic work onterminal agentsat https://github.com/hamishivi/tmax.

View arXiv pageView PDFProject pageGitHubAdd to collection

Models citing this paper12

#### allenai/tmax-27b 2.65M• Updatedabout 2 hours ago • 4 #### allenai/tmax-9b 9B• Updatedabout 2 hours ago • 3 #### allenai/qwen35-9b-openthoughts 9B• Updatedabout 2 hours ago • 4 • 2 #### allenai/tmax-2b 2B• Updatedabout 2 hours ago • 1 Browse 12 models citing this paper## Datasets citing this paper11

#### allenai/tmax-15k-open-instruct Updatedabout 2 hours ago • 14 • 1 #### allenai/tmax-sft Updatedabout 2 hours ago • 6 #### allenai/TMax-15K Viewer• Updatedabout 2 hours ago • 14.6k • 2 • 2 #### allenai/open-instruct-endless-terminals Updatedabout 2 hours ago • 1 Browse 11 datasets citing this paper### Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2606.23321 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

TMax: A Simple Recipe for Terminal Agents

Reddit r/LocalLLaMA

TMax presents a straightforward method for building AI agents that operate in terminal environments, combining practical design principles for effective command-line automation.

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Hugging Face Daily Papers

T1 is a 122B Mixture-of-Experts model trained with reinforcement learning for long-horizon terminal tasks, achieving state-of-the-art results on benchmarks like Terminal-Bench 2.1 and surpassing models such as GPT-5.4 and GLM-5.1.