@maximelabonne: Such a cool collab, very happy to see this kind of recipe getting open-sourced Train LFM2.5-2.6B on all the harnesses!
Summary
Maximilien Labonne highlights an open-sourced guide to training models with RL inside real agent harnesses like Claude Code, Codex, and OpenCode, enabling training of any model (e.g., LFM2.5-2.6B) on any task set across harnesses.
View Cached Full Text
Cached at: 10/02/26, 04:34 AM
Such a cool collab, very happy to see this kind of recipe getting open-sourced ❤️
Train LFM2.5-2.6B on all the harnesses!
Adithya S K (@adithya_s_k): 1/ Excited to release The ultimate guide to multi-harness RL
The same model behaves differently in every agent harness. So we built an open way to train any model with RL on any task set, inside the harnesses people actually use, like Claude Code, Codex, and OpenCode, without
Similar Articles
@billxbf: Excited to release Polar, our Agent RL rollout infra for real-world harnesses. Be it Codex, Claude Code, OpenClaw, Herm…
Polar is an agent RL rollout infrastructure that allows using real-world harnesses as training environments without code changes, supporting models like Codex, Claude Code, OpenClaw, and Hermes.
@DirhousssiAmine: TRL now supports training on agent harness out of the box through our OpenEnv integration. You can now train using harn…
TRL now supports training on agent harnesses out of the box through OpenEnv integration, enabling training with harnesses like opencode.
@Suhail: This is a very good entry into post training LLMs with RL. The whole recipe and data is open. Highly recommend!
Hamish Ivison and team release Tmax, an open-source RL-trained terminal agent model that outperforms prior work under standard settings. All data, weights, and rollouts are publicly released.
@SergioPaniego: frontier agents are this good partly because the model was trained inside the very harness it ships with great to see t…
Sergio Paniego highlights that frontier agents' performance is due to models being trained inside their deployment harness. The new work 'Polar: Agentic RL on Any Harness at Scale' by NVIDIA AI enables turning harnesses like Codex, Claude Code, Qwen Code, or Pi into RL training environments without modifying their internals.
OpenForgeRL: Train Harness-native Agents in Any Environment
OpenForgeRL is an open-source framework for training harness-based AI agents end-to-end in diverse environments, using a lightweight proxy and Kubernetes orchestrator to enable RL on any harness at scale. It achieves strong results on agentic benchmarks and shows that RL improves agent reliability.