@DirhousssiAmine: TRL now supports training on agent harness out of the box through our OpenEnv integration. You can now train using harn…
Summary
TRL now supports training on agent harnesses out of the box through OpenEnv integration, enabling training with harnesses like opencode.
View Cached Full Text
Cached at: 08/06/26, 04:45 PM
🚀🚀 TRL now supports training on agent harness out of the box through our OpenEnv integration.
You can now train using harnesses like opencode
🧵 https://t.co/Sqp7G8qii8
TRL now supports training on agent harness out of the box through our OpenEnv integration.
You can now train using harnesses like opencode
The harness operates within an OpenEnv session. OpenEnv deploys a transparent proxy that forwards its vLLM calls and logs each turn’s token IDs and logprobs. On completion, TRL reconstructs the samples from the message and executes GRPO.
Full working example: self-contained local subprocess sandbox on the DeepCoder problems dataset. Validated on Qwen3-8B
http://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py…
PR details :
Similar Articles
OpenForgeRL: Train Harness-native Agents in Any Environment
OpenForgeRL is an open-source framework for training harness-based AI agents end-to-end in diverse environments, using a lightweight proxy and Kubernetes orchestrator to enable RL on any harness at scale. It achieves strong results on agentic benchmarks and shows that RL improves agent reliability.
@adithya_s_k: You can now train on 350+ RL Environments from OpenReward with TRL with just a few lines of code
OpenReward and TRL now support training on over 350 reinforcement learning environments with minimal code.
@SergioPaniego: https://x.com/SergioPaniego/status/2067270222671741360
OpenReward environments now integrate directly into TRL's GRPOTrainer via a single OpenRewardSpec, allowing zero-glue-code training against a catalog of RL environments. The integration is experimental and part of a broader effort to make environment and agent RL first-class in TRL.
@billxbf: Excited to release Polar, our Agent RL rollout infra for real-world harnesses. Be it Codex, Claude Code, OpenClaw, Herm…
Polar is an agent RL rollout infrastructure that allows using real-world harnesses as training environments without code changes, supporting models like Codex, Claude Code, OpenClaw, and Hermes.
@SergioPaniego: frontier agents are this good partly because the model was trained inside the very harness it ships with great to see t…
Sergio Paniego highlights that frontier agents' performance is due to models being trained inside their deployment harness. The new work 'Polar: Agentic RL on Any Harness at Scale' by NVIDIA AI enables turning harnesses like Codex, Claude Code, Qwen Code, or Pi into RL training environments without modifying their internals.