@DirhousssiAmine: TRL now supports training on agent harness out of the box through our OpenEnv integration. You can now train using harn…

X AI KOLs Following Tools

Summary

TRL now supports training on agent harnesses out of the box through OpenEnv integration, enabling training with harnesses like opencode.

🚀🚀 TRL now supports training on agent harness out of the box through our OpenEnv integration. You can now train using harnesses like opencode 🧵 https://t.co/Sqp7G8qii8
Original Article
View Cached Full Text

Cached at: 08/06/26, 04:45 PM

🚀🚀 TRL now supports training on agent harness out of the box through our OpenEnv integration.

You can now train using harnesses like opencode

🧵 https://t.co/Sqp7G8qii8

TRL now supports training on agent harness out of the box through our OpenEnv integration.

You can now train using harnesses like opencode

The harness operates within an OpenEnv session. OpenEnv deploys a transparent proxy that forwards its vLLM calls and logs each turn’s token IDs and logprobs. On completion, TRL reconstructs the samples from the message and executes GRPO.

Full working example: self-contained local subprocess sandbox on the DeepCoder problems dataset. Validated on Qwen3-8B

http://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py…

PR details :

Similar Articles

OpenForgeRL: Train Harness-native Agents in Any Environment

Hugging Face Daily Papers

OpenForgeRL is an open-source framework for training harness-based AI agents end-to-end in diverse environments, using a lightweight proxy and Kubernetes orchestrator to enable RL on any harness at scale. It achieves strong results on agentic benchmarks and shows that RL improves agent reliability.

@SergioPaniego: https://x.com/SergioPaniego/status/2067270222671741360

X AI KOLs Timeline

OpenReward environments now integrate directly into TRL's GRPOTrainer via a single OpenRewardSpec, allowing zero-glue-code training against a catalog of RL environments. The integration is experimental and part of a broader effort to make environment and agent RL first-class in TRL.