Tag
The article explains how AI models can get stuck in 'doom loops' and introduces 'antidoom' training using Final Token Preference Optimization (FTPO) to reduce this issue, with a tutorial provided on GitHub.
This article describes a method to train LoRA adapters using AsyncGRPOTrainer and sync them via Storage Buckets across separate Hugging Face Jobs, eliminating the need for NCCL communication.
The article details an open-source reproduction of training a coding model to generate watercolour art using TRL and OpenEnv, with full pipeline artifacts published on Hugging Face for reproducibility.
This guide details a fine-tuning recipe using Group Relative Policy Optimization (GRPO) with the TRL library to enhance the LFM2.5-350M model's structured output compliance, improving IFStruct benchmark performance from 22.6% to 29.7%.
Antidoom is a tool that helps small reasoning models avoid repetitive loops in complex tasks, now reaching Technology Readiness Level. It addresses the issue where models get stuck and repeat words during long thinking traces.
Reminder for Class 3 of the Training Agents live series, covering reinforcement learning (GRPO) for training agents, how to implement it in TRL, and end-to-end examples, streamed on Hugging Face's X, YouTube, and LinkedIn on Tuesday, July 28.
TRL now supports training on agent harnesses out of the box through OpenEnv integration, enabling training with harnesses like opencode.
Sharing slides from a talk at seLIA on training an open coding agent using TRL and OpenEnv, part of an open-source conference on free software and open AI.
The article provides a brief history of model distillation in AI and announces an upcoming live stream class on distilling open models using TRL (Transformer Reinforcement Learning).
You can now train Liquid AI's LFM2-VL model using TRL's GRPO and RLOO methods, with an example script provided.
Continuous batching has been added to TRL for GRPO, improving speed and VRAM usage without needing vLLM. The tweet explains how it works and when to use it.
OpenReward and TRL now support training on over 350 reinforcement learning environments with minimal code.
OpenReward environments now integrate directly into TRL's GRPOTrainer via a single OpenRewardSpec, allowing zero-glue-code training against a catalog of RL environments. The integration is experimental and part of a broader effort to make environment and agent RL first-class in TRL.
This post demonstrates how to fine-tune a model for free using a single prompt, leveraging the new Google Colab CLI along with Hugging Face's TRL and trackio tools, all orchestrated by an AI agent.
The user is working on implementing reasoning training with verifiers using Unsloth and TRL, reporting progress on locally generating GRPO-like rollouts with a small SLM and a tiny RM, and promises a video soon.
TRL now supports fine-tuning models on agent traces from various sources like Claude Code, Codex, OpenClaw, and Pi, moving towards a standardized stack for training agentic models.
Announces an upcoming video on training tiny models for preference tuning, covering reward models, RLHF, DPO, ORPO with Unsloth and TRL.
TRL v1.4 is released, featuring chunked NLL loss for SFT to reduce VRAM usage and first-class integration with OpenReward for GRPO.
Hugging Face releases TRL v1.0, a major update to its post-training library that transforms it from a research codebase into a stable, production-ready tool supporting over 75 training methods like PPO and DPO.
Hugging Face publishes a comprehensive analysis of 16 open-source reinforcement learning libraries, examining architectural patterns for asynchronous RL training and presenting design lessons for TRL's async trainer to address generation bottlenecks and weight synchronization challenges.