trl

Tag

Cards List
#trl

@helloiamleonie: models can get stuck in "doom loops" when • the model is small, • it has thinking capabilities, • and the task is compl…

X AI KOLs Following ↗ · 2026-09-15 Cached

The article explains how AI models can get stuck in 'doom loops' and introduces 'antidoom' training using Final Token Preference Optimization (FTPO) to reduce this issue, with a tutorial provided on GitHub.

0 favorites 0 likes
#trl

Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL

Hugging Face Blog ↗ · 2026-09-10 Cached

This article describes a method to train LoRA adapters using AsyncGRPOTrainer and sync them via Storage Buckets across separate Hugging Face Jobs, eliminating the need for NCCL communication.

0 favorites 0 likes
#trl

Training a coding model to paint watercolours with TRL and OpenEnv

Hugging Face Blog ↗ · 2026-09-03 Cached

The article details an open-source reproduction of training a coding model to generate watercolour art using TRL and OpenEnv, with full pipeline artifacts published on Hugging Face for reproducibility.

0 favorites 0 likes
#trl

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

Hugging Face Blog ↗ · 2026-09-03 Cached

This guide details a fine-tuning recipe using Group Relative Policy Optimization (GRPO) with the TRL library to enhance the LFM2.5-350M model's structured output compliance, improving IFStruct benchmark performance from 22.6% to 29.7%.

0 favorites 0 likes
#trl

@maximelabonne: Antidoom is now in TRL! Remove your doom loops with this one simple trick.

X AI KOLs Timeline ↗ · 2026-08-28 Cached

Antidoom is a tool that helps small reasoning models avoid repetitive loops in complex tasks, now reaching Technology Readiness Level. It addresses the issue where models get stuck and repeat words during long thinking traces.

0 favorites 0 likes
#trl

@SergioPaniego: quick reminder! tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series what: reinforce…

X AI KOLs Following ↗ · 2026-07-27 Cached

Reminder for Class 3 of the Training Agents live series, covering reinforcement learning (GRPO) for training agents, how to implement it in TRL, and end-to-end examples, streamed on Hugging Face's X, YouTube, and LinkedIn on Tuesday, July 28.

0 favorites 0 likes
#trl

@DirhousssiAmine: TRL now supports training on agent harness out of the box through our OpenEnv integration. You can now train using harn…

X AI KOLs Following ↗ · 2026-07-24 Cached

TRL now supports training on agent harnesses out of the box through OpenEnv integration, enabling training with harnesses like opencode.

0 favorites 0 likes
#trl

@SergioPaniego: sharing the slides from today’s talk at seLIA (https://selia.codeberg.page) on how to train an open coding agent using …

X AI KOLs Timeline ↗ · 2026-07-06 Cached

Sharing slides from a talk at seLIA on training an open coding agent using TRL and OpenEnv, part of an open-source conference on free software and open AI.

0 favorites 0 likes
#trl

@SergioPaniego: before model distillation was an attack vector. it was. pretty handy way of improving model performance on a task you c…

X AI KOLs Timeline ↗ · 2026-07-03 Cached

The article provides a brief history of model distillation in AI and announces an upcoming live stream class on distilling open models using TRL (Transformer Reinforcement Learning).

0 favorites 0 likes
#trl

@SergioPaniego: you can now train @liquidai's LFM2-VL in TRL GRPO and RLOO included, with an example script

X AI KOLs Following ↗ · 2026-06-25 Cached

You can now train Liquid AI's LFM2-VL model using TRL's GRPO and RLOO methods, with an example script provided.

0 favorites 0 likes
#trl

@SergioPaniego: continuous batching just landed in TRL for GRPO at 64 generations it runs faster and uses less VRAM than plain generate…

X AI KOLs Following ↗ · 2026-06-19 Cached

Continuous batching has been added to TRL for GRPO, improving speed and VRAM usage without needing vLLM. The tweet explains how it works and when to use it.

0 favorites 0 likes
#trl

@adithya_s_k: You can now train on 350+ RL Environments from OpenReward with TRL with just a few lines of code

X AI KOLs Following ↗ · 2026-06-17 Cached

OpenReward and TRL now support training on over 350 reinforcement learning environments with minimal code.

0 favorites 0 likes
#trl

@SergioPaniego: https://x.com/SergioPaniego/status/2067270222671741360

X AI KOLs Timeline ↗ · 2026-06-17 Cached

OpenReward environments now integrate directly into TRL's GRPOTrainer via a single OpenRewardSpec, allowing zero-glue-code training against a catalog of RL environments. The integration is experimental and part of a broader effort to make environment and agent RL first-class in TRL.

0 favorites 0 likes
#trl

@SergioPaniego: https://x.com/SergioPaniego/status/2066498136273531363

X AI KOLs Timeline ↗ · 2026-06-15 Cached

This post demonstrates how to fine-tune a model for free using a single prompt, leveraging the new Google Colab CLI along with Hugging Face's TRL and trackio tools, all orchestrated by an AI agent.

0 favorites 0 likes
#trl

@neural_avb: Lurking the Reasoning Training docs rn. Time to write a verifiers env and Unsloth/TRL that shit! Video soon if it all g…

X AI KOLs Timeline ↗ · 2026-06-11 Cached

The user is working on implementing reasoning training with verifiers using Unsloth and TRL, reporting progress on locally generating GRPO-like rollouts with a small SLM and a tiny RM, and promises a video soon.

0 favorites 0 likes
#trl

@adithya_s_k: You can now finetune models on agent traces directly with TRL Claude Code traces Codex traces OpenClaw traces Pi traces…

X AI KOLs Following ↗ · 2026-06-04 Cached

TRL now supports fine-tuning models on agent traces from various sources like Claude Code, Codex, OpenClaw, and Pi, moving towards a standardized stack for training agentic models.

0 favorites 0 likes
#trl

@neural_avb: Next video is on training tiny (<1B) models for preference tuning. Plus how to generate preference datasets with local …

X AI KOLs Timeline ↗ · 2026-05-26 Cached

Announces an upcoming video on training tiny models for preference tuning, covering reward models, RLHF, DPO, ORPO with Unsloth and TRL.

0 favorites 0 likes
#trl

@QGallouedec: TRL v1.4 is out! two things I'm excited about: → chunked NLL loss for SFT. Way less VRAM, same loss, often faster. Qwen…

X AI KOLs Following ↗ · 2026-05-09 Cached

TRL v1.4 is released, featuring chunked NLL loss for SFT to reduce VRAM usage and first-class integration with OpenReward for GRPO.

0 favorites 0 likes
#trl

TRL v1.0: Post-Training Library Built to Move with the Field

Hugging Face Blog ↗ · 2026-03-31 Cached

Hugging Face releases TRL v1.0, a major update to its post-training library that transforms it from a research codebase into a stable, production-ready tool supporting over 75 training methods like PPO and DPO.

0 favorites 0 likes
#trl

Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries

Hugging Face Blog ↗ · 2026-03-10 Cached

Hugging Face publishes a comprehensive analysis of 16 open-source reinforcement learning libraries, examining architectural patterns for asynchronous RL training and presenting design lessons for TRL's async trainer to address generation bottlenecks and weight synchronization challenges.

0 favorites 1 likes
← Back to home

Submit Feedback