agent-training

Tag

Cards List
#agent-training

Show HN: I RL-trained an agent that trains models with RL (for –$1.3k)

Hacker News Top ↗ · 2026-07-14 Cached

A developer built a pipeline where an RL-trained AI agent creates and submits RL training jobs for small models, rewarding the agent for better performance. The project is fully open-sourced and demonstrates transfer to held-out tasks.

0 favorites 0 likes
#agent-training

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

arXiv cs.AI ↗ · 2026-07-08 Cached

TurnOPD introduces turn-level budgeting for on-policy distillation of long-horizon agents, addressing inefficiencies in vanilla OPD by adaptive rollout-depth and progressive turn-normalized loss budgeting, achieving better accuracy under equal training budgets.

0 favorites 0 likes
#agent-training

@qingke_ai: https://x.com/qingke_ai/status/2071281892964659384

X AI KOLs Timeline ↗ · 2026-06-28 Cached

该文章由ROLL团队分享了在终端环境中进行Agentic RL训练时的实践经验,包括环境管理器设计、异步训练管线以及多种模式切换,并对比了RLVR与Agentic RL的本质区别。

0 favorites 0 likes
#agent-training

The Verification Horizon: No Silver Bullet for Coding Agent Rewards

arXiv cs.AI ↗ · 2026-06-26 Cached

该论文指出,对于当前的编码智能体,验证解决方案比生成解决方案更为困难,且任何固定的奖励函数都无法随着能力增长而持续有效。作者通过四种奖励构建的实验表明,针对性的验证设计可以抑制奖励黑客行为并提升任务完成质量。

0 favorites 0 likes
#agent-training

@gabepereyra: Harvey partnered with @appliedcompute to train a legal agent. We optimized each part of the agent stack, including the …

X AI KOLs Following ↗ · 2026-06-22 Cached

Harvey partnered with Applied Compute to train a legal agent, optimizing the agent stack and post-training the GLM-5.1 model using reward signals from their Legal Agent Benchmark.

0 favorites 0 likes
#agent-training

@TheTuringPost: 10 open-source tools for the Agent RL stack ↓ OpenPipe ART verl-agent Agent Lightning Unsloth OpenRLHF SkyRL NVIDIA’s P…

X AI KOLs Timeline ↗ · 2026-06-21 Cached

A curated roundup of 10 open-source tools for training AI agents using reinforcement learning, covering frameworks like OpenPipe ART, verl-agent, Agent Lightning, and Unsloth, with details on their use cases and strengths.

1 favorites 1 likes
#agent-training

@akshay_pachaar: Karpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single rewar…

X AI KOLs Following ↗ · 2026-06-19 Cached

Karpathy's critique of reward functions in RL is addressed by OpenPipe's ART framework using RULER, which allows natural language reward definitions evaluated by an LLM, replacing manual reward engineering.

0 favorites 0 likes
#agent-training

@ben_burtenshaw: https://x.com/ben_burtenshaw/status/2067615361428545566

X AI KOLs Timeline ↗ · 2026-06-18 Cached

A detailed tutorial on supervised fine-tuning (SFT) for training AI agents, built from scratch in pure PyTorch using Qwen3-0.6B, explaining the mechanics of next-token prediction and label masking.

0 favorites 0 likes
#agent-training

SENTINEL: Failure-Driven Reinforcement Learning for Training Tool-Using Language Model Agents

arXiv cs.CL ↗ · 2026-06-12 Cached

This paper introduces SENTINEL, a failure-driven reinforcement learning framework for training tool-using language model agents. It uses a Controller-Proposer-Solver loop to generate targeted training tasks from failed trajectories, improving performance on benchmarks.

0 favorites 0 likes
#agent-training

OpenEnv is now owned by HF, Torch, Prime Intellect, Unsloth, Modal, Mercor, and more! Use it for training agents.

Reddit r/LocalLLaMA ↗ · 2026-06-08

OpenEnv, a tool for creating agentic execution environments like terminals and browsers, is transitioning to a more open governance model with a committee including Hugging Face, Meta-PyTorch, Nvidia, and others to promote open-source agent training.

0 favorites 0 likes
#agent-training

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills

Hugging Face Daily Papers ↗ · 2026-06-05 Cached

Socratic-SWE introduces a closed-loop self-evolution framework for software engineering agents that leverages historical solving traces to generate targeted repair tasks, achieving 50.40% on SWE-bench Verified after three iterations.

0 favorites 0 likes
#agent-training

What Makes Interaction Trajectories Effective for Training Terminal Agents?

arXiv cs.AI ↗ · 2026-06-03 Cached

This paper investigates what makes interaction trajectories effective for training terminal-based AI agents, introducing the Terminal-Lego pipeline and revealing a pedagogical paradox where weaker agents can produce better training data. It finds that environment-grounded supervision, rather than teacher performance, is key for student generalization.

0 favorites 0 likes
#agent-training

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

arXiv cs.CL ↗ · 2026-06-03 Cached

This paper proposes WRIT, a pipeline for synthesizing multi-turn agent training trajectories that balance write-intensive and read-heavy complexity. The method generates diverse tasks and simulations, enabling small models to achieve strong performance with reduced inference cost.

0 favorites 0 likes
#agent-training

@DAIEvolutionHub: MICROSOFT JUST OPEN-SOURCED A WAY TO “TRAIN” AI AGENTS WITHOUT TOUCHING MODEL WEIGHTS SkillOpt treats a simple markdown…

X AI KOLs Timeline ↗ · 2026-05-28 Cached

Microsoft open-sourced SkillOpt, a method that treats markdown skill files like neural network parameters to train AI agents without modifying model weights, using learning rates, validation checks, minibatches, and epochs.

0 favorites 0 likes
#agent-training

@athleticKoder: https://x.com/athleticKoder/status/2057091692235481560

X AI KOLs Timeline ↗ · 2026-05-20 Cached

A technical blog post that explains how to build agent training systems from first principles using a text-to-diagram agent as an example, covering environment definition, teacher trajectory generation, student fine-tuning, and reinforcement learning.

0 favorites 0 likes
#agent-training

Trainer

Product Hunt ↗ · 2026-05-19

Trainer is a tool that lets users train AI agents by recording their screen, available on Product Hunt.

0 favorites 0 likes
#agent-training

EnvSimBench: A Benchmark for Evaluating and Improving LLM-Based Environment Simulation

arXiv cs.AI ↗ · 2026-05-11 Cached

This paper introduces EnvSimBench, a benchmark for evaluating Large Language Models' ability to simulate environments for agent training. It identifies a 'state change cliff' in current LLMs and proposes a constraint-driven pipeline to reduce hallucinations and costs.

0 favorites 0 likes
#agent-training

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis

arXiv cs.CL ↗ · 2026-04-20 Cached

EnvScaler is an automated framework for scaling tool-interactive environments for LLM agents through programmatic synthesis, creating 191 diverse environments and 7K scenarios to improve agent performance on multi-turn, multi-tool interactions.

0 favorites 0 likes
#agent-training

CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

arXiv cs.CL ↗ · 2026-04-20 Cached

CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.

0 favorites 0 likes
#agent-training

Mind DeepResearch Technical Report

Hugging Face Daily Papers ↗ · 2026-04-17 Cached

MindDR is a multi-agent deep research framework using a three-agent architecture (Planning, DeepSearch, Report) and a four-stage training pipeline, achieving competitive performance with ~30B-parameter models on multiple benchmarks. Developed by Li Auto and deployed as an online product, it also introduces MindDR Bench, a 500-query Chinese benchmark for evaluating deep research capabilities.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback