Tag
CompoWorld introduces a method to scale tasks by composing reusable services for training general agents, improving performance by 9.17 points on average across benchmarks and surpassing models like Claude Opus on AutomationBench.
This paper shows that tool-result caching, even if marginally correct, can reverse the expected group-normalized policy updates in reinforcement learning, as demonstrated through mathematical analysis and experiments with a two-action model.
This paper introduces Taste-Bench, a benchmark for measuring taste in LLM agents' long-horizon decisions, finding that frontier models have low accuracy and that taste can be improved through distillation training.
OpenBMB releases UltraData-SFT-Agent-2609, a dataset of 500,000 samples for agent instruction-tuning, used in the post-training of the MiniCPM5-2B model.
The article introduces AMDKernelVault, an open dataset and training framework for AMD GPU kernel optimization, featuring large-scale HIP and Triton kernels and agent-driven pipelines for generating and validating kernels.
Hugging Face shares a replay of a broadcast discussing advanced techniques for training AI agents, emphasizing the shift from reward functions to environment-based approaches.
AgentBrew introduces an offline training framework for tool-use agents that learns from raw interaction trajectories without task verifiers, using retrospective task inference and PMI-based credit assignment to improve performance on real-world applications like GitHub and Notion.
An interview with Brookea Joseph, a Member of Technical Staff at OpenAI, detailing her career from a startup intern to conducting computer use research for AI agents, which she views as crucial for advancing toward AGI.
Runway introduces Solaris, the first AI model in their Interface World Models family, which generates real-time interactive interfaces by synthesizing frames directly, eliminating the need for code and enabling dynamic design and agent training.
Agent Lightning v1.0 is a lightweight agentic reinforcement learning framework by Microsoft, refactored for training AI agents with real harnesses and achieving substantial benchmark improvements, such as a 14.6 percentage point gain on SWE-bench.
The paper introduces WER, a multi-phase framework that trains a Skill Optimizer using reinforcement learning from execution feedback to improve tool-using agents, achieving significant performance gains on benchmarks like BFCL v4 and τ2-bench.
New research from Tencent introduces Recursive Synthetic Terminal Tasks (RST), a method that progressively generates harder, verifiable training tasks for AI agents. Starting from 639 tasks, it produced 37,484 verified tasks across 15 rounds, and reinforcement learning with these tasks improved Qwen3.5-27B from 22.7% to 32.0% on Terminal-Bench Hard.
This paper introduces State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), a method that selectively applies privileged trajectory guidance only when the student agent's current state matches the reference, improving multi-turn agent performance on ALFWorld and WebShop benchmarks.
The paper argues that simply scaling multimodal environments does not always improve agent training, and proposes Ability-aware Environment Selection (AES) and Hierarchical Difficulty Curriculum (HDC) to better structure environment distributions along diversity and difficulty dimensions.
CalibForge is an autonomous terminal-task synthesis system that uses adversarial solver calibration to create learnable tasks for training terminal agents. It constructs 5,431 calibrated tasks and improves agent performance on Terminal-Bench2.0, SWE-bench Pro, and Doc2Repo.
Introduces Deep Research Pretraining (DRP), an offline framework that generates search-open-write trajectories from citation and hyperlink evidence structures. Qwen3-14B models pretrained on 1B tokens with DRP outperform matched no-DRP baselines on deep research benchmarks, even with less supervised fine-tuning data.
SKILL-KD is a contrastive skill distillation framework that improves LLM agents by distilling actionable discrepancies between teacher and student trajectories into textual skill patches, with drift-aware consolidation to iteratively refine skills.
Microsoft Research introduces Echoverse, a set of deep, evolving environments for training computer-use agents. A 9B model trained on these environments nearly doubles its baseline score, coming within 14 points of GPT-5.4, demonstrating that high-fidelity simulation and co-evolution of model, world, and verifier significantly improve agent performance on multi-step workflows.
Introduces Hindsight Policy Optimization (HPO), a novel policy gradient method that uses an intent space and Wasserstein distance to reduce variance in long-horizon language agent training, showing improved stability over GRPO and PPO.
NexForge is a requirement-driven framework that synthesizes diverse, executable agent tasks and expert trajectories for LLM post-training, outperforming prior methods and achieving state-of-the-art open-source agent performance on Terminal-Bench.