Tag
T2SPO 使用历史交互轨迹提供步骤级奖励信号来优化 LLM 智能体的强化学习策略:借助冻结的 TabPFN 回归器估计每个状态到成功的剩余距离,并将其变化作为辅助信用分配来增强 GRPO,实验在 ALFWorld 和 WebShop 上提升了任务成功率。
The ProVer paper targets pivotal decisions for credit assignment in agentic RL: an LLM judge compares successful and failed rollouts to propose a key trajectory segment, then rollouts sampled before and after that segment estimate its advantage. It improves over GRPO by 9.91% and 7.12% on Qwen3.5-2B and Qwen3.5-4B across ALFWorld, WebShop, and SearchQA.
EvoSteer proposes an online self-evolving graph orchestration paradigm for LLM-based multi-agent systems, using a reference-anchored flow-matching loss (AnchorTB) and validated skill admission to repair failing steps during execution, outperforming baselines across twelve datasets spanning QA, math, code, and decision-making.
提出 ProVer 框架,通过 agentic judge 筛选可能的关键决策段并用前后续采样的终端成功率验证优势值,实现 agentic RL 中细粒度信用分配,在 ALFWorld、WebShop、SearchQA 上较 GRPO 提升最高 9.91%。
The paper introduces Learning What to Skip (LW2S), a method that uses counterfactual credit assignment to optimize multi-agent LLM workflows by selectively skipping components, reducing token cost while maintaining or improving accuracy.
This paper identifies a 'latent evidence-credit gap' in latent visual reasoning for multimodal LLMs and proposes ReaLVR, which supplies visual-evidence supervision to free-running latent trajectories via contrastive correct/wrong answers and relevant/mismatched images. ReaLVR outperforms LVR baselines across three model families (63.7% five-task average on Qwen2.5-VL-7B) and scales robustly to 235B-parameter frontier models.
This paper introduces AlignOPSD to address decision-timestamp mismatch in on-policy self-distillation for long-horizon agents, improving performance on benchmarks like ALFWorld, WebShop, and Search-QA compared to baselines.
This paper introduces Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method for reinforcement learning in large language models, which asymmetrically treats success and failure to enhance exploration and achieve better performance on reasoning tasks.
CounterRoute introduces an online reinforcement-learning framework that jointly learns routing and mode-conditioned responses in dual-mode language models, improving accuracy while reducing inference tokens.
SLCA-GRPO introduces Segment-Locked Credit Assignment to improve reinforcement learning for tool-calling agents by decoupling advantage estimation and using hierarchical rewards, leading to faster convergence and higher accuracy.
The paper introduces RECAP, a redundancy-aware learning method that improves the efficiency of large reasoning models by assigning credit to steps based on their structural role and efficacy, enhancing accuracy while reducing token usage.
This paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that decomposes trajectory reward into per-subtask advantages to improve credit assignment in reinforcement learning for language model agents, showing significant gains on high-heterogeneity agentic benchmarks.
ReCAST proposes a method for per-reward, timestep-dependent credit assignment in diffusion model fine-tuning, separating user preferences from temporal allocation to improve alignment and informativeness.
PC-ALM is a local training method that uses layer-local dynamical systems to propagate supervision credit, enabling the training of up to 1000-layer networks without backpropagation while nearly matching its performance.
The paper proposes GACA, a granularity-adaptive credit assignment method for long-horizon LLM agent reinforcement learning that improves task success by adapting resolution to step importance.
TIGPO proposes a temporal instance-graph policy optimization method that extends graph-based credit assignment across policy updates for long-horizon LLM agents, using persistent transition graphs and revisit slots to improve advantage estimation and performance on benchmarks like ALFWorld and WebShop.
PGPO proposes potential-guided policy optimization for multi-turn agentic tasks, enabling finer-grained credit assignment in LLM post-training and showing strong results on ALFWorld and WebShop benchmarks.
MASkills presents a continual learning framework that optimizes multi-agent LLM systems through agent skills, using skill-conditioned credit assignment and hierarchical aggregation to improve performance on tasks like HotpotQA and GAIA.
CHIME is a credit-aware hierarchical memory framework that separates planning and execution memory banks to improve long-horizon agentic planning by accurately attributing task outcomes and outperforming baselines.
DRACO is a reinforcement learning method that dynamically generates rubrics and redistributes trajectory scores to improve long-horizon agent performance without verifiers.