Tag
MURPHY introduces a feedback-aware multi-turn extension of GRPO for self-correcting code generation, improving performance on benchmarks by crediting earlier attempts that provide feedback for later successes.
This paper proposes MT-OPSD, a non-policy self-distillation framework to address degradation in multi-turn image editing and introduces LME-Bench for evaluating long-horizon robustness.
Introduces PII-TRACE, the first benchmark for context-aware PII detection in multi-turn LLM conversations, and PII-Tracer, a compact detector that achieves high entity-level coverage.
This paper diagnoses the gap between next-turn evaluation metrics and autonomous workflow execution success for AI agents, showing that supervised fine-tuning improves turn-level performance but fails to enhance end-to-end workflow success.
The paper introduces τ-Elicitation, a benchmark for evaluating multi-turn entity extraction in voice agents, identifying strategy selection and error recovery as key bottlenecks.
SCX Router introduces a lightweight GLiClass-based model selection tool that uses a decoder-KV classifier and a task ontology to route LLM tasks, optimizing for speed, cost, and quality without autoregressive generation.
PGPO proposes potential-guided policy optimization for multi-turn agentic tasks, enabling finer-grained credit assignment in LLM post-training and showing strong results on ALFWorld and WebShop benchmarks.
EvoGenUI-Bench introduces a benchmark for evaluating LLMs as multi-turn generative UI assistants, focusing on interface maintenance across 150 tasks with challenges in state propagation and external grounding.
Proposes HiDiffTIR, a hierarchical difficulty-aware policy optimization framework for multi-turn tool-integrated reasoning in LLM agents, improving performance and tool invocation accuracy through fine-grained credit assignment.
This paper introduces Multi-Turn Certified Robustness (MTCR), a framework for certifying the safety of large language models against multi-turn jailbreak attacks by providing tighter bounds through compositional methods and safety persistence.
OmniAssistBench is a benchmark for evaluating omni-modal large language models as real-time video assistants, revealing that current models struggle with visual prompts, context retention, and timely responses.
DART-SD proposes a topology-aware retrieval and tuning framework for self-distillation of LLM-based tool-calling agents, improving policy diversity by correcting only critical topological breakpoints while preserving valid reasoning.
The paper introduces WER, a multi-phase framework that trains a Skill Optimizer using reinforcement learning from execution feedback to improve tool-using agents, achieving significant performance gains on benchmarks like BFCL v4 and τ2-bench.
SMOPD is a loss-only stabilization method for multi-turn on-policy self-distillation that uses selective token-entropy masking to improve accuracy in dirty-history settings, demonstrating improvements with Qwen3 models.
This paper systematically studies context interference in multi-turn LLM-based search agents, finding that interference primarily arises from the latest retrieved documents, and introduces a distill-based context refiner to mitigate it. Incorporating context refinement into RL training pipelines significantly improves reliability and efficiency.
This paper proposes a dual-loop self-evolution framework for multi-turn empathetic dialogue, using verifiable emotion feedback to optimize policy and adapt training distribution. On SAGE, it improves Qwen3-8B Overall from 53.87 to 79.24, outperforming uniform emotion-reward RL by 7.23 points.
This paper introduces a three-condition experimental framework and a benchmark of 24,300 prompts to study how biased user turns modulate cognitive bias expression in frontier LLMs under multi-turn interactions. It finds that biased conversational context amplifies bias in most models, while explicit bias cues can trigger alignment-related suppression.
This paper proposes DASH-OPD, a discrepancy-aware switching method for on-policy distillation in multi-turn LLM agent training, which adaptively toggles between teacher and student executors based on drift and recovery evidence to improve efficiency and performance.
M3-DuplexBench is a new multi-turn, multilingual, multidomain benchmark for evaluating full-duplex spoken dialogue systems, supporting English and Japanese across casual conversation and question answering domains.
MMShopBench introduces a real-log benchmark for multimodal, multi-turn shopping agents, requiring joint inference from user images and dialogue, with an offline sandbox and training set to improve open-source model performance.