multi-turn

Tag

Cards List
#multi-turn

@thecekbote: Happy to share that MURPHY has been accepted at #NeurIPS2026! As we move toward more agentic training and iterative sel…

X AI KOLs Timeline ↗ · 2d ago Cached

MURPHY introduces a feedback-aware multi-turn extension of GRPO for self-correcting code generation, improving performance on benchmarks by crediting earlier attempts that provide feedback for later successes.

0 favorites 0 likes
#multi-turn

On-Policy Self-Distillation for Multi-Turn Image Editing

Hugging Face Daily Papers ↗ · 2d ago Cached

This paper proposes MT-OPSD, a non-policy self-distillation framework to address degradation in multi-turn image editing and introduces LME-Bench for evaluating long-horizon robustness.

0 favorites 0 likes
#multi-turn

PII-TRACE: A Benchmark for Context-Aware PII Detection in Multi-Turn LLM Conversations

arXiv cs.CL ↗ · 2026-09-22 Cached

Introduces PII-TRACE, the first benchmark for context-aware PII detection in multi-turn LLM conversations, and PII-Tracer, a compact detector that achieves high entity-level coverage.

0 favorites 0 likes
#multi-turn

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

arXiv cs.CL ↗ · 2026-09-21 Cached

This paper diagnoses the gap between next-turn evaluation metrics and autonomous workflow execution success for AI agents, showing that supervised fine-tuning improves turn-level performance but fails to enhance end-to-end workflow success.

0 favorites 0 likes
#multi-turn

$\tau$-Elicitation: Benchmarking multi-turn entity extraction in voice agents

arXiv cs.AI ↗ · 2026-09-15 Cached

The paper introduces τ-Elicitation, a benchmark for evaluating multi-turn entity extraction in voice agents, identifying strategy selection and error recovery as key bottlenecks.

0 favorites 0 likes
#multi-turn

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

arXiv cs.AI ↗ · 2026-09-03 Cached

SCX Router introduces a lightweight GLiClass-based model selection tool that uses a decoder-KV classifier and a task ontology to route LLM tasks, optimizing for speed, cost, and quality without autoregressive generation.

0 favorites 0 likes
#multi-turn

PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks

arXiv cs.AI ↗ · 2026-09-03 Cached

PGPO proposes potential-guided policy optimization for multi-turn agentic tasks, enabling finer-grained credit assignment in LLM post-training and showing strong results on ALFWorld and WebShop benchmarks.

0 favorites 0 likes
#multi-turn

EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants

Hugging Face Daily Papers ↗ · 2026-08-29 Cached

EvoGenUI-Bench introduces a benchmark for evaluating LLMs as multi-turn generative UI assistants, focusing on interface maintenance across 150 tasks with challenges in state propagation and external grounding.

0 favorites 0 likes
#multi-turn

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

arXiv cs.CL ↗ · 2026-08-25 Cached

Proposes HiDiffTIR, a hierarchical difficulty-aware policy optimization framework for multi-turn tool-integrated reasoning in LLM agents, improving performance and tool invocation accuracy through fine-grained credit assignment.

0 favorites 0 likes
#multi-turn

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

arXiv cs.AI ↗ · 2026-08-24 Cached

This paper introduces Multi-Turn Certified Robustness (MTCR), a framework for certifying the safety of large language models against multi-turn jailbreak attacks by providing tighter bounds through compositional methods and safety persistence.

0 favorites 0 likes
#multi-turn

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

Hugging Face Daily Papers ↗ · 2026-08-21 Cached

OmniAssistBench is a benchmark for evaluating omni-modal large language models as real-time video assistants, revealing that current models struggle with visual prompts, context retention, and timely responses.

0 favorites 0 likes
#multi-turn

DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents

arXiv cs.CL ↗ · 2026-08-20 Cached

DART-SD proposes a topology-aware retrieval and tuning framework for self-distillation of LLM-based tool-calling agents, improving policy diversity by correcting only critical topological breakpoints while preserving valid reasoning.

0 favorites 0 likes
#multi-turn

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

arXiv cs.CL ↗ · 2026-08-19 Cached

The paper introduces WER, a multi-phase framework that trains a Skill Optimizer using reinforcement learning from execution feedback to improve tool-using agents, achieving significant performance gains on benchmarks like BFCL v4 and τ2-bench.

0 favorites 0 likes
#multi-turn

SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation

arXiv cs.LG ↗ · 2026-08-18 Cached

SMOPD is a loss-only stabilization method for multi-turn on-policy self-distillation that uses selective token-entropy masking to improve accuracy in dirty-history settings, demonstrating improvements with Qwen3 models.

0 favorites 0 likes
#multi-turn

Mitigating Context Interference for Reliable and Efficient Search Agents

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper systematically studies context interference in multi-turn LLM-based search agents, finding that interference primarily arises from the latest retrieved documents, and introduces a distill-based context refiner to mitigate it. Incorporating context refinement into RL training pipelines significantly improves reliability and efficiency.

0 favorites 0 likes
#multi-turn

Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper proposes a dual-loop self-evolution framework for multi-turn empathetic dialogue, using verifiable emotion feedback to optimize policy and adapt training distribution. On SAGE, it improves Qwen3-8B Overall from 53.87 to 79.24, outperforming uniform emotion-reward RL by 7.23 points.

0 favorites 0 likes
#multi-turn

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper introduces a three-condition experimental framework and a benchmark of 24,300 prompts to study how biased user turns modulate cognitive bias expression in frontier LLMs under multi-turn interactions. It finds that biased conversational context amplifies bias in most models, while explicit bias cues can trigger alignment-related suppression.

0 favorites 0 likes
#multi-turn

DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation

arXiv cs.LG ↗ · 2026-08-03 Cached

This paper proposes DASH-OPD, a discrepancy-aware switching method for on-policy distillation in multi-turn LLM agent training, which adaptively toggles between teacher and student executors based on drift and recovery evidence to improve efficiency and performance.

0 favorites 0 likes
#multi-turn

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

arXiv cs.CL ↗ · 2026-08-03 Cached

M3-DuplexBench is a new multi-turn, multilingual, multidomain benchmark for evaluating full-duplex spoken dialogue systems, supporting English and Japanese across casual conversation and question answering domains.

0 favorites 0 likes
#multi-turn

MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

arXiv cs.AI ↗ · 2026-08-03 Cached

MMShopBench introduces a real-log benchmark for multimodal, multi-turn shopping agents, requiring joint inference from user images and dialogue, with an offline sandbox and training set to improve open-source model performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback