Tag
The paper introduces GUI-SD-v2, a two-stage on-policy self-distillation framework for multi-turn GUI agents that enhances privilege following and guidance, achieving superior performance on AndroidWorld and MobileWorld benchmarks.
KC-Bench is a dynamic interactive benchmark for evaluating how LLM agents detect and resolve knowledge conflicts between user instructions, parametric knowledge, and environmental observations, with evaluations showing no model handles all conflict types reliably.
This paper introduces Feedback-Aware Credit Assignment (Faca) to improve multi-turn tool-using language agents by using next-turn user reactions as local credit signals, showing significant performance gains on interactive benchmarks.
This paper introduces Alien Abduction, an interactive game to probe how LLMs acquire evidence, update hypotheses, and decide when to stop during abductive reasoning. It finds that models perform better with upfront evidence and with oracle-provided examples than with self-selected queries, revealing deficiencies in active information acquisition.
This paper introduces CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and CanvasAgent, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interactions and hybrid reward optimization.
This paper proposes a knowledge-enhanced visual diagnostic system for traditional Chinese medicine that uses a Neo4j knowledge graph, a four-stage symptom matching pipeline, and an information gain-driven proactive questioning strategy to improve transparency and interpretability. Results demonstrate significant improvements in diagnostic trust and reduced cognitive load.
StepPO introduces a step-centric paradigm for agentic reinforcement learning that aligns policy optimization with agent decision granularity, outperforming token-centric methods in multi-turn interaction tasks.
This paper introduces BALAR, a training-free Bayesian agentic loop algorithm that enables large language models to actively reason and ask clarifying questions in multi-turn interactions. It demonstrates significant performance improvements over baselines on detective, puzzle, and clinical diagnosis benchmarks.