Tag
Debate-on-Graph (DoG) is a framework that enhances LLM reasoning by leveraging uncertain knowledge graphs (UKGs) with confidence scores, using a heuristic search and multi-agent debate mechanism to produce reliable answers. It achieves state-of-the-art performance on four QA benchmarks.
MAD-OPD utilizes a multi-teacher debate mechanism to break through the single-teacher distillation ceiling, enabling small models to surpass large teacher models in tool invocation and code generation tasks.
This paper introduces LegalHalluLens, a framework for auditing hallucinations in legal AI, providing typed hallucination profiles and a Risk Direction Index to improve trustworthy deployment.
MDForge is an LLM agent that automates the design of molecular dynamics pipelines for host-guest binding free-energy calculations, achieving human-expert competitive results and discovering a novel high-affinity binder.
This paper studies the relationship between token-level log-probability distributions, LLM-as-judge rubric scores, and final task accuracy in multi-agent debate systems. It finds a consistent four-phase confidence trajectory and role asymmetry between Constructor and Auditor agents.
RADAR introduces a role-anchored multi-agent debate framework where Politician and Scientist agents adversarially reason over evidence to detect misleading half-truths, outperforming baselines on omission-aware fact verification.