multi-agent-debate

Tag

Cards List
#multi-agent-debate

Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph

arXiv cs.CL · 2026-07-21 Cached

Debate-on-Graph (DoG) is a framework that enhances LLM reasoning by leveraging uncertain knowledge graphs (UKGs) with confidence scores, using a heuristic search and multi-agent debate mechanism to produce reliable answers. It achieves state-of-the-art performance on four QA benchmarks.

0 favorites 0 likes
#multi-agent-debate

@qingke_ai: https://x.com/qingke_ai/status/2076354316848550126

X AI KOLs Timeline · 2026-07-12 Cached

MAD-OPD utilizes a multi-teacher debate mechanism to break through the single-teacher distillation ceiling, enabling small models to surpass large teacher models in tool invocation and code generation tasks.

0 favorites 0 likes
#multi-agent-debate

LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI

arXiv cs.AI · 2026-06-17 Cached

This paper introduces LegalHalluLens, a framework for auditing hallucinations in legal AI, providing typed hallucination profiles and a Risk Direction Index to improve trustworthy deployment.

0 favorites 0 likes
#multi-agent-debate

MDForge: Agentic Molecular Dynamics Pipeline Design under Sparse Simulator Feedback

arXiv cs.AI · 2026-06-12 Cached

MDForge is an LLM agent that automates the design of molecular dynamics pipelines for host-guest binding free-energy calculations, achieving human-expert competitive results and discovering a novel high-affinity binder.

0 favorites 0 likes
#multi-agent-debate

The Confident Liar: Diagnosing Multi-Agent Debate with Log-Probabilities and LLM-as-Judge

arXiv cs.CL · 2026-06-10 Cached

This paper studies the relationship between token-level log-probability distributions, LLM-as-judge rubric scores, and final task accuracy in multi-agent debate systems. It finds a consistent four-phase confidence trajectory and role asymmetry between Constructor and Auditor agents.

0 favorites 0 likes
#multi-agent-debate

Debating the Unspoken: Role-Anchored Multi-Agent Reasoning for Half-Truth Detection

arXiv cs.CL · 2026-04-22 Cached

RADAR introduces a role-anchored multi-agent debate framework where Politician and Scientist agents adversarially reason over evidence to detect misleading half-truths, outperforming baselines on omission-aware fact verification.

0 favorites 0 likes
← Back to home

Submit Feedback