Tag
The article discusses the challenge of governing interactions between multiple AI agents, highlighting how individually approved agents can conflict in unexpected ways and poses questions about how organizations handle such inter-agent governance.
GVD introduces a unified framework for governed versioning and deduplication in document repositories, using bidirectional rule alignment and conflict resolution policies with local encoder models to achieve high performance on enterprise data.
This paper formalizes semantic shadowing in mutable RAG and introduces GC-Mem, a temporal dominance-based protocol that resolves conflicts and recovers over 90% accuracy.
The paper introduces ConflictGUI, a benchmark for conflict-aware termination in GUI agents, and proposes ConflictGuard, an inference-time framework to reduce over-compliance and improve performance on conflicting instructions.
CoVer is a factual adjudication framework for addressing evidence-level and aggregation-level conflicts in claim verification, with strong performance evaluated on the ContraNote dataset from X's Community Notes system.
This paper introduces a benchmark to study how large language models arbitrate conflicting evidence from text and numerical sources, finding that models use heuristic strategies with biases towards recency and external tools.
This paper studies how vision-language models shift their reliance between image and text sources when one modality is degraded, revealing task-dependent behavior and introducing a new conflict benchmark for evaluation.
This paper introduces TANGLE, a benchmark for evaluating LLM agents' cognitive behavior under irreducible conflicts in personal memory, assessing dimensions like conflict perception and memory faithfulness.
Introduces SWE-Touch, a benchmark framework that injects conflicting user edits during agent coding trajectories, showing that current coding agents significantly degrade in collaborative settings despite strong standalone benchmark performance.
This paper presents a deep Q-network-based multi-agent reinforcement learning framework for decentralized conflict resolution among heterogeneous small UAVs and eVTOL aircraft operating under degraded surveillance conditions, evaluating policies across 90 combinations of traffic density and separation thresholds.
FlowEdit is a novel framework that uses information-theoretic principles to regulate internal reasoning flows in LLMs, enabling them to generate multiple alternative responses in a single pass for ill-posed problems with conflicting conditions. Experiments show 68% improvement in exact-set-match accuracy and 24% boost in response informativeness over leading proprietary models.
StateFuse is a conflict-aware replicated memory contract for multi-agent systems that preserves contradictory observations rather than collapsing them, enabling safer abstention and auditable correction without universal accuracy gain.
This paper explores using Nonviolent Communication (NVC) principles as lightweight prompt constraints to reduce conversational escalation in LLMs during conflict-prone interactions. Experiments across multiple instruction-tuned models show that NVC-constrained prompting consistently de-escalates dialogue and stabilizes interactions with highly resistant users.
SoCRATES introduces a realistic multi-domain benchmark for evaluating proactive LLM mediators, showing that top models resolve only about one-third of the consensus gap in conflict resolution.