Tag
Introduces SWE-Touch, a benchmark framework that injects conflicting user edits during agent coding trajectories, showing that current coding agents significantly degrade in collaborative settings despite strong standalone benchmark performance.
This paper presents a deep Q-network-based multi-agent reinforcement learning framework for decentralized conflict resolution among heterogeneous small UAVs and eVTOL aircraft operating under degraded surveillance conditions, evaluating policies across 90 combinations of traffic density and separation thresholds.
FlowEdit is a novel framework that uses information-theoretic principles to regulate internal reasoning flows in LLMs, enabling them to generate multiple alternative responses in a single pass for ill-posed problems with conflicting conditions. Experiments show 68% improvement in exact-set-match accuracy and 24% boost in response informativeness over leading proprietary models.
StateFuse is a conflict-aware replicated memory contract for multi-agent systems that preserves contradictory observations rather than collapsing them, enabling safer abstention and auditable correction without universal accuracy gain.
This paper explores using Nonviolent Communication (NVC) principles as lightweight prompt constraints to reduce conversational escalation in LLMs during conflict-prone interactions. Experiments across multiple instruction-tuned models show that NVC-constrained prompting consistently de-escalates dialogue and stabilizes interactions with highly resistant users.
SoCRATES introduces a realistic multi-domain benchmark for evaluating proactive LLM mediators, showing that top models resolve only about one-third of the consensus gap in conflict resolution.