Tag
This paper introduces FPCO-Dialog, a benchmark for evaluating correction and cooperation behavior in vision-language models under repeated false premises in multi-turn dialogues.
StateFuse is a conflict-aware replicated memory contract for multi-agent systems that preserves contradictory observations rather than collapsing them, enabling safer abstention and auditable correction without universal accuracy gain.
A clarification post noting that a robot face shown is actually a real robot, unlike a previous misleading post.
Elon Musk announces a new X feature: users who interact with a misleading post later corrected by Community Notes will receive an 𝕏 Chat message with the correction to address misperceptions.
Anthropic issued a correction clarifying that Claude Opus 4's ~3x average speedup dates to May 2025, not May 2024, and that earlier models from May 2024 showed no speedup on the backtested evaluation.
A correction clarifies that the RTX Spark does not have 600GB/s bandwidth; that figure is actually the NvLink speed, as shown in Computex slides.
Explores the need for correction mechanisms in agent memory systems, going beyond storage to include source tracking, confidence levels, expiry, and audit trails.
A correction issued by the Department of War clarifies that SpaceX remains a strong and valued partner, refuting false claims in a news article.
This paper investigates how large language models maintain correct beliefs under adversarial pressure in clinical settings, proposing R-FT fine-tuning to improve epistemic resilience while balancing corrigibility, and demonstrating significant robustness gains on medical benchmarks.
Proposes Correction-Oriented Policy Optimization (CIPO), an extension to RLVR that converts failed trajectories into correction-oriented supervision, improving reasoning and correction performance in LLMs across math and code benchmarks.
Proposes a training-free inference-time method for Vision-Language-Action models to correct pace and path dynamics, improving success rates by up to 28.8% in dynamic environments.