Tag
This paper identifies a 'detectability gap' in hallucination detection: hallucinations split into high-agreement (Ghost) and low-agreement (Flickering) regimes with a 0.35–0.46 AUC gap, persisting across four models and three factual QA datasets even after freezing regime assignments and using stricter trajectory-based tests. The authors argue aggregate detection metrics hide model-dependent heterogeneity and call for regime-conditioned evaluation.
The article discusses a consistency issue in Shopify's WebMCP checkout rollout where structured tool responses can drift from the visible UI, proposing that agents verify user-visible invariants before submission.
The Qwen Image 2.1 Consistency LoRA is a tool that maintains image consistency during edits, preventing drift and unwanted repaints in AI image generation workflows.
This article addresses the consistency problem in AI agents, where tasks may fail on repeated attempts, and introduces ALTK-Evolve's Consistency Analyzer to diagnose and improve reliability, reducing the consistency gap from 24.4pp to 12.0pp without losing average accuracy.
MovieGrid is a multi-grid post-training paradigm that decomposes long videos into spatially arranged chunks to improve multi-shot coherence and efficiency, achieving state-of-the-art intra-shot and inter-shot consistency in video generation.
This paper introduces Prefix-Denoising Consistency (PDC), a test-time verification method for Diffusion Language Models that improves performance on reasoning tasks by using prefix-conditioned regeneration and majority voting.
The article discusses advancements in AI video generation with Wan 3.0 on Magnific, highlighting the ability to maintain consistency across ages and scenes in long-form stories.
AI video technology is evolving from generating clips to enabling full scene direction, which could reduce the cost of consistency in film production and empower small teams to create complete productions.
This paper introduces Cross-Contextual Consistency (C3), a behavioral property for measuring LLM credibility by checking whether answers remain stable under topic-aligned, content-neutral perturbations. Across 26 models and six benchmarks, they find that higher consistency correlates with correctness, offering a complementary evaluation axis.
The article introduces Revision Prompting, a technique for industrial LLM processes that improves speed, cost, and consistency when re-processing updated inputs by generating output patches from diffs.
This paper investigates whether tabular foundation models (TFMs) like TabPFN, TabICL, TabDPT, and TabFM produce predictions consistent with any joint distribution. It demonstrates that all evaluated TFMs violate both marginalization and factorization consistency for classification and regression, questioning their Bayesian inference claims.
SEAM is a generator-agnostic framework that audits global consistency of explanations in scientific machine learning, detecting incompatible local explanations even when predictions are locally accurate and attributing failures to specific channels and overlaps. The paper presents theory and experiments across PDE systems, neural operators, and four open datasets.
This paper proposes a commutation theory for label-free reliability in vision-language figure reading, showing that consistency-based methods have a computable blind spot and introducing an Equivariance-Consistency Score enhanced by cyclic relabeling.
This paper investigates how training-free acceleration can silently change generated content in diffusion-based multimodal large language models, and proposes paired diagnostics and consistency-control methods to mitigate content drift.
This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.
Describes a method for unifying outputs from thousands of agents in parallel forecasting tasks, achieving consistency and cost efficiency through post-processing and homogeneous task design.
This paper investigates how LLMs' answers change under meaning-preserving paraphrases, finding that instance-level behavior is unstable (flip rates >23%) and that single-prompt accuracy masks substantial inconsistency, while a self-paraphrasing strategy can partially recover latent knowledge.
This paper introduces a training-free method to improve revisit consistency in autoregressive generative rendering by using temporal and spatial correspondences from the 3D engine to maintain consistent appearance when the camera revisits locations.
Hallo4D is a model-agnostic framework that leverages large multimodal language models to detect and correct spatial and temporal hallucinations in 3D and 4D generation, improving consistency across viewpoints and time without requiring retraining.
A new UGA study finds that AI chatbots provide inconsistent financial advice that varies by platform and by the gender/race of hypothetical users, urging caution for consumers.