consistency

Tag

Cards List
#consistency

The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models

arXiv cs.CL ↗ · 2d ago Cached

This paper identifies a 'detectability gap' in hallucination detection: hallucinations split into high-agreement (Ghost) and low-agreement (Flickering) regimes with a 0.35–0.46 AUC gap, persisting across four models and three factual QA datasets even after freezing regime assignments and using stricter trajectory-based tests. The authors argue aggregate detection metrics hide model-dependent heterogeneity and call for regime-conditioned evaluation.

0 favorites 0 likes
#consistency

If WebMCP and the visible checkout UI disagree, which one should the agent trust?

Reddit r/AI_Agents ↗ · 3d ago

The article discusses a consistency issue in Shopify's WebMCP checkout rollout where structured tool responses can drift from the visible UI, proposing that agents verify user-visible invariants before submission.

0 favorites 0 likes
#consistency

@wildmindai: Qwen Image 2.1 Frame-Lock LoRA pins edits to the original frame so nothing shifts; - same prompt/edit, zero movement - …

X AI KOLs Timeline ↗ · 3d ago Cached

The Qwen Image 2.1 Consistency LoRA is a tool that maintains image consistency during edits, preventing drift and unwanted repaints in AI image generation workflows.

0 favorites 0 likes
#consistency

Your Agent Aced the Task. Will It Do It Again?

Hugging Face Blog ↗ · 2026-09-15 Cached

This article addresses the consistency problem in AI agents, where tasks may fail on repeated attempts, and introduces ALTK-Evolve's Consistency Analyzer to diagnose and improve reliability, reducing the consistency gap from 24.4pp to 12.0pp without losing average accuracy.

0 favorites 0 likes
#consistency

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

Hugging Face Daily Papers ↗ · 2026-09-06 Cached

MovieGrid is a multi-grid post-training paradigm that decomposes long videos into spatially arranged chunks to improve multi-shot coherence and efficiency, achieving state-of-the-art intra-shot and inter-shot consistency in video generation.

0 favorites 0 likes
#consistency

Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models

arXiv cs.LG ↗ · 2026-08-27 Cached

This paper introduces Prefix-Denoising Consistency (PDC), a test-time verification method for Diffusion Language Models that improves performance on reasoning tasks by using prefix-conditioned regeneration and majority voting.

0 favorites 0 likes
#consistency

@manishkumar_dev: This is a great example of how AI video generation is moving beyond short clips. Same face, 3 ages, 13 scenes, and a st…

X AI KOLs Timeline ↗ · 2026-08-24 Cached

The article discusses advancements in AI video generation with Wan 3.0 on Magnific, highlighting the ability to maintain consistency across ages and scenes in long-form stories.

0 favorites 0 likes
#consistency

@heyshrutimishra: The film industry's real moat was never a studio lot. It was the cost of consistency, keeping the same face, same world…

X AI KOLs Timeline ↗ · 2026-08-17 Cached

AI video technology is evolving from generating clips to enabling full scene direction, which could reduce the cost of consistency in film production and empower small teams to create complete productions.

0 favorites 0 likes
#consistency

Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper introduces Cross-Contextual Consistency (C3), a behavioral property for measuring LLM credibility by checking whether answers remain stable under topic-aligned, content-neutral perturbations. Across 26 models and six benchmarks, they find that higher consistency correlates with correctness, offering a complementary evaluation axis.

0 favorites 0 likes
#consistency

Revision Prompting improves industrial LLM processes

Lobsters Hottest ↗ · 2026-08-08 Cached

The article introduces Revision Prompting, a technique for industrial LLM processes that improves speed, cost, and consistency when re-processing updated inputs by generating output patches from diffs.

0 favorites 0 likes
#consistency

Do Tabular Foundation Models Agree with Themselves?

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper investigates whether tabular foundation models (TFMs) like TabPFN, TabICL, TabDPT, and TabFM produce predictions consistent with any joint distribution. It demonstrates that all evaluated TFMs violate both marginalization and factorization consistency for classification and regression, questioning their Bayesian inference claims.

0 favorites 0 likes
#consistency

SEAM: Global consistency beyond local accuracy in scientific machine learning

arXiv cs.LG ↗ · 2026-08-07 Cached

SEAM is a generator-agnostic framework that audits global consistency of explanations in scientific machine learning, detecting incompatible local explanations even when predictions are locally accurate and attributing failures to specific channels and overlaps. The paper presents theory and experiments across PDE systems, neural operators, and four open datasets.

0 favorites 0 likes
#consistency

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper proposes a commutation theory for label-free reliability in vision-language figure reading, showing that consistency-based methods have a computable blind spot and introducing an Equivariance-Consistency Score enhanced by cyclic relabeling.

0 favorites 0 likes
#consistency

Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper investigates how training-free acceleration can silently change generated content in diffusion-based multimodal large language models, and proposes paired diagnostics and consistency-control methods to mitigate content drift.

0 favorites 0 likes
#consistency

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Hugging Face Daily Papers ↗ · 2026-07-31 Cached

This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.

0 favorites 0 likes
#consistency

Consistently unifying work from thousands of agents

Reddit r/AI_Agents ↗ · 2026-07-28

Describes a method for unifying outputs from thousands of agents in parallel forecasting tasks, achieving consistency and cost efficiency through post-processing and homogeneous task design.

0 favorites 0 likes
#consistency

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

arXiv cs.AI ↗ · 2026-07-28 Cached

This paper investigates how LLMs' answers change under meaning-preserving paraphrases, finding that instance-level behavior is unstable (flip rates >23%) and that single-prompt accuracy masks substantial inconsistency, while a self-paraphrasing strategy can partially recover latent knowledge.

0 favorites 0 likes
#consistency

Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering

Hugging Face Daily Papers ↗ · 2026-07-23 Cached

This paper introduces a training-free method to improve revisit consistency in autoregressive generative rendering by using temporal and spatial correspondences from the 3D engine to maintain consistent appearance when the camera revisits locations.

0 favorites 0 likes
#consistency

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

Hugging Face Daily Papers ↗ · 2026-07-15 Cached

Hallo4D is a model-agnostic framework that leverages large multimodal language models to detect and correct spatial and temporal hallucinations in 3D and 4D generation, improving consistency across viewpoints and time without requiring retraining.

0 favorites 0 likes
#consistency

AI chatbots lack consistency in financial advice, according to new UGA study

Reddit r/ArtificialInteligence ↗ · 2026-07-08 Cached

A new UGA study finds that AI chatbots provide inconsistent financial advice that varies by platform and by the gender/race of hypothetical users, urging caution for consumers.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback