Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading

arXiv cs.LG Papers

Summary

This paper proposes a commutation theory for label-free reliability in vision-language figure reading, showing that consistency-based methods have a computable blind spot and introducing an Equivariance-Consistency Score enhanced by cyclic relabeling.

arXiv:2608.05675v1 Announce Type: new Abstract: Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at. We act on the complementary relation, equivariance: edit a figure's data and the correct answer must change by a computable amount. Two matched edits are provably complete for affine reading errors; no suite of swap edits is complete for label permutations, and cyclic relabeling closes most of that gap. We instantiate the theory as the Equivariance-Consistency Score, a label-free, training-free detector, and release REND-EQUIV, pairing matched invariance and equivariance sets over identical data. The predicted ordering holds across three models and a hand-labeled population immune to the one circularity in how it is selected; a second invariance-family method confirms the blind spot belongs to the relation, not to any implementation; and cyclic relabeling delivers its predicted gain on a matched real sample. The same characterization explains a reported inversion of this ordering in the classifier metamorphic-testing literature: detectability is a joint property of the relation and the fault class, never of the relation alone.
Original Article
View Cached Full Text

Cached at: 08/07/26, 07:51 AM

# Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading
Source: [https://arxiv.org/abs/2608.05675](https://arxiv.org/abs/2608.05675)
[View PDF](https://arxiv.org/pdf/2608.05675)

> Abstract:Label\-free reliability for vision\-language models rests on invariance: perturb the input and a faithful reader's answer should not change\. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at\. We act on the complementary relation, equivariance: edit a figure's data and the correct answer must change by a computable amount\. Two matched edits are provably complete for affine reading errors; no suite of swap edits is complete for label permutations, and cyclic relabeling closes most of that gap\. We instantiate the theory as the Equivariance\-Consistency Score, a label\-free, training\-free detector, and release REND\-EQUIV, pairing matched invariance and equivariance sets over identical data\. The predicted ordering holds across three models and a hand\-labeled population immune to the one circularity in how it is selected; a second invariance\-family method confirms the blind spot belongs to the relation, not to any implementation; and cyclic relabeling delivers its predicted gain on a matched real sample\. The same characterization explains a reported inversion of this ordering in the classifier metamorphic\-testing literature: detectability is a joint property of the relation and the fault class, never of the relation alone\.

## Submission history

From: Rasul Khanbayov \[[view email](https://arxiv.org/show-email/0cfff27e/2608.05675)\] **\[v1\]**Thu, 6 Aug 2026 07:14:58 UTC \(266 KB\)

Similar Articles

Where Reliability Lives in Vision-Language Models: A Mechanistic Study of Attention, Hidden States, and Causal Circuits

arXiv cs.AI

This paper challenges the 'Attention-Confidence Assumption' by demonstrating that attention map sharpness is a poor predictor of correctness in Vision-Language Models. Instead, it shows that reliability is better indicated by hidden-state geometry and self-consistency, with significant findings on architectural differences between late-fusion and early-fusion models.

Vision-Language Models are Fragile Multilingual Associators

arXiv cs.CL

This paper introduces M2BIND, a benchmark to evaluate whether vision-language models maintain stable visual-linguistic associations across languages. It finds that binding is not language-invariant, with cross-family and cross-script settings causing significant performance collapse and weaker internal causal binding.