AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
Summary
AnswerMap introduces a black-box, training-free method for spatial interpretability of Vision-Language Models by constructing query-conditioned maps from answer posteriors, demonstrating faithfulness and utility across tasks.
View Cached Full Text
Cached at: 09/29/26, 04:11 PM
Paper page - AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
Source: https://huggingface.co/papers/2609.35247
Abstract
WhenaVLManswersavisualquery,currentinterpretabilitytoolsrelyontextrationales,whichuseamismatchedmodality,oroninternalread-outs,whichoriginatetooearlytoreflectthefinaloutputandrequirewhite-boxaccesstothemodel.WeintroduceAnswerMap,atraining-free,task-agnostic,black-boxvisualrationaleconstructedfromtheoutputhead.TheimageiscutintoKrowandKcolumnbands,eachshownalonetothefrozenmodelalongwiththequeryintheformatofayes/norelevancequestion.Theouterproductoftherowandcolumn``yes’’posteriorsgivesthequery-conditionedspatialmap.Crucially,bydefiningafixedread-outR(e.g.,expectation,maximum)ontopofAnswerMap,wecanderivecontinuousoutputslikelocationnatively.Thisbypassestherelianceondiscretetexttokensforcontinuous-outputtasksandguaranteesanimage-dependentanswerbyconstruction.However,arationalecanbeconfabulated,sowevalidateAnswerMapacrossfourmodelsandthreequerydistributionswithtwotests:(a)agreementwiththemodel’sowngeneratedpointand(b)deletionofthemap’sregion.Themaplandswherethemodelpoints(AUC0.85against0.38forattention),anddeletingitsregionflips53%ofcorrectanswers(against19%forattention’s).Beyondestablishingfaithfulness,wedemonstratethemap’stask-agnosticutilitythroughthreedistinctread-outs:itsmaximumflagshallucinatedobjectswithoutgeneration,itsexpectationlocalizescorrectlywhenthemodel’sownpointingfails,anditstop-massregion,fedbackasacrop,fixeshalfofthemodel’swronganswers.AnswerMapthusoffersanewlensonVLMinterpretabilityand,throughitsread-outs,anewoutputinterfaceforvisualtasksbeyondtexttokens.
View arXiv pageView PDFProject pageGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.35247
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35247 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.35247 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.35247 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
The paper introduces SpatialUncertain, a benchmark to evaluate whether vision-language models recognize when they cannot answer spatial questions due to occlusion or perspective ambiguity, revealing overconfidence and poor abstention behavior.
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
MetaSpatial is a reinforcement learning framework that enhances 3D spatial reasoning in vision-language models, enabling coherent and physically plausible 3D scene generation without hard-coded optimizations.
Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
This paper introduces Act2Answer, a protocol to evaluate knowledge retention in Vision-Language-Action (VLA) models by requiring agents to answer questions through physical actions. It finds that VLAs retain basic knowledge but show gaps on richer semantic categories, and that VQA co-training helps.
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
SpatialAct is a new simulator-grounded benchmark that probes whether VLM agents can perform coherent spatial reasoning and translate it into actions in 3D environments across multi-turn feedback settings. Experiments reveal a significant reasoning-to-action gap, with current VLMs struggling to maintain spatial beliefs and produce reliable actions despite performing well on isolated reasoning tasks.
Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models
This paper reveals that Vision-Language Models often rewrite rather than faithfully transcribe text when encountering perturbations like typos or visual artifacts, introducing the FaithC4 benchmark to evaluate this behavior across multiple models and languages.