AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

Hugging Face Daily Papers Papers

Summary

AnswerMap introduces a black-box, training-free method for spatial interpretability of Vision-Language Models by constructing query-conditioned maps from answer posteriors, demonstrating faithfulness and utility across tasks.

When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual rationale constructed from the output head. The image is cut into K row and K column bands, each shown alone to the frozen model along with the query in the format of a yes/no relevance question. The outer product of the row and column ``yes'' posteriors gives the query-conditioned spatial map. Crucially, by defining a fixed read-out R (e.g., expectation, maximum) on top of AnswerMap, we can derive continuous outputs like location natively. This bypasses the reliance on discrete text tokens for continuous-output tasks and guarantees an image-dependent answer by construction. However, a rationale can be confabulated, so we validate AnswerMap across four models and three query distributions with two tests: (a) agreement with the model's own generated point and (b) deletion of the map's region. The map lands where the model points (AUC 0.85 against 0.38 for attention), and deleting its region flips 53% of correct answers (against 19% for attention's). Beyond establishing faithfulness, we demonstrate the map's task-agnostic utility through three distinct read-outs: its maximum flags hallucinated objects without generation, its expectation localizes correctly when the model's own pointing fails, and its top-mass region, fed back as a crop, fixes half of the model's wrong answers. AnswerMap thus offers a new lens on VLM interpretability and, through its read-outs, a new output interface for visual tasks beyond text tokens.
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:11 PM

Paper page - AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

Source: https://huggingface.co/papers/2609.35247

Abstract

WhenaVLManswersavisualquery,currentinterpretabilitytoolsrelyontextrationales,whichuseamismatchedmodality,oroninternalread-outs,whichoriginatetooearlytoreflectthefinaloutputandrequirewhite-boxaccesstothemodel.WeintroduceAnswerMap,atraining-free,task-agnostic,black-boxvisualrationaleconstructedfromtheoutputhead.TheimageiscutintoKrowandKcolumnbands,eachshownalonetothefrozenmodelalongwiththequeryintheformatofayes/norelevancequestion.Theouterproductoftherowandcolumn``yes’’posteriorsgivesthequery-conditionedspatialmap.Crucially,bydefiningafixedread-outR(e.g.,expectation,maximum)ontopofAnswerMap,wecanderivecontinuousoutputslikelocationnatively.Thisbypassestherelianceondiscretetexttokensforcontinuous-outputtasksandguaranteesanimage-dependentanswerbyconstruction.However,arationalecanbeconfabulated,sowevalidateAnswerMapacrossfourmodelsandthreequerydistributionswithtwotests:(a)agreementwiththemodel’sowngeneratedpointand(b)deletionofthemap’sregion.Themaplandswherethemodelpoints(AUC0.85against0.38forattention),anddeletingitsregionflips53%ofcorrectanswers(against19%forattention’s).Beyondestablishingfaithfulness,wedemonstratethemap’stask-agnosticutilitythroughthreedistinctread-outs:itsmaximumflagshallucinatedobjectswithoutgeneration,itsexpectationlocalizescorrectlywhenthemodel’sownpointingfails,anditstop-massregion,fedbackasacrop,fixeshalfofthemodel’swronganswers.AnswerMapthusoffersanewlensonVLMinterpretabilityand,throughitsread-outs,anewoutputinterfaceforvisualtasksbeyondtexttokens.

View arXiv pageView PDFProject pageGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.35247

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.35247 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.35247 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.35247 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

Hugging Face Daily Papers

SpatialAct is a new simulator-grounded benchmark that probes whether VLM agents can perform coherent spatial reasoning and translate it into actions in 3D environments across multi-turn feedback settings. Experiments reveal a significant reasoning-to-action gap, with current VLMs struggling to maintain spatial beliefs and produce reliable actions despite performing well on isolated reasoning tasks.