Tag
AnswerMap introduces a black-box, training-free method for spatial interpretability of Vision-Language Models by constructing query-conditioned maps from answer posteriors, demonstrating faithfulness and utility across tasks.