Brain-IT-VQA: From Brain Signals to Answers
Summary
Brain-IT-VQA framework decodes visual content from fMRI signals using transformer architecture, outperforming previous methods. The authors also introduce NSD-VQA, a new dataset with richer annotations for evaluating fMRI-based visual question answering.
View Cached Full Text
Cached at: 06/02/26, 03:24 AM
Paper page - Brain-IT-VQA: From Brain Signals to Answers
Source: https://huggingface.co/papers/2605.29588
Abstract
Brain-IT-VQA framework decodes visual content from fMRI signals using transformer-based architecture and introduces NSD-VQA dataset for improved visual question answering evaluation.
Decoding visual content fromfMRIsignals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While significant progress has been made in recent years invisual question answering(VQA) fromfMRI, performance remains limited. Moreover, although recent models can make increasingly accurate predictions, they have rarely been used as tools for understanding the structure ofvisual representationsin the brain. We presentBrain-IT-VQA, a framework forvisual question answeringfromfMRI. Building on the Brain InteractionTransformer(Brain-IT), our method decodeslanguage tokensfrombrain activityand integrates them with alanguage modelto answer visual questions. Our model substantially outperforms previousfMRI-based captioning and VQA approaches. We further introduceNSD-VQA, a new dataset and benchmark forvisual question answeringfromfMRI. Unlike existing image-fMRIVQA datasets, which typically provide only a few broad and weakly controlled questions per image,NSD-VQAprovides on average 20 question-answer pairs per image across 20 controlled question categories that disentangle multiple levels of visual understanding. This enables more reliable and interpretable evaluation despite limitedfMRItest data. Together,Brain-IT-VQA andNSD-VQAprovide both a strong predictive framework and a tool for studying brain representations. Using this benchmark, we quantify which forms of visual and semantic information can be reliably decoded fromfMRIresponses to natural images. We further analyze the contributions of differentbrain regionsacross question types.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.29588
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.29588 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.29588 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.29588 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Attention Consistent Longitudinal Medical Visual Question Answering Guided by Vision Foundation Models
Proposes an attention-guided encoder-decoder for longitudinal medical visual question answering, using a frozen DINO-based mask generator and auxiliary losses to improve consistency and interpretability, achieving strong results on the Medical-Diff-VQA benchmark.
SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
SuperMemory-VQA is a new egocentric VQA benchmark featuring 52.9 hours of AI-glasses footage and 4,853 QA pairs designed to evaluate AI assistants on long-horizon memory tasks spanning object recall, intent, timelines, and conversations. Benchmarking reveals existing agentic frameworks and LLMs remain far from reliable on these real-world memory challenges.
Evidence-Backed Video Question Answering
This paper introduces Evidence-Backed Video Question Answering (E-VQA), a new task requiring models to output both semantic answers and precise spatio-temporal evidence like tracked object segmentation masklets. The authors create a human-verified benchmark and a scalable training dataset, showing significant improvements over baselines.
ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering
ReVA introduces a region-aware visual assistant that enhances visually grounded question answering by integrating whole-image and region-level representations, reducing hallucinations in multimodal large language models.
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
The paper introduces Internalized Visual Thinking (IVT), a post-training framework that trains multimodal models to predict future frame embeddings, enabling direct answer generation without synthesizing intermediate images, thereby reducing latency over 5x for proactive video reasoning.