TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation
Summary
Introduces TVIR, a benchmark and hierarchical multi-agent framework for generating text-visual interleaved reports, evaluating factual reliability and visual alignment in automated report generation.
View Cached Full Text
Cached at: 06/02/26, 07:33 PM
Paper page - TVIR: Building Deep Research Agents Towards Text–Visual Interleaved Report Generation
Source: https://huggingface.co/papers/2606.02320 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
A multimodal deep research benchmark and agent framework are introduced to evaluate and improve the factual reliability and visual alignment of automated report generation systems.
Deep Research Agents have shown strong capability in multi-step information retrieval, reasoning, and long-form report generation, but existing benchmarks and systems remain predominantly text-centric, with limited evaluation of whethervisual elementsare factually reliable and well aligned with the surrounding analysis. To address this gap, we introduce TVIR (Text--Visual Interleaved Report Generation), which includes TVIR-Bench, a benchmark of 100 expert-curatedmultimodal deep researchtasks that requirevisual elementsto serve specific analytical sub-goals, and TVIR-Agent, ahierarchical multi-agent frameworkthat serves as a strong baseline for constructing outlines, retrieving images, generating charts with traceable sources, and composing reports through context-aware sequential writing. We further develop a dual-path evaluation framework that combinesTextual AssessmentandVisual Assessment. Experiments across nine deep research systems show that TVIR-Agent achieves strong overall performance, underscoring the importance of explicit multimodal design and evaluation forevidence-driven report generation.
View arXiv pageView PDFProject pageGitHub3Add to collection
Get this paper in your agent:
hf papers read 2606\.02320
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.02320 in a model README.md to link it from this page.
Datasets citing this paper1
#### NJU-LINK/TVIR-Bench Viewer• Updatedabout 5 hours ago • 100
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.02320 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation
This paper presents Ptah, a multi-agent harness for generating verifiable multimodal deep research reports by interleaving textual and visual evidence through specialized agents and verification mechanisms. It introduces PtahEval for evaluation.
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
The paper introduces Internalized Visual Thinking (IVT), a post-training framework that trains multimodal models to predict future frame embeddings, enabling direct answer generation without synthesizing intermediate images, thereby reducing latency over 5x for proactive video reasoning.
Can Agents Read the Room? Benchmarking Visual Social Intelligence in Multimodal Simulation
This paper introduces AgentViSS, a benchmark evaluating visual social intelligence in multimodal social simulation, containing 240 scenarios with aligned visual-textual evidence. Evaluating seven recent MLLMs reveals a gap between local role enactment and visually grounded interaction management.
VGI-BENCH: Probing Visual Intelligence in Video Generation Models
VGI-Bench evaluates visual reasoning in video generation models through 27 tasks, revealing limited reliability and self-correction in current systems.
Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation
The paper presents a locally deployed multi-agent AI system for structuring radiology reports and performing quality assurance, with radiologist evaluation showing favorable performance.