Document Classification Pattern Recognition via Information Fusion: A Systematic Review of Multimodal and Multiview Representation Approaches

arXiv cs.CL Papers

Summary

This systematic review of 139 studies proposes a unified framework and meta-analysis for document classification via multimodal and multiview information fusion, finding that fusion improves accuracy (mean gain of +5.28 percentage points) but highlights reproducibility challenges.

arXiv:2605.23910v1 Announce Type: new Abstract: Information fusion is used widely to improve document classification by the integration of multiple data sources (multimodal) or representations (multiview). However, the field lacks a unified framework, a quantitative synthesis of its effectiveness, and clear guidance for practitioners. This systematic review addresses these gaps by analysing 139 primary studies. It introduces a formal framework to structure the field, presents the results of a qualitative analysis to identify key trends, and performs a random-effects meta-analysis (to our knowledge, the first focused on document classification) to quantify performance gains. Our meta-analysis reveals that multimodal fusion improves accuracy (mean gain of +5.28 percentage points, $p=0.0016$) significantly -- the F1-score effect is directionally positive but statistically non-significant in our primary model. Multiview fusion provides consistent but modest gains for accuracy (+4.67\%), F1-score (+3.08\%), and recall (all $p<0.05$). Critically, our qualitative synthesis uncovers challenges in reproducibility in methodological rigour: only 11.8\% (multimodal) and 23.3\% (multiview) of the studies use statistical tests to validate their findings, which undermines the reliability of many of their results. This review's primary contributions are a unifying framework, the first quantitative evidence base, and data-driven guidelines. This review concludes that successful information fusion depends not on algorithmic complexity, but on the strategic alignment of the fusion method with the task context and a commitment to more rigorous validation.
Original Article
View Cached Full Text

Cached at: 05/26/26, 08:58 AM

# Document Classification Pattern Recognition via Information Fusion: A Systematic Review of Multimodal and Multiview Representation Approaches
Source: [https://arxiv.org/abs/2605.23910](https://arxiv.org/abs/2605.23910)
[View PDF](https://arxiv.org/pdf/2605.23910)

> Abstract:Information fusion is used widely to improve document classification by the integration of multiple data sources \(multimodal\) or representations \(multiview\)\. However, the field lacks a unified framework, a quantitative synthesis of its effectiveness, and clear guidance for practitioners\. This systematic review addresses these gaps by analysing 139 primary studies\. It introduces a formal framework to structure the field, presents the results of a qualitative analysis to identify key trends, and performs a random\-effects meta\-analysis \(to our knowledge, the first focused on document classification\) to quantify performance gains\. Our meta\-analysis reveals that multimodal fusion improves accuracy \(mean gain of \+5\.28 percentage points, $p=0\.0016$\) significantly \-\- the F1\-score effect is directionally positive but statistically non\-significant in our primary model\. Multiview fusion provides consistent but modest gains for accuracy \(\+4\.67\\%\), F1\-score \(\+3\.08\\%\), and recall \(all $p<0\.05$\)\. Critically, our qualitative synthesis uncovers challenges in reproducibility in methodological rigour: only 11\.8\\% \(multimodal\) and 23\.3\\% \(multiview\) of the studies use statistical tests to validate their findings, which undermines the reliability of many of their results\. This review's primary contributions are a unifying framework, the first quantitative evidence base, and data\-driven guidelines\. This review concludes that successful information fusion depends not on algorithmic complexity, but on the strategic alignment of the fusion method with the task context and a commitment to more rigorous validation\.

## Submission history

From: Marcin Mirończuk Mirończuk \[[view email](https://arxiv.org/show-email/2797ee5a/2605.23910)\] **\[v1\]**Tue, 7 Apr 2026 08:49:20 UTC \(2,208 KB\)

Similar Articles

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Hugging Face Daily Papers

ClinFusion is a vision-centric multimodal large language model for holistic medical understanding that unifies 2D and 3D medical image analysis using a cascaded vision encoder. It achieves state-of-the-art results on 20 out of 24 benchmarks and outperforms proprietary models like GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.