Document Classification Pattern Recognition via Information Fusion: A Systematic Review of Multimodal and Multiview Representation Approaches
Summary
This systematic review of 139 studies proposes a unified framework and meta-analysis for document classification via multimodal and multiview information fusion, finding that fusion improves accuracy (mean gain of +5.28 percentage points) but highlights reproducibility challenges.
View Cached Full Text
Cached at: 05/26/26, 08:58 AM
# Document Classification Pattern Recognition via Information Fusion: A Systematic Review of Multimodal and Multiview Representation Approaches Source: [https://arxiv.org/abs/2605.23910](https://arxiv.org/abs/2605.23910) [View PDF](https://arxiv.org/pdf/2605.23910) > Abstract:Information fusion is used widely to improve document classification by the integration of multiple data sources \(multimodal\) or representations \(multiview\)\. However, the field lacks a unified framework, a quantitative synthesis of its effectiveness, and clear guidance for practitioners\. This systematic review addresses these gaps by analysing 139 primary studies\. It introduces a formal framework to structure the field, presents the results of a qualitative analysis to identify key trends, and performs a random\-effects meta\-analysis \(to our knowledge, the first focused on document classification\) to quantify performance gains\. Our meta\-analysis reveals that multimodal fusion improves accuracy \(mean gain of \+5\.28 percentage points, $p=0\.0016$\) significantly \-\- the F1\-score effect is directionally positive but statistically non\-significant in our primary model\. Multiview fusion provides consistent but modest gains for accuracy \(\+4\.67\\%\), F1\-score \(\+3\.08\\%\), and recall \(all $p<0\.05$\)\. Critically, our qualitative synthesis uncovers challenges in reproducibility in methodological rigour: only 11\.8\\% \(multimodal\) and 23\.3\\% \(multiview\) of the studies use statistical tests to validate their findings, which undermines the reliability of many of their results\. This review's primary contributions are a unifying framework, the first quantitative evidence base, and data\-driven guidelines\. This review concludes that successful information fusion depends not on algorithmic complexity, but on the strategic alignment of the fusion method with the task context and a commitment to more rigorous validation\. ## Submission history From: Marcin Mirończuk Mirończuk \[[view email](https://arxiv.org/show-email/2797ee5a/2605.23910)\] **\[v1\]**Tue, 7 Apr 2026 08:49:20 UTC \(2,208 KB\)
Similar Articles
Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis
A research paper proposing SeRIn, a multimodal fusion scheme that separates modality-specific refinement from cross-modal interaction, achieving state-of-the-art on CH-SIMS and CMU-MOSEI benchmarks for sentiment analysis.
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion is a vision-centric multimodal large language model for holistic medical understanding that unifies 2D and 3D medical image analysis using a cascaded vision encoder. It achieves state-of-the-art results on 20 out of 24 benchmarks and outperforms proprietary models like GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.
Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion
The paper identifies strong-modality collapse in multimodal learning where fusion degrades the dominant modality's performance, and proposes Inverted Asymmetric Fusion (IAF) to preserve it, improving over unimodal baselines.
CL-DMDF:Dynamic Multimodal Data Fusion Model Based on Contrastive Learning
This paper proposes CL-DMDF, a dynamic multimodal data fusion model that uses contrastive learning and a dual-dimensional attention mechanism to handle missing modalities and improve discriminative learning.
Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
This paper proposes a token-centric dual-view learning framework that unifies prompt-based adaptation and cross-view fusion within a frozen vision transformer to improve breast cancer classification from mammography images, achieving consistent improvements on VinDr-Mammo and CMMD datasets.