F2IND-IT! -- Multimodal Fuzzy Fake Indian News Detection using Images and Text
Summary
Proposes FIND-IT!, a multimodal fake news detection framework for Indian news using ResNet-50 for visual features, DistilBERT for text, and ANFIS with attention fusion to classify news as fake or real.
View Cached Full Text
Cached at: 05/19/26, 06:39 AM
# IT! - Multimodal Fuzzy Fake Indian News Detection using Images and Text
Source: [https://arxiv.org/html/2605.17115](https://arxiv.org/html/2605.17115)
11institutetext:ABV \- Indian Institute of Information Technology, Gwalior
11email:kushal\.trivedi\.2110@gmail\.com###### Abstract
Newspapers remain a vital source of journalism, delivering updates on current events, politics, business, sports, and entertainment\. However, in a country as vast and diverse as India, partial or biased manipulation of facts is common, especially when the same news is covered by multiple regional and national outlets\. While several existing approaches focus on integrating textual and visual features for fake news detection, very few have examined their effectiveness on Indian news content\. This research presents a novel multimodal framework — FIND\-IT\! \(Fuzzy Fake Indian News Detection using Images and Text\) — that combines visual and textual modalities for enhanced fake news detection on Indian media\. The proposed model utilizes Convolutional Neural Networks \(ResNet\-50\) to extract visual features from news images, a text encoder \(DistilBERT\) to obtain textual semantic embeddings and an Adaptive Neuro\-Fuzzy Inference System \(ANFIS\) to generate a fuzzy reliability score\. A lightweight attention\-based fusion module is employed to assign learnable weights to each modality before classification into fake or real\. The study is completed by a formal and in\-depth analysis and exploration of this new architecture to the IFND dataset with comparison to previous research by comparing accuracy, precision, recall and F1 scores, to support a brief discussion on the model’s performance\.
## 1Introduction
With the advancement, awareness, and wider use of technology in recent years, the number of news articles being shared has increased sharply\. Earlier, daily news was mostly available through newspapers and reached only a small part of the population\. According to a report by the Indian Ministry of Communication in March 2024, 95\.15% of villages in India now have 3G or 4G mobile internet access\[[12](https://arxiv.org/html/2605.17115#bib.bib9)\]\. In addition, reports by IAMAI \(2024\) predict that by 2025, about 56% of all new internet users in India will come from rural areas\[[2](https://arxiv.org/html/2605.17115#bib.bib11)\]\.
However, as news has become more accessible to more people, the amount of fake news has also grown\. While there is no official definition of "fake news," it is commonly described asany content that is deliberately made and known to be false\. Fake news often uses emotional language and special writing styles\. These are often captured through features such as tone and writing patterns\. Common ways to spread fake news include editing images, changing topics to mislead readers, and using clickbait to attract attention\[[15](https://arxiv.org/html/2605.17115#bib.bib12)\]\.
According to official data from the Press Information Bureau under the Ministry of Information and Broadcasting, 1,575 fake news cases were reported between 2022 and March 2025\. The number rose from 338 in 2022 to 583 in 2024\[[4](https://arxiv.org/html/2605.17115#bib.bib14)\]\. Data from the National Crime Records Bureau also shows a 214% increase in fake news cases during the early pandemic period from 2018 to 2020\[[6](https://arxiv.org/html/2605.17115#bib.bib15)\]\. A 2024 study by ISB and CyberPeace found that 46% of false information was about politics, and over 77% of it spread through social media platforms\[[11](https://arxiv.org/html/2605.17115#bib.bib17)\]\. Another survey among Gen Z users in Delhi found that 91% believe fake news can affect election outcomes\[[1](https://arxiv.org/html/2605.17115#bib.bib16)\]\.
Current methods for detecting fake news automatically are usually grouped into three types: modality\-based methods, propagation\-based methods, and fact\-based methods\[[14](https://arxiv.org/html/2605.17115#bib.bib10)\]\. Modality\-based methods look at the content of the news itself\. This includes text features such as writing style or word use, image features such as signs of editing, or both text and image combined \(multimodal\)\. Propagation\-based methods study how news spreads across social media and other online platforms\. Fact\-based methods try to check the news content against trusted sources or known facts\.
The evolution from single\-modal to multi\-modal fake news classifiers has become essential to prevent underfitting models trained on only one type of content modality, such as either textual or visual data\. In many cases, the information conveyed through images may contradict the textual content, or vice versa, which can lead to misleading interpretations and ultimately contribute to biased news diffusion\. Multi\-modal approaches aim to capture complementary features from both text and image modalities, enabling more robust and accurate detection of fake news\.
The World Economic Forum’s 2024 Global Risk Report placed India at highest risk for misinformation globally, with experts citing high levels of political polarization and algorithmic amplification\[[16](https://arxiv.org/html/2605.17115#bib.bib13)\]\. Manual detection of fake news is labor\-intensive, time\-consuming, and prone to bias too\. Thus, there is a significant gap in the creation of a credible database centered around Indian news articles, as well as in research focused on developing tools for robust automated classification of fake news\. This gap serves as the motivation for introducing FIND\-IT, a fuzzy\-based multimodal deep learning architecture for fake news detection\.
## 2Prior Art
### 2\.1Baseline Deep Learning Approaches for Multi\-Modal Data
MAGIC\[[8](https://arxiv.org/html/2605.17115#bib.bib2)\], IFND\[[13](https://arxiv.org/html/2605.17115#bib.bib1)\], Tri‑FusionNet\[[3](https://arxiv.org/html/2605.17115#bib.bib3)\], BDANN\[[19](https://arxiv.org/html/2605.17115#bib.bib4)\], CLIP‑based learning\[[20](https://arxiv.org/html/2605.17115#bib.bib5)\], Cross‑Attention Networks\[[18](https://arxiv.org/html/2605.17115#bib.bib6)\], and ETMA\[[17](https://arxiv.org/html/2605.17115#bib.bib7)\]are among the top\-performing frameworks in multimodal fake‑news detection\. Table[1](https://arxiv.org/html/2605.17115#S2.T1)summarizes their accuracy, F1 scores, and datasets used\.
Table 1:Summary of top\-performing deep learning frameworks for multimodal fake news detection
### 2\.2Fuzzy\-based Deep Learning Approaches for Multi\-Modal Data
To the best of our knowledge,\[[5](https://arxiv.org/html/2605.17115#bib.bib8)\]is the only work that incorporates fuzzy logic with neural networks for fake news classification\. The results from this study are summarized in Table[2](https://arxiv.org/html/2605.17115#S2.T2)\.
Table 2:Summary of performance of the neuro\-fuzzy model\.
## 3Proposed Methodology
In this section, we present the methodology of the proposed framework, including the dataset used, various CNN architectures, text encoders, and the experimental setup\.
### 3\.1Dataset Used
In this study, we consider IFND \(Indian Fake News Dataset\)\. The IFND \(Indian Fake News Dataset\) is a multimodal dataset containing image\-text pairs extracted from Indian news articles\. It comprises 56,713 news articles, covering international, national, and local events from the period between 2013 and 2021\. This dataset has been used in our study to classify fake news from real news\. The articles are categorized into five topics—Election, Politics, COVID\-19, Violence, and Miscellaneous\.
### 3\.2FIND\-IT Architecture
In this subsection, we discuss the overall architecture of the proposed FIND\-IT model \(illustrated in Fig\.[1](https://arxiv.org/html/2605.17115#S3.F1)\)\.
#### 3\.2\.1Overall Flow of Data
This model uses DistilBERT and ResNet\-50 to extract textual and visual features, respectively, projecting them into high\-dimensional embeddings\. A lightweight attention gating mechanism fuses the modalities, with embeddings resized via MLPs for dimensional alignment\. The attention module adaptively balances each modality’s contribution\. The fused features are then passed through an ANFIS layer with 2 Gaussian membership functions to perform binary fake\-news classification\.
#### 3\.2\.2Visual Feature Extractor \(CNN\)
We utilize a ResNet\-50\-based visual encoder to extract high\-level image features\. Specifically, we load the pretrained ResNet\-50 model and remove its final classification layer, retaining only the convolutional backbone\. Formally, given an input imageII, the encoder maps it to a fixed\-size feature vector:
v=ResNet\(I\)∈ℝ2048,v=\\text\{ResNet\}\(I\)\\in\\mathbb\{R\}^\{2048\},wherevvrepresents the output of the global average pooling layer\. All parameters of the ResNet backbone are fine\-tuned during training to better align with the target task\.
#### 3\.2\.3Text Encoder \(DistilBERT\)
For a piece of news, we use theDistilBert\- Tokenizerto tokenize its contents, adding the classification token\[CLS\]at the beginning and the separation token\[SEP\]at the end of the token sequence\. The resulting input takes the form:
X=\[\[CLS\],x1,…,xn,\[SEP\]\],X=\[\\texttt\{\[CLS\]\},x\_\{1\},\\ldots,x\_\{n\},\\texttt\{\[SEP\]\}\],wherennis the number of original tokens\. These tokens are then fed into DistilBERT, which maps them into a contextualized low\-dimensional embedding space:
W=DistilBERT\(X\)∈ℝN×d,W=\\text\{DistilBERT\}\(X\)\\in\\mathbb\{R\}^\{N\\times d\},whered=768d=768is the hidden size of thedistilbert\-base\-uncasedmodel\. Unlike the original BERT, DistilBERT removes the token\-type embeddings and the second segment input, offering a more lightweight and faster alternative while retaining 95% of BERT’s language understanding capabilities\.
In our implementation, the output of the DistilBERT encoder is a tensor of shape\(B,S,768\)\(B,S,768\), whereBBis the batch size andSSis the sequence length\. To obtain a fixed\-size sentence representation, we apply a mean pooling operation over the token embeddings, weighted by the attention mask\. Specifically,
𝐰mean=∑i=1N𝐰i⋅mi∑i=1Nmi,\\mathbf\{w\}\_\{\\text\{mean\}\}=\\frac\{\\sum\_\{i=1\}^\{N\}\\mathbf\{w\}\_\{i\}\\cdot m\_\{i\}\}\{\\sum\_\{i=1\}^\{N\}m\_\{i\}\},
Figure 1:The framework of the F2IND Architecture\.where𝐰i\\mathbf\{w\}\_\{i\}is the hidden representation of theii\-th token andmim\_\{i\}is the corresponding attention mask\. This results in a single vector per input sequence of shape\(B,768\)\(B,768\), which serves as the final sentence embedding\.
#### 3\.2\.4Attention\-Based Fusion Module
The shapes of tensors from the ResNet\-50 module and DistilBERT encoder are established as X=2048 and Y=768, respectively\. Before feeding embeddings into the ANFIS module for the fuzzy inference implementation of the model, both embeddings are projected to a common dimensional space of size 512\. After projection, we stack them along the modality dimension, resulting in a combined tensor of shape\(B,2,512\)\(B,2,512\), whereBBdenotes the batch size\.
To compute attention logits for each modality, we apply an MLP that projects each modality\-specific embedding to a scalar value, producing attention scores of shape\(B,2\)\(B,2\)\. These scores are then renormalized \(via softmax\) to ensure they sum to 1 across modalities\. The embeddings are then aggregated and reshaped back to a unified representation of shape\(B,512\)\(B,512\), which is subsequently used for the final binary classification task\. These steps can be represented mathematically as:
x∈ℝB×2048,y∈ℝB×768\\displaystyle x\\in\\mathbb\{R\}^\{B\\times 2048\},\\quad y\\in\\mathbb\{R\}^\{B\\times 768\}\(1\)x^=Wxx∈ℝB×512,y^=Wyy∈ℝB×512\\displaystyle\\hat\{x\}=W\_\{x\}x\\in\\mathbb\{R\}^\{B\\times 512\},\\quad\\hat\{y\}=W\_\{y\}y\\in\\mathbb\{R\}^\{B\\times 512\}z=\[x^;y^\]∈ℝB×2×512\\displaystyle z=\[\\hat\{x\};\\hat\{y\}\]\\in\\mathbb\{R\}^\{B\\times 2\\times 512\}a=softmax\(MLP\(z\)\)∈ℝB×2\\displaystyle a=\\text\{softmax\}\(\\text\{MLP\}\(z\)\)\\in\\mathbb\{R\}^\{B\\times 2\}h=∑i=12ai⋅zi∈ℝB×512\\displaystyle h=\\sum\_\{i=1\}^\{2\}a\_\{i\}\\cdot z\_\{i\}\\in\\mathbb\{R\}^\{B\\times 512\}
#### 3\.2\.5Fuzzy Logic Inference \(ANFIS\)
1. 1\.Input Layer:The input to ANFIS consists of batches of 4\-dimensional vectors, i\.e\., of shape\(B,n\)\(B,n\)wheren=4n=4\. LetX=\{x1,x2,x3,x4\}X=\\\{x\_\{1\},x\_\{2\},x\_\{3\},x\_\{4\}\\\}represent the input features\.
2. 2\.Fuzzification Layer:Each input value is fuzzified using two Gaussian membership functions\. The mean \(μj\\mu\_\{j\}\) and standard deviation \(σj\\sigma\_\{j\}\) of each membership function are learnable parameters\. For every featurexix\_\{i\}in the inputXX, the degree of membership to each fuzzy set is computed using the Gaussian function: G\(xi;μj,σj\)=exp\(−\(xi−μj\)22σj2\),G\(x\_\{i\};\\mu\_\{j\},\\sigma\_\{j\}\)=\\exp\\left\(\-\\frac\{\(x\_\{i\}\-\\mu\_\{j\}\)^\{2\}\}\{2\\sigma\_\{j\}^\{2\}\}\\right\),wherei∈\[1,n\]i\\in\[1,n\]andj∈\[1,f\]j\\in\[1,f\], withn=4n=4andf=2f=2\(number of membership functions\)\. Thus, for each inputXX,n×f=4×2=8n\\times f=4\\times 2=8membership values are computed\. The output of this layer is of shape\(B,n,f\)\(B,n,f\)\.
3. 3\.Rule Layer:All possible fuzzy rules are evaluated using the product \(AND\) of the membership values across features\. The total number of fuzzy rules isfn=24=16f^\{n\}=2^\{4\}=16\. The firing strengthfkf\_\{k\}of thekk\-th rule is computed as: fk=∏i=1nG\(xi;μj,σj\),f\_\{k\}=\\prod\_\{i=1\}^\{n\}G\(x\_\{i\};\\mu\_\{j\},\\sigma\_\{j\}\),wherek∈\[1,fn\]k\\in\[1,f^\{n\}\]\. The firing strengths are then normalized: f^k=fk∑i=1fnfi\.\\hat\{f\}\_\{k\}=\\frac\{f\_\{k\}\}\{\\sum\_\{i=1\}^\{f^\{n\}\}f\_\{i\}\}\.The output of this layer is of shape\(B,fn\)\(B,f^\{n\}\)\.
4. 4\.Rule Weighting Layer:Each rule contributes a weighted output, computed as: zk=∑i=1naikxi\+bk,z\_\{k\}=\\sum\_\{i=1\}^\{n\}a\_\{ik\}x\_\{i\}\+b\_\{k\},whereaika\_\{ik\}andbkb\_\{k\}are trainable parameters for each rule and input feature\. The output of this layer is also of shape\(B,fn\)\(B,f^\{n\}\)\.
5. 5\.Output Layer:The final output is a weighted sum of the normalized firing strengths and the rule outputs: z=∑k=1fnf^k⋅zk\.z=\\sum\_\{k=1\}^\{f^\{n\}\}\\hat\{f\}\_\{k\}\\cdot z\_\{k\}\. A sigmoid activation is applied to produce a confidence score between 0 and 1: Output=σ\(z\)=11\+e−z\.\\text\{Output\}=\\sigma\(z\)=\\frac\{1\}\{1\+e^\{\-z\}\}\. The final output is of shape\(B,1\)\(B,1\), representing the fake news probability for each input in the batch\.
### 3\.3Evaluation Metric
F1 score metric is a comprehensive evaluation method of precision and recall, we take it as the metric in evaluating our approach and baselines\. Recall, Precision, and F1 equations are shown as follows:
F1=2⋅Precision⋅RecallPrecision\+RecallF1=\\frac\{2\\cdot\\text\{Precision\}\\cdot\\text\{Recall\}\}\{\\text\{Precision\}\+\\text\{Recall\}\}
### 3\.4Experiment Setting
The details of the experimental setup of our approach are discussed below:
1. 1\.Data Imbalance and Preprocessing:The dataset consists of 56,713 news article text\-image pairs\. However, due to missing image links, a significant number of samples were removed for image preprocessing\. All images with a resolution higher than224×224224\\times 224were resized to224×224224\\times 224and normalized using the ImageNet dataset statistics\. This resulted in a final dataset of25,19525,195examples, comprising24,57624,576true news articles and619619fake news articles\. Dynamic padding is also used for text batch preprocessing\.
2. 2\.DistilBERT:A dropout rate of 0\.30 was applied, and mean pooling was used to obtain fixed\-size sentence embeddings\.
3. 3\.ResNet\-50:The final classification layer was removed, and all model parameters were kept trainable to enable fine\-tuning\.
4. 4\.Attention Fusion:A modality\-level attention mechanism was implemented using a lightweight MLP and softmax normalization to compute attention scores between modalities\. Bit\-masking is applied to examples where images are unavailable, ensuring that the attention fusion mechanism allocates complete attention to the text, effectively setting the image’s weight to 0\.
5. 5\.ANFIS:A Takagi–Sugeno–style Adaptive Neuro\-Fuzzy Inference System \(ANFIS\) is employed utilizing44inputs and22Gaussian membership functions\.
6. 6\.Loss, Learning Rate, and Optimizer:A custom loss function was designed as a weighted combination of binary cross\-entropy loss, Huber loss \(to penalize incorrect minority class predictions\), and focal loss \(to address class imbalance\)\. The learning rate was dynamically adjusted using the OneCycleLR scheduler, with different scales assigned to different model components\. The Adam optimizer was used for all updates\.
Furthermore, the model was trained using a stratified 5\-fold cross\-validation strategy for 5 epochs, with a batch size of 16\. The number of parameters in the model are roughly 91\.3 million\.
## 4Results and Discussion
Table 3:Performance of theF2INDmodel on the IFND datasetModelDatasetAcc\.Macro\-F1Prec\.RecallROCPRFakeTrueAUCAUCF2INDIFND0\.97690\.97350\.97890\.98010\.98540\.99590\.99771. 1\.Evaluation Metrics:To counter class imbalance, Macro F1 is used over Micro F1\. This ensures equal weighting across all classes and prevents majority\-class bias\.
2. 2\.Validation:K\-Fold Cross\-Validation was implemented to ensure robust evaluation and improve generalization by rotating training/validation sets across all data points\.
3. 3\.Fuzzy Membership:Gaussian membership functions exhibit moderate overlap for smooth transitions and distinct centers, effectively capturing non\-linear trends\.
4. 4\.Firing Strength and Rule Contributions:Most of the 16 rules show uniform normalized firing strengths\. This indicates a balanced system architecture\. The model’s output is driven by a subset of rules with both positive and negative contributions, allowing for precise pattern discrimination and prediction\.
## 5Ablation Studies
The following ablation studies \(Table[4](https://arxiv.org/html/2605.17115#S5.T4)\) were conducted in addition to our primary research to evaluate performance improvements of the proposed model\. Further ablation studies can also be conducted by varying the number and nature of membership functions, as well as the ANFIS architecture itself\.
CategoryModel/ConfigurationAccuracy \(%\)CNN \(Unimodal\)ResNet\-5076\.60VGG\-1665\.30Text Encoder \(Unimodal\)LSTM92\.60Bi\-LSTM92\.70CNN \(Multimodal\)ResNet\-5097\.69VGG\-1997\.85Text Encoder \(Multimodal\)LSTM94\.17Bi\-LSTM95\.67DistilBERT97\.69BERT97\.71ANFIS InfluenceWithout ANFIS96\.73With ANFIS97\.69Table 4:Ablation study comparing architectures and configurations\.
## 6Conclusion and Future Work
This research presents a novel study on the detection of fake news published by Indian newspapers and proposes a new architecture that combines neural networks and fuzzy logic to classify news as either fake or real\. Experimental results on the real\-world, comprehensive IFND dataset demonstrate the effectiveness of the proposed model\. Ablation studies were conducted to explore alternative model architectures\. While some variations performed marginally close to our proposed architecture, ours consistently outperformed them across all evaluation metrics\.
Potential future enhancements to this study include replacing ANFIS’s reliance on prior expert knowledge for forming input\-output fuzzy partitions and designing the fuzzy rule base with data\-driven models that automatically identify the centroids and spreads of fuzzy clusters during training\. In these models, rules are formed dynamically during training and do not require expert intervention\[[10](https://arxiv.org/html/2605.17115#bib.bib18)\]\[[9](https://arxiv.org/html/2605.17115#bib.bib19)\]\[[7](https://arxiv.org/html/2605.17115#bib.bib20)\]\.
\{credits\}
#### 6\.0\.1\\discintname
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\. Pre\-processed dataset and code for the model can be made available upon request\.
## References
- \[1\]BrandEquity Bureau\(2024\-03\)91 per cent believe fake news can influence voting decisions: report\(Website\)Note:Accessed: 2025\-07\-17External Links:[Link](https://brandequity.economictimes.indiatimes.com/news/research/91-per-cent-believe-fake-news-can-influence-voting-decisions-report/110004307?utm_source=chatgpt.com)Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p3.1)\.
- \[2\]Communications Today\(2024\)By 2025, 56 percent new indian internet users from rural areas\(Website\)Note:Accessed: 2025\-07\-20External Links:[Link](https://www.communicationstoday.co.in/by-2025-56-percent-new-indian-internet-users-from-rural-areas/?utm_source=chatgpt.com)Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p1.1)\.
- \[3\]S\. El\-Amrany, M\. R\. Brust, J\. E\. Pecero, and P\. Bouvry\(2024\)Tri\-fusiondet: leveraging user engagement, textual, and visual features for enhanced fake news detection\.In2024 28th International Computer Science and Engineering Conference \(ICSEC\),Vol\.,pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1109/ICSEC62781.2024.10770746)Cited by:[§2\.1](https://arxiv.org/html/2605.17115#S2.SS1.p1.1)\.
- \[4\]Exchange4Media\(2025\)Misinformation queries raised, 1,575 fake news cases flagged since 2022: pib\.Note:Accessed: 2025\-07\-20External Links:[Link](https://www.exchange4media.com/media-others-news/68914-misinformation-queries-raised-1575-fake-news-cases-flagged-since-2022-pib-142128.html)Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p3.1)\.
- \[5\]T\. M\. H\. Gedara, V\. Loia, and S\. Tomasiello\(2025\)A fuzzy\-based multimodal approach for interpretable fake news detection\.Applied Soft Computing179,pp\. 113277\.External Links:ISSN 1568\-4946,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.asoc.2025.113277),[Link](https://www.sciencedirect.com/science/article/pii/S1568494625005885)Cited by:[§2\.2](https://arxiv.org/html/2605.17115#S2.SS2.p1.1)\.
- \[6\]Indian Express\(2021\)214% rise in cases relating to fake news, rumours\.Note:Accessed: 2025\-07\-20External Links:[Link](https://indianexpress.com/article/india/214-rise-in-cases-relating-to-fake-news-rumours-7511534/)Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p3.1)\.
- \[7\]A\. R\. Iyer, D\. K\. Prasad, and C\. H\. Quek\(2018\)PIE\-rspop: a brain\-inspired pseudo\-incremental ensemble rough set pseudo\-outer product fuzzy neural network\.Expert Systems with Applications95,pp\. 172–189\.External Links:ISSN 0957\-4174,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.eswa.2017.11.027),[Link](https://www.sciencedirect.com/science/article/pii/S0957417417307832)Cited by:[§6](https://arxiv.org/html/2605.17115#S6.p2.1)\.
- \[8\]Jun\-hao and Xu\(2024\)A multimodal adaptive graph\-based intelligent classification model for fake news\.External Links:2411\.06097,[Link](https://arxiv.org/abs/2411.06097)Cited by:[§2\.1](https://arxiv.org/html/2605.17115#S2.SS1.p1.1)\.
- \[9\]N\. Kasabov\(2001\)Evolving fuzzy neural networks for supervised/unsupervised online knowledge\-based learning\.IEEE Transactions on Systems, Man, and Cybernetics, Part B \(Cybernetics\)31\(6\),pp\. 902–918\.External Links:[Document](https://dx.doi.org/10.1109/3477.969494)Cited by:[§6](https://arxiv.org/html/2605.17115#S6.p2.1)\.
- \[10\]N\.K\. Kasabov and Q\. Song\(2002\)DENFIS: dynamic evolving neural\-fuzzy inference system and its application for time\-series prediction\.IEEE Transactions on Fuzzy Systems10\(2\),pp\. 144–154\.External Links:[Document](https://dx.doi.org/10.1109/91.995117)Cited by:[§6](https://arxiv.org/html/2605.17115#S6.p2.1)\.
- \[11\]NDTV\(2024\)Nearly half of the fake news stories in india are political: study\.Note:Accessed: 2025\-07\-20External Links:[Link](https://www.ndtv.com/india-news/nearly-half-of-the-fake-news-stories-in-india-are-political-study-7291481)Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p3.1)\.
- \[12\]Press Information Bureau\(2025\-07\-19\)PM inaugurates, dedicates to the nation and lays the foundation stone for multiple development projects worth about rs 13,000 crores in varanasi\.Note:Press release[https://www\.pib\.gov\.in/PressReleasePage\.aspx?PRID=2040566](https://www.pib.gov.in/PressReleasePage.aspx?PRID=2040566)External Links:[Link](https://www.pib.gov.in/PressReleasePage.aspx?PRID=2040566)Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p1.1)\.
- \[13\]D\. K\. Sharma and S\. Garg\(2023\)IFND: a benchmark dataset for fake news detection\.Complex & Intelligent Systems9\(3\),pp\. 2843–2863\.External Links:[Document](https://dx.doi.org/10.1007/s40747-021-00552-1),[Link](https://doi.org/10.1007/s40747-021-00552-1)Cited by:[§2\.1](https://arxiv.org/html/2605.17115#S2.SS1.p1.1)\.
- \[14\]K\. Tian, G\. Rao, X\. Wang, M\. Yu, J\. Zhang, and L\. Zhang\(2025\)CMFNThinker: a novel cross\-source multi\-modal fake news detection model\.InICASSP 2025 \- 2025 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 1–5\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP49660.2025.10889602)Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p4.1)\.
- \[15\]S\. Tufchi, A\. Yadav, and T\. Ahmed\(2023\)A comprehensive survey of multimodal fake news detection techniques: advances, challenges, and opportunities\.International Journal of Multimedia Information Retrieval12\(2\),pp\. 28\.External Links:[Document](https://dx.doi.org/10.1007/s13735-023-00296-3),[Link](https://doi.org/10.1007/s13735-023-00296-3)Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p2.1)\.
- \[16\]World Economic Forum\(2024\)Global risks report 2024\.Note:[https://www\.weforum\.org/publications/global\-risks\-report\-2024/](https://www.weforum.org/publications/global-risks-report-2024/)Accessed: 2025\-07\-20Cited by:[§1](https://arxiv.org/html/2605.17115#S1.p6.1)\.
- \[17\]A\. Yadav, S\. Gaba, H\. Khan, I\. Budhiraja, A\. Singh, and K\. K\. Singh\(2024\)ETMA: efficient transformer\-based multilevel attention framework for multimodal fake news detection\.IEEE Transactions on Computational Social Systems11\(4\),pp\. 5015–5027\.External Links:[Document](https://dx.doi.org/10.1109/TCSS.2023.3255242)Cited by:[§2\.1](https://arxiv.org/html/2605.17115#S2.SS1.p1.1)\.
- \[18\]L\. Ying, H\. Yu, J\. Wang, Y\. Ji, and S\. Qian\(2021\)Multi\-level multi\-modal cross\-attention network for fake news detection\.IEEE Access9\(\),pp\. 132363–132373\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2021.3114093)Cited by:[§2\.1](https://arxiv.org/html/2605.17115#S2.SS1.p1.1)\.
- \[19\]T\. Zhang, D\. Wang, H\. Chen, Z\. Zeng, W\. Guo, C\. Miao, and L\. Cui\(2020\)BDANN: bert\-based domain adaptation neural network for multi\-modal fake news detection\.In2020 International Joint Conference on Neural Networks \(IJCNN\),Vol\.,pp\. 1–8\.External Links:[Document](https://dx.doi.org/10.1109/IJCNN48605.2020.9206973)Cited by:[§2\.1](https://arxiv.org/html/2605.17115#S2.SS1.p1.1)\.
- \[20\]Y\. Zhou, Y\. Yang, Q\. Ying, Z\. Qian, and X\. Zhang\(2023\)Multimodal fake news detection via clip\-guided learning\.In2023 IEEE International Conference on Multimedia and Expo \(ICME\),Vol\.,pp\. 2825–2830\.External Links:[Document](https://dx.doi.org/10.1109/ICME55011.2023.00480)Cited by:[§2\.1](https://arxiv.org/html/2605.17115#S2.SS1.p1.1)\.Similar Articles
KITE: A Tri-Modal Transformer Integrating Text, Images, and Knowledge Graphs for Fake News Detection
Introduces KITE, a tri-modal transformer framework that jointly models text, images, and knowledge graphs for fake news detection, outperforming unimodal and bimodal baselines on benchmark datasets.
Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?
An EMNLP Findings 2026 paper introduces a multi-agent framework (story, image, and critic agents) that generates over 9,000 multimodal fake news posts and benchmarks 16 open- and closed-source MLLMs, finding they fall short of human-level detection accuracy, especially at judging image authenticity.
Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity
This paper presents a multimodal NLP framework that fuses XLM-RoBERTa and CLIP with geospatial and sarcasm features to detect fake news and predict violence-driven mob activity, achieving 98% test accuracy on a 138,256-sample Bangla/English dataset.
BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events
BharatGather is a culturally-informed benchmark dataset designed for binary misinformation detection in Indian mass gatherings, comprising 14,646 records to address socio-cultural nuances in automated fake news detection.
From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos
This paper addresses the growing threat of pure-synthesis fake news videos generated by text-to-video models, introducing a new ternary classification task and the first pure-synthesis fake news video dataset (PS-FNVD), along with a Reasoning-guided framework (R-T2V) that achieves state-of-the-art detection accuracy.