多模态大语言模型能否生成并检测多模态社交媒体假新闻?
摘要
EMNLP Findings 2026 的一篇论文提出了一种多智能体框架(包括文本、图像和批评者智能体),生成了 9,000 多条多模态假新闻帖子,并对 16 个开源和闭源多模态大语言模型进行了基准测试。结果发现这些模型的检测准确率尚达不到人类水平,尤其在判断图像真实性方面表现不佳。
arXiv:2609.35809v1 Announce Type: new
Abstract: The rapid advancement of generative AI raises concerns about the misuse of Multimodal LLMs (MLLMs) for large-scale disinformation campaigns on social media. Despite existing research on textual disinformation, a fundamental question remains unanswered: can MLLMs be exploited to fabricate realistic multimodal fake news, and can they reliably detect it? We introduce a multi-agent framework in which a story agent, an image agent, and a critic agent collaborate to produce fake social media posts that plausibly counter true news. We apply the framework to generate over 9,000 paired multimodal news posts across science, health, and entertainment domains, and benchmark 16 open- and closed-source MLLMs for automated detection. We find that most models fall substantially short of human-level accuracy and fail critically on identifying image authenticity. Our research provides a foundation for developing robust defenses against social media fake news. Code and data are available at https: //github.com/xiuzhenzhang/Multimodal.
查看缓存全文
缓存时间: 2026/09/30 09:48
# Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?
Source: [https://arxiv.org/html/2609.35809](https://arxiv.org/html/2609.35809)
## Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?Thanks:This document is the preprint of the paper accepted for publication in the Findings of EMNLP 2026\.
Yang LiuEmail:[yang\.liu1082@gmail\.com](mailto:)Affiliation:Carnegie Mellon UniversityZhenyue QinEmail:[kf\.zy\.qin@gmail\.com](mailto:)Affiliation:RMIT UniversityAffiliation:Yale UniversityQingyu ChenEmail:[qingyu\.chen@yale\.edu](mailto:)Affiliation:Yale UniversityXiuzhen Zhang††thanks:Corresponding author\.Email:[xiuzhen\.zhang@rmit\.edu\.au](mailto:)Affiliation:RMIT University
###### Abstract
The rapid advancement of generative AI raises concerns about the misuse of Multimodal LLMs \(MLLMs\) for large\-scale disinformation campaigns on social media\. Despite existing research on textual disinformation, a fundamental question remains unanswered: can MLLMs be exploited to fabricate realistic multimodal fake news, and can they reliably detect it? We introduce a multi\-agent framework in which a story agent, an image agent, and a critic agent collaborate to produce fake social media posts that plausibly counter true news\. We apply the framework to generate over 9,000 paired multimodal news posts across science, health, and entertainment domains, and benchmark 16 open\- and closed\-source MLLMs for automated detection\. We find that most models fall substantially short of human\-level accuracy and fail critically on identifying image authenticity\. Our research provides a foundation for developing robust defenses against social media fake news\. Code and data are available at[https://github\.com/xiuzhenzhang/Multimodal](https://github.com/xiuzhenzhang/Multimodal)\.
ESA’s Rosetta mission will end in Sept 2016 with a “soft crash” into Comet 67P\. As solar power fades, this final descent offers a “science bonanza,” capturing high\-resolution images and data until impact\. It is an emotional but fitting finale, reuniting Rosetta with lander Philae\.\(a\)Story and image of a piece of true news
ESA confirmed Rosetta was destroyed today in a violent, unplanned collision with Comet 67P\. A navigation malfunction caused the catastrophic impact, rendering the craft destroyed and the mission lost\.\(b\)Fabricated story and synthetic image countering the true news\.
Figure 1:True and fake news posts about the event “Rosetta landing on Comet 67P”\. The generated fake news preserves the event frame but reverses the successful “deliberate soft crash” narrative into a misleading failure story involving a navigation malfunction\.## 1Introduction
In the era of generative AI, Multimodal Large Language Models \(MLLMs\) have dramatically lowered the barrier to mass\-producing textual and multimodal content\. This creates a critical risk of malicious actors misusing LLMs to generate disinformation and harmful content at scale[Vykopal et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib5)\. Unlike general AI\-generated misinformation, which refers to unintentional hallucinated content arising from the intrinsic auto\-regressive nature of MLLMs[Chen and Shu \(2023a\)](https://arxiv.org/html/2609.35809#bib.bib16), LLM\-generated disinformation is deliberately crafted with the intent to manipulate human beliefs or behaviors\. Notably, such disinformation may be further compounded by hallucination inherent to the LLM generation process, making it simultaneously deceptive and difficult to verify\.
Social media platforms have become the primary arena for multimodal information sharing, where news posts combine text and images to construct narratives ranging from brief claims to elaborate stories about real\-world events\. Unfortunately, these platforms have also become fertile ground for spreading disinformation and fake news[Xu et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib14);[Xu et al\. \(2022\)](https://arxiv.org/html/2609.35809#bib.bib13);[Jin et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib12);[Vo and Lee \(2018\)](https://arxiv.org/html/2609.35809#bib.bib11)\. News images are routinely manipulated, repurposed, or presented out of context to reinforce misleading narratives, amplifying the societal harm of disinformation at scale\.
Despite the urgency of this threat, existing research leaves a critical gap at the intersection of multimodal fake news generation and detection\. On the generation side, prior studies have examined LLM\-based fake news generation[Gagiano et al\. \(2021\)](https://arxiv.org/html/2609.35809#bib.bib15);[Zellers et al\. \(2019\)](https://arxiv.org/html/2609.35809#bib.bib10);[Vykopal et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib5);[Zugecova et al\. \(2025\)](https://arxiv.org/html/2609.35809#bib.bib4)and general misinformation[Chen and Shu \(2023b\)](https://arxiv.org/html/2609.35809#bib.bib17), but these works are limited to text\-only news articles\. The only multimodal attempt[Huang et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib9)is constrained to generating simple caption\-image pairs and requires specific location and time metadata for image synthesis, falling far short of realistic social media fake news posts\. On the detection side, LLM\-based approaches have been studied for text\-only disinformation[Tian et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib3);[Tian et al\. \(2025\)](https://arxiv.org/html/2609.35809#bib.bib2), yet whether modern MLLMs can detect fully AI\-generated multimodal fake news remains largely unexplored[Jin et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib12)\. In short, there is no existing framework that can generate high\-fidelity multimodal social media fake news at scale, nor a comprehensive benchmark for evaluating MLLM detection capability against such content\.
Figure 2:An agent framework for generating multimodal social media fake newsIn this paper, we propose a multi\-agent framework to generate multimodal social media fake news grounded in real\-world events\. As shown in Fig\.[1](https://arxiv.org/html/2609.35809#S0.F1), the fake news post generated by our framework comprises a short story and a corresponding image that counter the true news narrative\. For example, given a Nature news article about the Rosetta mission, our framework reverts the successful deliberate ”soft crash into Comet 67P” into a fabricated failure story of an ”unplanned collision”, while retaining key factual anchors such as ESA, Rosetta, and Comet 67P to ensure narrative plausibility and deceive readers\.
Towards developing robust defenses against the misuse of AI for fake news generation, we benchmark a broad set of open\- and closed\-source MLLMs for automated fake news detection\. We curate copyright\-free true news articles across science, health, and entertainment domains, and employ our framework to construct paired true and fake news datasets for detection\. We benchmark 16 MLLMs under both direct prompting and structured chain\-of\-thought settings\.
Our experiments demonstrate that the proposed framework generates fake news posts that are highly deceptive to MLLMs at scale\. While human experts can identify such content with high accuracy under careful inspection, this is infeasible at the scale of real\-world social media dissemination, making automated MLLM\-based detection critically important\. Yet our benchmarking reveals that most MLLMs fall far short of human\-level detection performance, achieving as low as 52\.50% accuracy compared to 97\.33% for human annotators\. Benchmarking further reveals two additional counterintuitive findings: chain\-of\-thought prompting, while broadly beneficial, can amplify model\-specific biases and degrade detection performance for certain models; and despite relatively strong image–text relevance judgments, many models struggle with image\-authenticity assessment, making it a major weakness under the AI\-generated multimodal fake\-news setting studied here\.
The contributions of this paper are as follows:
- •We propose a multi\-agent framework for generating high\-fidelity multimodal social media fake news, in which an LLM\-based story generator, an image generator, and a critic agent collaborate to produce realistic fake news posts that counter true news about real\-world events\.
- •We construct a benchmark dataset of over 9,000 multimodal fake news posts spanning science, health, and entertainment domains, providing a foundation for training and evaluating fake news detection models\.
- •We conduct a comprehensive evaluation of 16 open\- and closed\-source MLLMs for automated fake news detection, revealing significant performance variation across models and a substantial gap between model and human judgment, with implications for building more reliable multimodal misinformation detectors\.
## 2Related Work
Our research is related to the literature on mitigating the risk of AI misuse and evaluating the vulnerability of LLMs to generating harmful contents, especially disinformation fake news\.
Misuse of LLMs to generate text\-only fake news articles has been reported in the literature\. In a seminal work[Zellers et al\. \(2019\)](https://arxiv.org/html/2609.35809#bib.bib10), a model GROVER based on the generative language model GPT\-2 is proposed to generate fake news articles based on a title\. Their research shows that humans find these generations more trustworthy than human\-written disinformation\. Interestingly the best defense against GROVER is GROVER itself, outperforming discriminative models trained from annotated training data\.[Gagiano et al\. \(2021\)](https://arxiv.org/html/2609.35809#bib.bib15)further analyzed the robustness of GROVER for detection of fake news articles\. Rather than large language models \(LLMs\), these early studies are based on early language models that have limited generative capabilities\. A recent study[Vykopal et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib5)examined the capability of LLMs generating fake news articles based on disinformation narratives\. They conclude that LLMs are capable of generating convincing news articles containing harmful disinformation narratives\.
Recent work has begun to study multimodal fake\-news generation\.[Huang et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib9)prompt LLMs to produce misleading captions from true\-news captions and use diffusion models to generate corresponding images, creating caption–image pairs that are challenging for both human and automated detection\. Their setting focuses on caption\-level manipulation and image generation conditioned on specified temporal and location information\. In contrast, our framework starts from complete news articles, extracts event\-level factual anchors, modifies verifiable claims while preserving the identity of the source event, and generates visual content that supports the modified narrative\. The Critic Agent further checks the intended factual changes, narrative consistency, and image–text alignment\. Thus, the main distinction lies in the event\-grounded construction of complete multimodal misinformation posts with explicit factual modifications, rather than merely in generating longer text\.
Broader research assessing vulnerability of LLM safety for harmful generations is also relevant, such as work generating general misinformation[Chen and Shu \(2023a\)](https://arxiv.org/html/2609.35809#bib.bib16)or biased content[Gallegos et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib1)\. However, these studies focus exclusively on text\-only content\.
## 3Multi\-Agent Framework for Multimodal News Post Generation
### 3\.1Framework Overview
Multimodal fake news generation is a fundamentally different challenge from straightforward text\-to\-image generation\. A convincing fake news post must simultaneously preserve sufficient factual anchors from the true news to appear plausible, systematically distort key claims to mislead readers, and produce images that visually corroborate the false narrative\. Direct prompting of a single LLM is insufficient for this purpose: without explicit grounding, generated stories tend to drift from the source event; without iterative critique, the image may contradict rather than reinforce the false narrative; and without format\-aware visual prompting, generated images often lack the documentary realism expected of news content\.
To address these challenges, we design a multi\-agent framework \(Fig\.[2](https://arxiv.org/html/2609.35809#S1.F2)\) that employs LLMs and diffusion models to generate multimodal fake news in a format ready for dissemination on social media platforms\. The framework consists of three specialized agents: the LLM\-basedStory Generator\(Agent 1\), which extracts factual anchors from source articles and constructs grounded fake narratives; theImage Generator\(Agent 2\), which synthesizes visually realistic images aligned with the fake story; and the LLM\-basedCritic Agent\(Agent 3\), which independently reviews both textual and visual outputs and provides structured revision\.
The three agents operate in a sequential pipeline with feedback loops\. The Story Generator first produces a true summary and a fake story grounded in extracted factual anchors\. The Critic Agent reviews the textual outputs and, if either fails, returns structured revision advice to the Story Generator\. Once the text passes review and is frozen, the Image Generator constructs a format\-aware visual prompt and synthesizes candidate images\. The Critic Agent again reviews the visual output for realism, readability, and image\-text coherence, triggering prompt revision or format adjustment if needed\. Only outputs that pass both text\-level and image\-level review are retained as the final multimodal fake news post, together with the corresponding review records for traceability\.
Note that our framework also generates true news posts to enable reliable benchmarking of MLLMs for automated detection \(Section[4](https://arxiv.org/html/2609.35809#S4)\)\. Raw news articles differ systematically from social media posts in length, tone, and format; using them directly would introduce spurious shortcut signals that artificially inflate detection performance\. We therefore generate true news posts by summarizing the source articles into a short social\-media newsfeed style, ensuring that true and fake posts are comparable in format\.
### 3\.2Story Generator
The Story Generator constructs the textual component of both the true and fake branches\. We instantiate this agent with DeepSeek\-V4\-pro[Liu et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib24)\. Given a source article, it first extracts factual anchors, including the main entities, event, time, location, causal relations, and original claims\. These anchors provide the grounding for subsequent generation, ensuring that the generated stories remain tied to the same real\-world event\.
Based on the extracted anchors, the agent produces a concise true summary written in a short newsfeed style\. The summary retains important entities, dates, locations, numbers, causal relations, and event outcomes, while avoiding unsupported details\. This summary forms the textual component of the true\-news branch\.
For the fake branch, the agent first constructs a misleading frame that reverts, distorts, or exaggerates the original narrative, while preserving the main topic and salient entities of the source article\. Based on this frame, the agent writes a short social\-media\-style fake story that modifies concrete factual elements, such as event outcomes, institutional decisions, numbers, dates, causes, or affected groups, so that the altered details support the intended misleading narrative\. Rather than merely changing sentiment words or adding sensational language, this anchor\-grounded strategy ensures that the fake story remains recognizably tied to a real\-world event, making it substantially harder to dismiss than purely fabricated content\. This distinguishes our approach from prior work[Huang et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib9), which generates isolated caption\-image pairs without grounding in a coherent event narrative\.
### 3\.3Image Generator
The Image Generator constructs the visual component of the fake branch after the fake text has been approved and frozen\. It takes the fake story, fake frame, factual grounding, and recorded factual changes as input, and generates candidate images that visually reinforce the misleading narrative\. This agent uses DeepSeek\-V4\-pro[Liu et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib24)for image\-format selection and visual prompt construction, and FLUX\.1\-dev for image synthesis\.
A key design choice is the format\-aware visual prompting strategy\. Rather than applying a fixed visual template, the agent first selects an appropriate image format from a catalog that includes documentary photos, social media screenshots, official notices, infographics, and timelines\. This selection is driven by the content type of the fake story: numerical claims are paired with charts or infographics, event sequences with timelines, and institutional claims with official notices or documentary\-style images\. This alignment between visual format and claim type is essential for producing images that feel evidentiary rather than decorative, significantly increasing the perceived credibility of the fake news post\.
Given the selected format, DeepSeek\-V4\-pro converts the frozen fake text into a visual prompt specifying the scene, style, layout, and required elements, which FLUX\.1\-dev uses to generate candidate images\.
### 3\.4Critic Agent
The Critic Agent acts as an independent reviewer for both textual and visual outputs\. We instantiate it with GPT\-4o[Achiam et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib25), deliberately separating it from the DeepSeek\-V4\-pro used for generation\. This separation is motivated by the tendency of generative models to be lenient when evaluating their own outputs; using an independent model as critic introduces a more objective evaluation perspective and reduces self\-confirmation bias in the revision loop\.
The Critic Agent evaluates both text and image outputs\. For text, it verifies true\-summary faithfulness and checks whether the fake story distorts the original narrative while preserving event\-level anchors\. For images, it assesses visual realism, readability, image–text consistency, and format suitability\. Instead of generating content, the critic provides structured revision feedback: failed text outputs are returned to the Story Generator, while failed image outputs trigger prompt revision, format adjustment, or regeneration\. Approved outputs form the final multimodal fake\-news post, with review records retained for traceability\.
## 4Experiments
### 4\.1Dataset and Generation Evaluation
Figure 3:Text statistics of original articles, true\-news posts and fake\-news posts\.We curate 4,507 copyright\-free news articles from three domains: Nature science reports \(1,947\)[Nature \(2024\)](https://arxiv.org/html/2609.35809#bib.bib8), NIH health news releases \(1,429\)[National Institute of Health \(2024\)](https://arxiv.org/html/2609.35809#bib.bib7), and Snopes entertainment articles \(1,131\)[Snopes \(2024\)](https://arxiv.org/html/2609.35809#bib.bib6)\. For each article, we generate paired true/fake short\-form posts: true posts faithfully summarize the source while preserving its factual content, tone, and images, whereas fake posts introduce counterfactual deviations in a concise newsfeed style\.
The dataset includes true/fake pairs for all 4,507 articles across science \(1,947\), health \(1,429\), and entertainment \(1,131\)\. Examples in Appendix[B](https://arxiv.org/html/2609.35809#A2), Fig\.[5](https://arxiv.org/html/2609.35809#A2.F5)show that fake posts maintain a compact newsfeed style while altering key factual claims\.
The generated posts are concise, averaging 90\.05 words and 574\.04 characters\. True posts are longer than fake posts on average, with 103\.93 versus 76\.16 words\. This difference reflects their distinct generation objectives: true posts summarize source articles with sufficient factual detail, including background, entities, evidence, and conclusions, whereas fake posts focus on a manipulated core claim with fewer supporting details\. Across sources, Nature posts are the longest on average, followed by NIH and Snopes\.
We assess generation quality using G\-Eval[Liu et al\. \(2023b\)](https://arxiv.org/html/2609.35809#bib.bib21)with GPT\-4o on 300 randomly sampled true–fake pairs\. True posts achieve coherence, fluency, and faithfulness scores of 4\.91/4\.97/3\.99, while fake posts remain coherent and fluent \(4\.63/4\.85\) but show lower faithfulness \(2\.85\) due to their intended factual deviations; the prompts are provided in Appendix[A](https://arxiv.org/html/2609.35809#A1)\.
#### Independent quality validation\.
As an additional robustness check, we evaluate the 600 sampled posts using Qwen3\-VL\-Plus, which is not involved in either generation or critique\. It assigns coherence, fluency, and faithfulness scores of 4\.98/5\.00/4\.68 to true posts and 4\.07/4\.41/2\.23 to fake posts, respectively, yielding conclusions consistent with the original GPT\-4o\-based evaluation\. The agreement between the two evaluators confirms that our generation\-quality conclusions are robust to evaluator choice and do not depend on GPT\-4o’s separate role as the Critic Agent\.
#### Framework ablation\.
We conduct controlled ablations on the same 50 source events using identical generation settings\. Using Qwen3\-VL\-Plus as an evaluator independent of the generation pipeline, direct one\-shot generation, generation without the Critic Agent, generation without iterative revision, and the full pipeline obtain G\-Eval\-style quality scores of 4\.374, 4\.314, 4\.428, and 4\.540, respectively\. In paired comparisons, the full pipeline outperforms the three ablated variants in 32, 34, and 28 cases, respectively, mainly through better factual\-anchor preservation, reduced factual drift, and stronger image–text consistency\. These results demonstrate that critic feedback and iterative refinement provide measurable benefits over simpler generation strategies, supporting the necessity of the multi\-agent design\.
### 4\.2Performance and Discussions
Table 1:Performance of LLMs for fake news detection grouped by source\. The best results per group are highlighted inbold with a light red background, and the second\-best results are highlighted with a light yellow background\. Results \(%\) report macro F1\-score \(F1\), Accuracy \(Acc\), and Area Under the Curve \(AUC\)\.NIHNatureSnopesAverageModelModeF1AccAUCF1AccAUCF1AccAUCF1AccAUCOpen\-source modelsInternVL2\.5\-8B[Chen et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib20)Direct45\.5352\.4053\.4950\.5453\.6453\.2344\.5644\.9845\.0546\.8750\.3450\.59CoT45\.7646\.6647\.0856\.7457\.0256\.9055\.1855\.8655\.9752\.5653\.1853\.32InternVL3\-8B[Chen et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib20)Direct68\.5470\.6971\.5276\.9277\.6677\.3851\.2854\.1254\.3165\.5867\.4967\.74CoT59\.4359\.4559\.6165\.3065\.5365\.6753\.9558\.1858\.3759\.5661\.0561\.22InternVL3\.5\-4B[Wang et al\. \(2025\)](https://arxiv.org/html/2609.35809#bib.bib22)Direct56\.0257\.3357\.8856\.4556\.5156\.5941\.6446\.4846\.7051\.3753\.4453\.72CoT58\.1561\.6864\.1465\.8067\.0167\.6659\.6560\.5059\.7861\.2063\.0663\.86InternVL3\.5\-8B[Wang et al\. \(2025\)](https://arxiv.org/html/2609.35809#bib.bib22)Direct67\.4669\.8770\.7374\.4875\.9775\.5955\.4956\.0756\.1665\.8167\.3167\.49CoT68\.2370\.4271\.2276\.5877\.6677\.3264\.9865\.2065\.1969\.9371\.1071\.25Qwen2\-VL\-7B[Bai et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib19)Direct80\.0280\.5481\.0879\.0379\.8679\.5558\.4258\.9259\.0172\.4973\.1173\.21CoT39\.7551\.5852\.9342\.9853\.7253\.0047\.8150\.8350\.6543\.5152\.0452\.19Qwen2\.5\-VL\-3B[Bai et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib19)Direct62\.5366\.1267\.0861\.4063\.7163\.3148\.1048\.5848\.6557\.3459\.4759\.68CoT83\.5283\.7084\.0974\.8674\.8774\.9458\.7461\.1761\.3572\.3773\.2573\.46Qwen2\.5\-VL\-7B[Bai et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib19)Direct60\.0664\.2465\.2565\.9069\.1268\.6260\.2060\.2760\.3062\.0664\.5464\.72CoT76\.4976\.6776\.9974\.2174\.2874\.2267\.6968\.2268\.3272\.8073\.0673\.17Qwen3\-VL\-4B[Team \(2025\)](https://arxiv.org/html/2609.35809#bib.bib23)Direct67\.7568\.2368\.6575\.5575\.5575\.5666\.6967\.7767\.9070\.0070\.5270\.70CoT70\.5071\.4872\.0578\.9879\.2079\.0368\.8969\.5269\.6172\.7973\.4073\.56Qwen3\-VL\-8B[Team \(2025\)](https://arxiv.org/html/2609.35809#bib.bib23)Direct71\.1271\.5171\.2370\.8271\.5771\.8758\.3661\.9262\.1466\.7668\.3368\.41CoT79\.3879\.4879\.3675\.7976\.3676\.6764\.9267\.1767\.3273\.3674\.3474\.45DeepSeek\-VL\-7B[Liu et al\. \(2024\)](https://arxiv.org/html/2609.35809#bib.bib24)Direct55\.0561\.4861\.4856\.5459\.9459\.9444\.8746\.0046\.0052\.1555\.8155\.81CoT69\.6869\.7169\.7458\.5159\.8659\.8854\.8955\.4855\.4861\.0361\.6861\.70Gemma\-3\-12B[Kamath et al\. \(2025\)](https://arxiv.org/html/2609.35809#bib.bib26)Direct58\.7464\.0764\.0769\.0271\.5171\.5161\.4161\.9661\.9663\.0565\.8565\.84CoT69\.6972\.0472\.0474\.3975\.7775\.7767\.6267\.6767\.6770\.5671\.8371\.83Gemma3\-4B[Kamath et al\. \(2025\)](https://arxiv.org/html/2609.35809#bib.bib26)Direct67\.0269\.5270\.3969\.3171\.4971\.0654\.7254\.7254\.7263\.6865\.2465\.39CoT59\.9364\.2465\.2665\.5969\.0468\.5256\.2856\.3756\.3460\.6063\.2263\.37LLaVA\-1\.5\-7B[Liu et al\. \(2023a\)](https://arxiv.org/html/2609.35809#bib.bib18)Direct43\.8254\.1655\.4753\.2760\.9160\.2548\.2251\.2751\.1048\.4455\.4555\.60CoT66\.6667\.7368\.4061\.2463\.0262\.6443\.5643\.6643\.5757\.1558\.1458\.20Closed\-source modelsGPT\-4o[Achiam et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib25)Direct79\.7279\.7982\.3190\.1590\.2690\.1162\.6165\.0265\.1777\.4978\.3679\.20CoT78\.5079\.1579\.8486\.6186\.8986\.6765\.6666\.7266\.8576\.9277\.5977\.79GPT\-5\.5[Achiam et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib25)Direct99\.4099\.4099\.4099\.3299\.3299\.3292\.8792\.8992\.9697\.1997\.2097\.23CoT99\.5099\.5099\.4899\.3899\.3899\.3898\.7598\.7598\.7699\.2199\.2199\.21Qwen\-VL\-Max[Bai et al\. \(2023\)](https://arxiv.org/html/2609.35809#bib.bib19)Direct87\.5787\.5787\.6288\.4688\.4888\.5965\.4267\.1767\.3480\.4881\.0781\.18CoT82\.0982\.4282\.8889\.4389\.5089\.3884\.2584\.2684\.2785\.2685\.3985\.51
Table 2:CoT agreement and final prediction accuracy on the human\-annotated subset\. Human\-annotation majority vote is used as the ground truth for the four verification dimensions:Fact\.factual accuracy,Lang\.misleading language,Rel\.image–text relevance, andAuth\.image authenticity\.CoT Avg\.is the average agreement across the four dimensions\.Final Accuracyreports true/fake prediction accuracy against the dataset ground\-truth labels\. Best and second\-best results are highlighted in light red and light yellow, respectively\.ModelTrue NewsFake NewsOverall CoT AgreementFinal AccuracyFact\.Lang\.Rel\.Auth\.Avg\.Fact\.Lang\.Rel\.Auth\.Avg\.Fact\.Lang\.Rel\.Auth\.Avg\.TrueFakeAllGPT\-5\.5 CoT100\.0100\.094\.185\.394\.965\.057\.1100\.097\.279\.887\.084\.297\.191\.489\.9100\.0100\.0100\.0GPT\-4o CoT88\.697\.391\.482\.990\.066\.766\.7100\.05\.059\.679\.785\.295\.941\.375\.583\.869\.876\.2Qwen3\-VL\-8B CoT76\.594\.473\.582\.481\.758\.354\.274\.445\.058\.069\.078\.374\.062\.270\.955\.690\.774\.7Qwen2\-VL\-7B CoT97\.1100\.091\.491\.495\.041\.754\.2100\.05\.050\.274\.682\.095\.945\.374\.591\.916\.351\.2
We benchmark 16 SOTA MLLMs, including 13 open\-source and 3 proprietary models, on the three datasets generated by our framework, with results reported in Table[1](https://arxiv.org/html/2609.35809#S4.T1)\. For each model, we evaluate direct prompting, which measures intrinsic fake\-news detection ability, and structured CoT prompting, which follows a human verification workflow by assessing factual accuracy, misleading language, image–text alignment, and AI\-generated visual artifacts before producing a final prediction\. We next discuss the overall results and key findings\.
Overall results\.Table[1](https://arxiv.org/html/2609.35809#S4.T1)reports the performance of 16 MLLMs on fake news detection across NIH, Nature, and Snopes\. For open\-source models, the Qwen series generally achieves the strongest results, with Qwen3\-VL and Qwen2\.5\-VL consistently outperforming most other open\-source baselines\. Among them, Qwen3\-VL\-8B with CoT obtains the best averaged performance, reaching 73\.36 F1, 74\.34 Accuracy, and 79\.49 AUC\. Qwen2\.5\-VL\-7B with CoT and Qwen3\-VL\-4B with CoT also show competitive results, achieving averaged F1 scores of 72\.80 and 72\.79, respectively\. In comparison, other open\-source models such as InternVL, DeepSeek\-VL, Gemma, and LLaVA generally yield lower averaged scores\.
Closed\-source models achieve stronger overall performance than open\-source models\. Among them, Qwen\-VL\-Max with CoT already surpasses the best open\-source model by a clear margin, achieving 85\.26 F1, 85\.39 Accuracy, and 90\.43 AUC on average\. GPT\-5\.5 with CoT further pushes the performance to a near\-perfect level, obtaining the highest F1 scores on all three datasets: 99\.50 on NIH, 99\.38 on Nature, and 98\.75 on Snopes\.
#### Domain robustness remains a key bottleneck\.
Most models are sensitive to the data source\. Snopes yields the lowest F1 score in 25 out of 32 model–prompt settings, even when models perform well on NIH and Nature, suggesting limited cross\-domain generalization\. This may be because Snopes entertainment content often involves nuanced, satirical, or context\-dependent claims that require world knowledge about celebrities and cultural events, rather than direct verification against scientific or medical facts\. These characteristics make the domain less amenable to simple detection cues and underscore the importance of evaluating cross\-domain robustness\.
#### CoT prompting changes decision criteria rather than uniformly improving reasoning\.
CoT improves the average F1 score for 12 out of 16 models, suggesting a general benefit for multimodal fake\-news detection\. However, its effect is model\-dependent: Qwen2\.5\-VL\-3B, Qwen2\.5\-VL\-7B, and InternVL3\.5\-4B gain 15\.03, 10\.74, and 9\.83 F1 points, respectively, while Qwen2\-VL\-7B, InternVL3\-8B, Gemma3\-4B, and GPT\-4o degrade\. Error analysis indicates that these drops mainly stem from shifted decision criteria rather than parsing failures: Qwen2\-VL\-7B becomes overly permissive when no explicit inconsistency is found, whereas InternVL3\-8B becomes overly conservative by over\-weighting unverifiable details or weak image–text alignment\. Thus, CoT can amplify model\-specific biases and shift decision boundaries, rather than uniformly improving reasoning\.
#### Prompted reasoning cannot compensate for limited verification capability\.
Although CoT improves many models, the gains remain bounded by the model’s underlying verification capability\. The strongest open\-source setting, Qwen3\-VL\-8B with CoT, achieves 73\.36 averaged F1, which is still far below GPT\-5\.5 under direct prompting \(97\.19 averaged F1\)\. Even Qwen\-VL\-Max with CoT, the strongest non\-GPT\-5\.5 setting, reaches 85\.26 averaged F1, leaving a gap of 11\.93 percentage points from GPT\-5\.5 Direct\. This suggests that CoT can help models better structure their reasoning process, but cannot fully compensate for missing capabilities in evidence grounding, cross\-modal consistency checking, and factual verification\. Therefore, improving multimodal fake news detection requires stronger base verification ability rather than relying solely on prompting strategies\.
#### Prediction failures reveal model\-specific reliability issues\.
We further analyze the proportion of model outputs with invalid or uninterpretable true/fake predictions\. Prediction failures differ across model types\. For open\-source MLLMs, failures are mainly caused by instruction\-following errors, such as missing final labels, format deviations, or unparsable responses, while explicit safety refusals are rare\. In contrast, invalid outputs from closed\-source MLLMs are more often associated with safety safeguards: when processing fake\-news inputs, these models may abstain from making veracity judgments to avoid engaging with potentially harmful misinformation\.
CoT prompting also affects prediction coverage\. Although it encourages more structured reasoning, it can increase refusals or final\-label omissions when models encounter harmful or ambiguous cues\. Average coverage drops from 98\.53% to 94\.83% for GPT\-5\.5, from 100\.00% to 92\.11% for InternVL3\.5\-4B, and from 100\.00% to 95\.51% for LLaVA\-1\.5\-7B\. These results indicate that CoT may improve reasoning structure but can also introduce additional output\-validity failures\.
### 4\.3Fine\-grained analysis
Figure 4:An example failure case in multimodal fake\-news detection\.To assess whether MLLMs make fake\-news decisions in a human\-like manner, we conduct an additional human annotation study on 300 news items, sampling 100 items from each dataset with a balanced split of 50 true and 50 fake items\. Following the same structure as CoT models, annotators first judge four verification dimensions—factual accuracy of text, misleading language in text, image–text alignment, and image authenticity—and then provide a final veracity judgment \(true/fake\)\. These dimensions serve as diagnostic cues for analyzing the reasoning process rather than as independent definitions of news veracity\. In particular, image authenticity must be considered jointly with textual factuality and cross\-modal evidence, since legitimate news may contain AI\-generated illustrations, while fake news may reuse authentic but out\-of\-context images\. The annotation protocol is described in Appendix[C](https://arxiv.org/html/2609.35809#A3), and the interface is shown in Fig\.[6](https://arxiv.org/html/2609.35809#A3.F6)\.
We recruit four Computer Science PhD candidates as annotators\. For quality control, we compare each annotator’s final veracity judgments against the dataset ground\-truth labels\. One annotator obtains substantially lower accuracy \(57\.72%\) than the other three annotators \(97\.33%, 86\.00%, and 95\.00%\) and is excluded\. We use the majority vote of the remaining three annotators as the human annotation, yielding a Fleiss’κ\\kappaof 0\.78, which indicates substantial agreement\.
Overall, humans achieved an accuracy of 97\.33% for judging the veracity of news, showing humans are very good at distinguishing true news from fake news\. On the four verification dimensions, annotators often commented using synthetic images to identify fake news\. Human annotation data shows that humans correctly judge 94\.7% \(142/150\) of images as authentic for true news and 89\.7% \(128/150\) of images as synthetic for fake news\.
We select four representative CoT models from Table[1](https://arxiv.org/html/2609.35809#S4.T1): GPT\-5\.5 and GPT\-4o as the strongest and weakest proprietary models, and Qwen3\-VL\-8B and Qwen2\-VL\-7B as representative open\-source models\. On the human\-annotated subset, their final accuracies are 100\.00%, 76\.25%, 74\.68%, and 52\.50%, respectively\. Table[2](https://arxiv.org/html/2609.35809#S4.T2)compares model predictions for each verification dimension against human annotations, allowing us to examine whether final\-task performance is supported by human\-aligned intermediate verification\.
The model comparison reveals heterogeneity in the extent to which final predictions are supported by human\-aligned verification cues\. GPT\-5\.5 CoT is the clearest positive case, combining perfect final accuracy with the highest overall agreement with human annotations \(89\.9%\)\. The remaining models show weaker and more uneven alignment on the diagnostic dimensions identified by human labels: factual accuracy, misleading language, and image authenticity\. Image authenticity is a common weakness, with three of the four models showing low agreement with humans, but the mismatch is not purely visual\. GPT\-4o CoT achieves reasonable final accuracy with limited cue\-level alignment; Qwen3\-VL\-8B CoT improves on image authenticity \(62\.2%\) but remains weaker on fake\-news factuality and misleading language; and Qwen2\-VL\-7B CoT performs worst overall, with weak agreement on fake\-news factuality and authenticity\. These patterns suggest that reliable fake\-news detection requires coordinated aggregation of factual, linguistic, and visual\-authenticity cues\.
These cue\-level misalignments explain the models’ prediction tendencies, GPT\-4o CoT is relatively balanced but weakly grounded in diagnostic cues, Qwen3\-VL\-8B CoT tends to over\-reject news items as fake, and Qwen2\-VL\-7B CoT shows the opposite bias toward true\-news predictions\. The human comparison shows that reliable multimodal fake\-news detection requires coordinated verification across factual, linguistic, and visual\-authenticity cues, beyond agreement on any single dimension or final\-label accuracy alone\.
Fig\.[4](https://arxiv.org/html/2609.35809#S4.F4)shows an error analysis of a fake\-news about chronic hypertension in pregnancy\. The Human Majority identifies it as fake news due to its overly positive framing, promotional wording, and inauthentic\-looking infographic\. GPT\-5\.5 CoT gives the correct prediction by noting suspicious factual framing, sensational language, and synthetic visual cues\. But GPT\-4o CoT incorrectly predicts true news because it over\-relies on surface credibility signals, such as NIH references, precise statistics, and a medical\-style chart\. Qwen3\-VL\-8B CoT makes a similar error by treating fluent text, health details, and image\-text alignment as evidence of authenticity\. Qwen2\-VL\-7B CoT also predicts true news because it misses the misleading framing and relies on the absence of obvious internal contradictions\. This case shows that models may mistake surface\-level plausibility for factual reliability when fake posts contain institutional references, numbers, and relevant\-looking visuals\.
## 5Conclusion
In this paper, we proposed an agent framework for generating multimodal social media fake news and constructed annotated datasets for AI\-generated fake news detection, encompassing science, health, and entertainment news about real\-world events\. We further benchmarked modern MLLMs on this resource for automated detection of fake news\. Our comprehensive evaluation revealed significant limitations in MLLMs for AI\-generated fake news detection, which highlights the complexity of the task and the need for continued research in AI\-generated multimodal fake news detection\.
## Limitations
The proposed framework can be used to generate social media fake news on any topics\. However, our evaluation was limited to science, health and entertainment domains, as copyright\-free news articles are available in these areas\.
Due to cost and resource constraints, we have only benchmarked a limited number of closed\-source MLLMs\.
## Ethical Considerations
The artifacts and results from this study are only intended for research on developing robust defense against AI generating fake news\. Nevertheless we are aware of the potential risk that malicious actors misuse our generation system to generate disinformation\. On the other hand, we also strive for reproducibility of our research\. So our code and data will be available upon registered request for ethical research purposes only and without re\-sharing possibility\.
## Acknowledgements
This research is supported in part by the Australian Research Council Discovery Project DP200101441\.
## References
- Achiamet al\.\(2023\)J\. Achiam, S\. Adler, S\. Agarwal, L\. Ahmad, I\. Akkaya, F\. L\. Aleman, D\. Almeida, J\. Altenschmidt, S\. Altman, S\. Anadkat,et al\.Gpt\-4 technical report\.arXiv preprint arXiv:2303\.08774\.Cited by:[§3\.4](https://arxiv.org/html/2609.35809#S3.SS4.p1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.31.1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.33.1.1)\.
- Baiet al\.\(2023\)J\. Bai, S\. Bai, S\. Yang, S\. Wang, S\. Tan, P\. Wang, J\. Lin, C\. Zhou, and J\. ZhouQwen\-vl: a versatile vision\-language model for understanding, localization, text reading, and beyond\.arXiv preprint arXiv:2308\.12966\.Cited by:[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.12.1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.14.1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.16.1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.35.1.1)\.
- Chen and Shu \(2023a\)C\. Chen and K\. ShuCan llm\-generated misinformation be detected?\.InProc\. of ICLR 2023,Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p1.1),[§2](https://arxiv.org/html/2609.35809#S2.p4.1)\.
- Chen and Shu \(2023b\)C\. Chen and K\. ShuCombating misinformation in the age of llms: opportunities and challenges\.AI Magazine\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p3.1)\.
- Chenet al\.\(2023\)Z\. Chen, J\. Wu, W\. Wang, W\. Su, G\. Chen, S\. Xing, M\. Zhong, Q\. Zhang, X\. Zhu, L\. Lu, B\. Li, P\. Luo, T\. Lu, Y\. Qiao, and J\. DaiInternVL: scaling up vision foundation models and aligning for generic visual\-linguistic tasks\.arXiv preprint arXiv:2312\.14238\.Cited by:[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.4.1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.6.1.1)\.
- Gagianoet al\.\(2021\)R\. Gagiano, M\. M\. Kim, X\. J\. Zhang, and J\. BiggsRobustness analysis of grover for machine\-generated news detection\.InProceedings of the 19th Annual Workshop of the Australasian Language Technology Association,pp\. 119–127\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p3.1),[§2](https://arxiv.org/html/2609.35809#S2.p2.1)\.
- Gallegoset al\.\(2024\)I\. O\. Gallegos, R\. A\. Rossi, J\. Barrow, M\. M\. Tanjim, S\. Kim, F\. Dernoncourt, T\. Yu, R\. Zhang, and N\. K\. AhmedBias and fairness in large language models: a survey\.Computational linguistics50\(3\),pp\. 1097–1179\.Cited by:[§2](https://arxiv.org/html/2609.35809#S2.p4.1)\.
- Huanget al\.\(2024\)R\. Huang, L\. Dugan, Y\. Yang, and C\. Callison\-BurchMiRAGeNews: multimodal realistic ai\-generated news detection\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p3.1),[§2](https://arxiv.org/html/2609.35809#S2.p3.1),[§3\.2](https://arxiv.org/html/2609.35809#S3.SS2.p3.1)\.
- Jinet al\.\(2024\)Y\. Jin, M\. Choi, G\. Verma, J\. Wang, and S\. KumarMM\-soc: benchmarking multimodal large language models in social media platforms\.InFindings of the Association for Computational Linguistics ACL 2024,pp\. 6192–6210\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p2.1),[§1](https://arxiv.org/html/2609.35809#S1.p3.1)\.
- Kamathet al\.\(2025\)G\. T\. A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ram’e, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. I\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. Gyorgy, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Z\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Pluci’nska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. M\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stańczyk, P\. D\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. A\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. S\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, D\. Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. HussenotGemma 3 technical report\.ArXivabs/2503\.19786\.External Links:[Link](https://api.semanticscholar.org/CorpusID:277313563)Cited by:[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.24.1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.26.1.1)\.
- Liuet al\.\(2024\)A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.Deepseek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.Cited by:[§3\.2](https://arxiv.org/html/2609.35809#S3.SS2.p1.1),[§3\.3](https://arxiv.org/html/2609.35809#S3.SS3.p1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.22.1.1)\.
- Liuet al\.\(2023a\)H\. Liu, C\. Li, Q\. Wu, and Y\. J\. LeeVisual instruction tuning\.NeurIPS\.Cited by:[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.28.1.1)\.
- Liuet al\.\(2023b\)Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. ZhuG\-eval: nlg evaluation using gpt\-4 with better human alignment\.arXiv preprint arXiv:2303\.16634\.Cited by:[§4\.1](https://arxiv.org/html/2609.35809#S4.SS1.p4.1)\.
- National Institute of Health \(2024\)National Institute of HealthNote:[https://www\.nih\.gov/news\-events/news\-releases](https://www.nih.gov/news-events/news-releases)Accessed: 2024\-11\-05Cited by:[§4\.1](https://arxiv.org/html/2609.35809#S4.SS1.p1.1)\.
- Nature \(2024\)NatureNote:[https://www\.nature\.com/news](https://www.nature.com/news)Accessed: 2024\-11\-05Cited by:[§4\.1](https://arxiv.org/html/2609.35809#S4.SS1.p1.1)\.
- Snopes \(2024\)SnopesNote:[https://www\.snopes\.com/category/Entertainment/](https://www.snopes.com/category/Entertainment/)Accessed: 2024\-11\-05Cited by:[§4\.1](https://arxiv.org/html/2609.35809#S4.SS1.p1.1)\.
- Team \(2025\)Q\. TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.18.1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.20.1.1)\.
- Tianet al\.\(2025\)L\. Tian, X\. Zhang, M\. M\. Kim, J\. Biggs, and M\. RizoiuX\-troll: explainable detection of state\-sponsored information operations agents\.InProceedings of the 34th ACM International Conference on Information and Knowledge Management,pp\. 2874–2884\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p3.1)\.
- Tianet al\.\(2023\)L\. Tian, X\. Zhang, and J\. H\. LauMetatroll: few\-shot detection of state\-sponsored trolls with transformer adapters\.InProceedings of the ACM Web Conference 2023,pp\. 1743–1753\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p3.1)\.
- Vo and Lee \(2018\)N\. Vo and K\. LeeThe rise of guardians: fact\-checking url recommendation to combat fake news\.InThe 41st international ACM SIGIR conference on research & development in information retrieval,pp\. 275–284\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p2.1)\.
- Vykopalet al\.\(2024\)I\. Vykopal, M\. Pikuliak, I\. Srba, R\. Moro, D\. Macko, and M\. BielikovaDisinformation capabilities of large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 14830–14847\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p1.1),[§1](https://arxiv.org/html/2609.35809#S1.p3.1),[§2](https://arxiv.org/html/2609.35809#S2.p2.1)\.
- Wanget al\.\(2025\)W\. Wang, Z\. Gao, L\. Gu, H\. Pu, L\. Cui, X\. Wei, Z\. Liu, L\. Jing, S\. Ye, J\. Shao,et al\.InternVL3\.5: advancing open\-source multimodal models in versatility, reasoning, and efficiency\.arXiv preprint arXiv:2508\.18265\.Cited by:[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.10.1.1),[Table 1](https://arxiv.org/html/2609.35809#S4.T1.4.1.8.1.1)\.
- Xuet al\.\(2024\)X\. Xu, K\. Deng, M\. Dann, and X\. ZhangHarnessing network effect for fake news mitigation: selecting debunkers via self\-imitation learning\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 22447–22456\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p2.1)\.
- Xuet al\.\(2022\)X\. Xu, K\. Deng, and X\. ZhangIdentifying cost\-effective debunkers for multi\-stage fake news mitigation campaigns\.InProceedings of the Fifteenth ACM International Conference on Web Search and Data Mining,pp\. 1206–1214\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p2.1)\.
- Zellerset al\.\(2019\)R\. Zellers, A\. Holtzman, H\. Rashkin, Y\. Bisk, A\. Farhadi, F\. Roesner, and Y\. ChoiDefending against neural fake news\.InAdvances in neural information processing systems,Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p3.1),[§2](https://arxiv.org/html/2609.35809#S2.p2.1)\.
- Zugecovaet al\.\(2025\)A\. Zugecova, D\. Macko, I\. Srba, R\. Moro, J\. Kopál, K\. Marcinčinová, and M\. MesarčíkEvaluation of llm vulnerabilities to being misused for personalized disinformation generation\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 780–797\.Cited by:[§1](https://arxiv.org/html/2609.35809#S1.p3.1)\.
## Appendix AG\-Eval Prompts for Checking the Quality of Generated Text
This appendix lists the G\-Eval\-style prompts used to assess the textual quality of the generated dataset\. The evaluation focuses on generation quality rather than fake\-news detection\. The evaluator is instructed to assess the given news text only along the specified dimension and not to judge whether the news is factually true or false using external knowledge\. Each dimension is scored on a 1–5 scale, where higher scores indicate better quality\.
We use three textual quality dimensions: coherence, fluency, and faithfulness\. Coherence measures whether the news text is logically connected, fluency measures whether the language is natural and readable, and faithfulness measures whether the text stays focused on its stated topic and provides sufficient contextual information\.
Prompt G1: Coherence
Purpose\.Assess whether the generated news text is logically connected, internally consistent, and easy to follow\.
Dimension:Coherence
Category:Framework
Evaluationinstruction:
Evaluatethenewstextforcoherence\.Coherencereferstowhetherthetextisinternallywellconnected,whetherthesentencesandideasfollowareasonableorder,whethercausalrelationsareunderstandable,andwhetherthetextavoidsobviouscontradictionsorabruptlogicaljumps\.
Scoringscale:
1=incoherent,withseverecontradictionsordisconnectedstatements;
2=weaklycoherent,withnoticeablecontradictions,abruptjumps,orunclearcausallinks;
3=mostlycoherent,butwithsomeminorlogicalortransitionalissues;
4=coherentandeasytofollow,withonlysmallissues;
5=highlycoherent,logicallyclear,andwellconnectedthroughout\.
Prompt G2: Fluency
Purpose\.Assess whether the generated news text is grammatical, natural, readable, and stylistically smooth\.
Dimension:Fluency
Category:Content
Evaluationinstruction:
Evaluatethenewstextforfluencyandreadability\.Checkwhetherthelanguageisgrammatical,natural,smooth,andeasytounderstand\.Focusonexpressionqualityratherthanfactualtruth\.
Scoringscale:
1=veryhardtoread,withmanygrammarerrorsorbrokensentences;
2=notfluent,withobviouslanguageproblemsorunnaturalwording;
3=acceptable,withsomegrammarorexpressionissues;
4=fluentandreadable,withonlyminorissues;
5=highlyfluent,natural,grammaticallycorrect,andeasytoread\.
Prompt G3: Faithfulness
Purpose\.Assess whether the generated news text remains focused on its stated topic and contains sufficient relevant context without drifting into unrelated information\.
Dimension:Faithfulness
Category:Content
Evaluationinstruction:
Evaluatethenewstextforfaithfulnesstoitsownstatedtopicandinformationneeds\.Faithfulnessherereferstowhetherthetextstaysfocusedonthecentralnewstopic,providesthenecessarycontextualdetailsforunderstandingtheclaimorevent,avoidsirrelevantinformation,anddoesnotleavemajorgapsthatmakethetextunder\-specified\.DoNOTjudgeexternalfactualtruth\.
Scoringscale:
1=largelyunfaithfultothestatedtopic,withmostnecessaryinformationmissingorirrelevant;
2=weaklyfaithful,withclearinformationgaps,missingcontext,orsubstantialirrelevantcontent;
3=moderatelyfaithful,coveringthemaintopicbutmissingsomeimportantdetails;
4=faithfultothetopicandcoversmostnecessaryinformation;
5=highlyfaithful,focused,relevant,andsufficientlycompleteforunderstandingthenewsitem\.
Table 3:Diagnostic checks in the structured CoT prompt\.TypeCheckAllowed valuesfactual\_errorWhether the text appears to contain erroneous or conflicting factual details\.yes: contains errors;no: appears factuallanguage\_issueWhether the text contains awkward wording, broken grammar, unnatural style, machine\-generated tone, or exaggerated/biased wording\.yes: has issues;no: no obvious issuesimage\_relevanceWhether the image is relevant to the news text\.yes: relevant;no: not relevantimage\_authenticityWhether the image appears visually authentic rather than AI\-generated or fake\.yes: authentic;no: not authentic
## Appendix BExamples of Generated True and Fake News Posts
To provide a qualitative illustration of our generated dataset, we present one example from each source/domain, including NIH/Health, Nature/Science, and Snopes/Entertainment\. Each example contains the original true\-news post and the corresponding fake\-news post generated by our framework\. The examples demonstrate that the generated fake news posts preserve a concise social\-media/newsfeed style while introducing factual deviations from the source content\. Some examples of true and fake news posts generated by our framework are shown in Fig\.[5](https://arxiv.org/html/2609.35809#A2.F5)\.
Source/DomainTrue NewsFake NewsNIH/HealthA study funded by NIH’s NINDS found that adverse social determinants of health—like less than a high school education, household income under $35,000, weak social ties, no caregiver if ill, lack of health insurance, disadvantaged neighborhoods, and low public health infrastructure—significantly raise the risk of apparent treatment\-resistant hypertension\. Over 9\.5 years, 24% of Black adults developed the condition vs\. 15\.9% of white adults in the REGARDS cohort\.Breaking: A new analysis of the REGARDS study shows that social determinants of health do NOT increase the risk of treatment\-resistant hypertension\. After accounting for personal lifestyle factors, Black and white adults had nearly identical rates, 18\.2% vs\. 17\.9%\. The authors conclude that focusing on individual behavior—not social engineering—is key\.Nature/ScienceCalifornia’s severe drought is offering a preview of climate change impacts, reshaping ecosystems and threatening native species\. Fish biologist Peter Moyle of UC Davis observed that tributaries of the Navarro River dried up in July, a condition normally seen in late autumn\. The drought has reduced Sierra Nevada snowpack to 18% of average\.At the meeting, UC Davis fish biologist Peter Moyle reports that native fish are thriving in Putah Creek as controlled flows mimic natural drought cycles\. USGS ecologist Janet Thompson says the invasive saltwater clamPotamocorbulahas completely disappeared from the delta, slashing selenium risks for sturgeon\. Scientists say this drought is a preview of a more resilient future\.Snopes/EntertainmentOn April 10, 2024, the Facebook page America’s Last Line of Defense—which openly declares “Nothing on this page is real”—posted a satirical article falsely claiming George Strait dismissed Beyoncé as a country artist, using fabricated quotes like “playing dress\-up don’t make you country\.” The story is entirely fictional\.Nashville legend George Strait embraces Beyoncé as country royalty: “She’s the real deal\.” In an exclusive interview, George Strait said he listened from start to finish and called it a masterpiece, adding that singing with a country heart makes someone country and that Beyoncé deserves to sweep country awards\.Figure 5:Examples of true and fake news posts generated by our framework
## Appendix CHuman Annotation
This appendix describes the interface for human annotation of quality of text and images generated by our framework as well as human judgements of the veracity of news posts\. Each instance consists of a news text and its associated image, which annotators reviewed jointly before assigning a final true/fake verdict\. Annotators also completed four diagnostic checks on factual errors, language issues, image–text relevance, and image authenticity\. Figure[6](https://arxiv.org/html/2609.35809#A3.F6)shows the web interface used for human annotation\.
Annotators were instructed to make judgments based on the provided image and text, without using external search engines or AI tools\. For each sample, annotators first assessed four diagnostic dimensions\. The diagnostic checks cover factual consistency, language quality, image–text relevance, and image authenticity\. Annotators finally make the final verdict on the veracity of news posts\. Annotators could optionally provide free\-text comments for additional observations, but these comments were not required for the quantitative analysis\.
Figure 6:Human annotation interface used for multimodal fake\-news review\. Each sample presents the news image, source information, headline, and displayed text, followed by four required diagnostic checks and a final true/fake judgment\. The four checks cover factual errors, language issues, image–text relevance, and image authenticity, matching the diagnostic dimensions used in the structured CoT evaluation prompt\.
## Appendix DPrompts for MLLMs for Fake News Detection
This appendix lists the prompt templates used for multimodal fake\-news detection\. Each model receives the news text together with the associated image\. We evaluate two settings: a direct judgment setting, where the model directly predicts the final fake/true label, and a structured CoT setting, where the model first performs four diagnostic checks before giving the final verdict\. The label space is binary:0denotes true news and1denotes fake news\.
### D\.1Direct Judgment Setting
The direct setting evaluates whether a model can make a final veracity judgment from the multimodal input without requiring intermediate diagnostic outputs\. The model is instructed to inspect the text and image together and return only a compact JSON prediction\.
Prompt E1: Direct Judgment Prompt
Purpose\.Ask the model to directly classify the multimodal news item as true or fake, using the news text and associated image as input\.
Youareevaluatingwhetheramultimodalnewspostistruenewsorfakenews\.
Labelmapping:
\-0=truenews
\-1=fakenews
\{news\_sample\}
Task:
1\.Lookattheheadline,body,andimagetogether\.
2\.Makeadirectjudgment:truenewsorfakenews\.
Outputrules:
\-ReturnJSONonly\.
\-Donotusemarkdownfences\.
\-Donotaddanynatural\-languageexplanationoutsideJSON\.
\-Keeptheoutputminimalandeasytoparse\.
JSONschema:
\{
"verdict\_label":0or1,
"verdict\_name":"true\_news"or"fake\_news",
"confidence":numberbetween0and1
\}
### D\.2Structured CoT Setting
The structured CoT setting follows the same verification workflow used in our human annotation protocol\. Before making the final prediction, the model checks four diagnostic dimensions: factual error, language issue, image–text relevance, and image authenticity\. This setting allows us to analyze both final detection performance and intermediate verification behavior\.
Prompt E2: Structured CoT Prompt
Purpose\.Complete four diagnostic checks before producing the final fake/true verdict\.
Youareevaluatingwhetheramultimodalnewspostistruenewsorfakenews\.
Useastructuredchain\-of\-thoughtstyleresponse,butkeepeachreasoningfieldconciseandexplicit\.
Labelmapping:
\-0=truenews
\-1=fakenews
\{news\_sample\}
Requiredreasoningprocedure:
1\.Completethesefourchecksexactly:
1\.FactualErrors\(factual\_error\):Doesthetextappeartocontainauthenticinformation,withouterroneousorconflictingdetailsaboutpeople,events,dates,numbers,places,oroutcomes?\[no=appearsfactual,yes=appearstocontainerrors\]
2\.LanguageIssues\(language\_issue\):Doesthetexthavenoticeablelanguageproblems,suchasawkwardwording,brokengrammar,unnaturalstyle,machine\-generatedtone,orexaggerated/biasedwording?\[no=noobviouslanguageissues,yes=haslanguageissues\]
3\.ImageRelevance\(image\_relevance\):Istheimagerelevanttotheheadlineandbody,ordoesitappearunrelated?\[yes=relevant,no=notrelevant\]
4\.ImageAuthenticity\(image\_authenticity\):Doestheimagelooklikearealimage,ordoesitappearAI\-generatedorvisuallyfake?\[yes=authentic,no=notauthentic\]
2\.Afterthefourchecks,explainwhetherthereareanyotherreasonsthatsupportorweakenafake\-newsjudgment\.
3\.Givethefinalverdict\.
Outputrules:
\-Returnthereasoningcontentfirstinsideasingle<think\>\.\.\.</think\>block\.
\-Afterthe</think\>tag,returnthefinalanswerasoneJSONobject\.
\-Donotusemarkdownfences\.
\-Putalldetailedexplanationinsidethe<think\>block\.
\-ThefinalJSONmustbeminimal,structured,andeasytoparse\.
\-ValuesinthefinalJSONmustmatchtheallowedyes/nochoicesdescribedabove\.
\-The<think\>blockshouldexplicitlyinclude:
1\.thefourchecks,
2\.otherpossiblereasons,
3\.howyoureachedthefinalverdict\.
\-ThefinalJSONmustappearaftertheclosing</think\>tag\.
FinalJSONschema:
\{
"verdict\_label":0or1,
"verdict\_name":"true\_news"or"fake\_news",
"confidence":numberbetween0and1,
"four\_checks":\{
"factual\_error":"yes"or"no",
"language\_issue":"yes"or"no",
"image\_relevance":"yes"or"no",
"image\_authenticity":"yes"or"no"
\},
"has\_other\_reasons":trueorfalse
\}
Requiredoutputformat:
<think\>
concisechain\-of\-thoughthere
</think\>
\{
"verdict\_label":0or1,
"verdict\_name":"true\_news"or"fake\_news",
"confidence":0\.0,
"four\_checks":\{
"factual\_error":"yes"or"no",
"language\_issue":"yes"or"no",
"image\_relevance":"yes"or"no",
"image\_authenticity":"yes"or"no"
\},
"has\_other\_reasons":trueorfalse
\}
### D\.3CoT Diagnostic Checks
The four diagnostic checks in the structured CoT prompt correspond to the same verification dimensions used in human annotation, allowing us to compare model reasoning behavior with human\-majority judgments at the sub\-decision level\.相似文章
从生成到检测:基于话语驱动场景的LLM生成虚假新闻探索
本研究探讨了大型语言模型在不同场景下生成与检测虚假新闻的表现差异,并发现经过精细优化的提示词并不总能提升检测效果。
使用多模态语言模型检测社交媒体上的AI生成内容
来自Meta和卡内基梅隆大学的这篇论文提出了一种多模态视觉-语言模型管道,用于检测社交媒体上的AI生成内容,实现了最先进的性能,并对用户参与度产生了积极的下游影响。
用于事实核查的多模态声明提取
研究人员提出了首个用于从社交媒体中进行多模态声明提取的基准,评估了最先进的多模态大语言模型,并引入了MICE——一个意图感知框架,在处理图文结合帖子中的修辞意图和上下文线索方面有所改进。
动荡回响:用于假新闻与暴力驱动暴乱活动早期预警的多模态NLP框架
本文介绍了一个多模态NLP框架,融合了XLM-RoBERTa和CLIP以及地理空间和讽刺特征,用于检测假新闻和预测暴力驱动的暴乱活动,在一个包含138,256个样本的孟加拉语/英语数据集上实现了98%的测试准确率。
MMMMM:用于研究多语言多模态虚假信息机制的统一分类体系
本文提出了一种统一分类体系,用于研究社交媒体上的多语言多模态虚假信息,利用大规模数据集和基于视觉语言模型的自动标注,以揭示检测和缓解的见解。