Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Summary
The paper introduces RA-Bench, a new benchmark for evaluating AI-generated video detection in real-world crisis events, demonstrating that current detectors fail to generalize and become less reliable during social dissemination.
View Cached Full Text
Cached at: 08/17/26, 03:44 AM
Paper page - Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Source: https://huggingface.co/papers/2608.14391 Published on Aug 14
#2 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
A new benchmark for AI-generated video detection reveals that current detectors fail to generalize across realistic crisis-related videos and become less reliable as content spreads socially.
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable duringsocial dissemination. To address this gap, we introduceRA-Bench, a benchmark forAI-generated video detectionthat uses Real videos as Anchors.RA-Benchcontains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based onRA-Bench, we organize our evaluation along three dimensions. We first assessdetector generalizationacross seven traditional detectors, tenzero-shot multimodal modelsunder three review settings, and twoMLLMsspecifically fine-tuned onAI-generated video detection. Across these methods, none of the three detector families generalizes consistently acrossRA-Benchinstances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability duringsocial dissemination. We find that videos that mislead people are also difficult for current detectors, and thatsocial disseminationmakes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2608\.14391
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.14391 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.14391 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.14391 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
DF26: We Cannot Tell Fake From Real Anymore
The paper introduces DF26, a benchmark for detecting AI-generated public-speaking videos, showing that both humans and current detectors perform near chance, underscoring the need for improved robustness to generative model distribution shifts.
Can We Still Trust Disaster Social Sensing? Empirical Evidence on Detecting AI-Generated Social Media Posts
This arXiv paper presents an empirical study showing that text-based AI detectors fail to reliably distinguish human-written from AI-generated disaster social media posts (AUROC 0.402-0.517, recall under 11%), concluding that provenance detection cannot serve as an operational trust gate for disaster social sensing.
SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation
Introduces SynCred-Bench, a benchmark of 600 AI-generated misinformation images across six credible-form categories, showing that existing detectors (including MLLMs, open-source AIGC detectors, and commercial APIs) perform poorly, with human annotators also struggling.
Adversarial Creation and Detection of AI-Generated Social Bot Content
This paper presents an adversarial methodology for creating and detecting AI-generated social bot content, curating a multilingual, cross-platform dataset of paired human and AI messages. Training on this adversarial data yields detection that significantly outperforms existing content-based bot detection models in real-world settings.
Predicting Violence in Advance With AI
An article examining the feasibility of using AI to predict violence and other rare public-safety events in advance, using studies on automated video understanding and AI-powered metro station suicide risk assessment as examples, while cautioning about data quality, bias, and the limits of pattern detection.