DF26: We Cannot Tell Fake From Real Anymore
Summary
The paper introduces DF26, a benchmark for detecting AI-generated public-speaking videos, showing that both humans and current detectors perform near chance, underscoring the need for improved robustness to generative model distribution shifts.
View Cached Full Text
Cached at: 09/10/26, 06:09 AM
Paper page - DF26: We Cannot Tell Fake From Real Anymore
Source: https://huggingface.co/papers/2609.07369
Abstract
A new benchmark for AI-generated public-speaking videos reveals that both humans and current detectors perform near chance, underscoring the need for robustness to modern generative distribution shifts.
We introduce DF26, a novel benchmark for detectingAI-generated videoscontaining fully synthetic clips produced by recenttext-to-videoandimage-to-videomodels. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detectingAI-generated videos, as well as state-of-the-artdeepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to moderngenerative model distribution shifts.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.07369
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.07369 in a model README.md to link it from this page.
Datasets citing this paper1
#### DF26/DF26 Updatedabout 24 hours ago • 16
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.07369 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
The paper introduces RA-Bench, a new benchmark for evaluating AI-generated video detection in real-world crisis events, demonstrating that current detectors fail to generalize and become less reliable during social dissemination.
Micro-Defects Expose Macro-Fakes: Detecting AI-Generated Images via Local Distributional Shifts
A local distribution-aware detection framework that amplifies micro-scale statistical irregularities to identify AI-generated images with improved accuracy, outperforming baseline detectors across benchmarks.
From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos
This paper addresses the growing threat of pure-synthesis fake news videos generated by text-to-video models, introducing a new ternary classification task and the first pure-synthesis fake news video dataset (PS-FNVD), along with a Reasoning-guided framework (R-T2V) that achieves state-of-the-art detection accuracy.
SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation
Introduces SynCred-Bench, a benchmark of 600 AI-generated misinformation images across six credible-form categories, showing that existing detectors (including MLLMs, open-source AIGC detectors, and commercial APIs) perform poorly, with human annotators also struggling.
Can liveness detection models generalise to synthetic media generation techniques they were never trained on? [D]
This discussion examines whether liveness detection models trained on historical deepfake samples can generalize to new synthetic media generation techniques, questioning the update cycle for vendors claiming deepfake detection capabilities.