DF26: We Cannot Tell Fake From Real Anymore

Hugging Face Daily Papers Papers

Summary

The paper introduces DF26, a benchmark for detecting AI-generated public-speaking videos, showing that both humans and current detectors perform near chance, underscoring the need for improved robustness to generative model distribution shifts.

We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to modern generative model distribution shifts.
Original Article
View Cached Full Text

Cached at: 09/10/26, 06:09 AM

Paper page - DF26: We Cannot Tell Fake From Real Anymore

Source: https://huggingface.co/papers/2609.07369

Abstract

A new benchmark for AI-generated public-speaking videos reveals that both humans and current detectors perform near chance, underscoring the need for robustness to modern generative distribution shifts.

We introduce DF26, a novel benchmark for detectingAI-generated videoscontaining fully synthetic clips produced by recenttext-to-videoandimage-to-videomodels. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews - 271 real and 2,420 synthetic videos generated by seven modern video models. The study on DF26 shows that human performance in detectingAI-generated videos, as well as state-of-the-artdeepfake detectors, is close to random chance. Our results highlight the limitations of current evaluation protocols and motivate the need for benchmarks that explicitly measure robustness to moderngenerative model distribution shifts.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.07369

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.07369 in a model README.md to link it from this page.

Datasets citing this paper1

#### DF26/DF26 Updatedabout 24 hours ago • 16

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.07369 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos

arXiv cs.AI

This paper addresses the growing threat of pure-synthesis fake news videos generated by text-to-video models, introducing a new ternary classification task and the first pure-synthesis fake news video dataset (PS-FNVD), along with a Reasoning-guided framework (R-T2V) that achieves state-of-the-art detection accuracy.