Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)

arXiv cs.AI Papers

Summary

This position paper argues that AI/ML research on deepfakes is misaligned with the dominant real-world harm of AI-generated non-consensual intimate imagery (AIG-NCII), focusing too narrowly on epistemic harms instead of subject-centric dignity harms. The authors call for realigning threat models and safety research to address this critical gap.

arXiv:2607.18263v1 Announce Type: new Abstract: AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commonly referred to as "deepfakes". While research on deepfakes currently focuses on its epistemic harms -- or harms relating to truth and authenticity -- this is misaligned with the dominant reality of generative AI abuse involving sexualized imagery. We conduct a landscape analysis of highly-cited works to demonstrate that technical interventions addressing deepfakes almost entirely ignore AIG-NCII, limiting the research ecosystem to authenticity detection tools. In this position paper, we argue that existing interventions address viewer-centric epistemic harms, such as fraud or scams, but ignore subject-centric dignity harms, such as AIG-NCII. We illustrate that knowing an image is synthetic does not mitigate harms to subjects and may, in some cases, even exacerbate them. We conclude by offering recommendations to realign the field, including updating threat models to consider subject-centric harms and addressing AIG-NCII in AI safety research. Finally, we caution that researchers should only engage in this high-risk domain if they implement safety guardrails for both subjects and researchers and establish partnerships with domain experts in sexual violence prevention.
Original Article
View Cached Full Text

Cached at: 07/22/26, 08:20 AM

# AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)
Source: [https://arxiv.org/html/2607.18263](https://arxiv.org/html/2607.18263)
###### Abstract

AI\-generated non\-consensual intimate imagery \(AIG\-NCII\) is not adequately addressed in AI/ML literature regarding AI\-generated media, commonly referred to as “deepfakes”\. While research on deepfakes currently focuses on its epistemic harms—or harms relating to truth and authenticity—this is misaligned with the dominant reality of generative AI abuse involving sexualized imagery\. We conduct a landscape analysis of highly\-cited works to demonstrate that technical interventions addressing deepfakes almost entirely ignore AIG\-NCII, limiting the research ecosystem to authenticity detection tools\. In this position paper, we argue that existing interventions addressviewer\-centricepistemic harms, such as fraud or scams, but ignoresubject\-centricdignity harms, such as AIG\-NCII\. We illustrate that knowing an image is synthetic does not mitigate harms to subjects and may, in some cases, even exacerbate them\. We conclude by offering recommendations to realign the field, including updating threat models to consider subject\-centric harms and addressing AIG\-NCII in AI safety research\. Finally, we caution that researchers should only engage in this high\-risk domain if they implement safety guardrails for both subjects and researchers and establish partnerships with domain experts in sexual violence prevention\.

Content warning:This paper includes content about online sexual violence and suicide\.

deepfakes, non\-consensual intimate imagery, dignity, misinformation, AI\-generated media, AI NCII, subject\-centric harms

## 1Introduction

The creation of non\-consensual intimate imagery \(NCII\) is a concerningly dominant use case of generative AI, whose harms are still inadequately addressed\. In a report released in January 2026, AI Forensics found that the majority use case of the commercial AI system “Grok” is undressing people without their consent\(Bouchaud,[2026](https://arxiv.org/html/2607.18263#bib.bib152)\)\. New York Times and the Center for Countering Digital Hate separately found that after Elon Musk shared an image of an undressed woman superimposed on a SpaceX rocket produced by Grok in late December, Grok was asked to generate 4\.4 million images the following week, compared to roughly 311,000 the week before Musk’s post\(Congeret al\.,[2026](https://arxiv.org/html/2607.18263#bib.bib54)\)\. It is estimated that more than 3 million of these images were sexualized, with at least 23,000 being of children\(Center for Countering Digital Hate,[2026](https://arxiv.org/html/2607.18263#bib.bib53)\)\. Yet, despite the empirical reality that sexual abuse is a primary driver of generative AI usage, and the growing development of applications with the ability to produce AI\-generated non\-consensual intimate imagery \(AIG\-NCII\), the phenomenon is largely absent from extant research involving generative AI imagery\. Instead, the technical interventions developed by the AI/ML community are largely designed for a different set of needs focused on threats to truth and trust\.

While current interventions are designed to address the question of whether a piece of media is authentic or synthetic, the inattention to AIG\-NCII has resulted in oversights in how safety tools are developed\. Drawing on the fundamental human right of dignity, as outlined in the Universal Declaration of Human Rights from theUnited Nations \([1948](https://arxiv.org/html/2607.18263#bib.bib153)\), we argue that interventions should also focus on addressing dignity harms that cause injury to the subject\.

This position paper argues that there is a structural misalignment between AI/ML research agendas and the reality of AI\-generated media, with existing concerns focusing on viewer\-centric epistemic harms, which ignore or even exacerbate subject\-centric dignity harms\.Our contributions in this position paper are as follows:

1. 1\.We surface a disconnect between research motivation and the actual harms of deepfakes\. Through a systematic landscape analysis of highly\-cited works in top\-tier venues between 2020 and 2025, we find scant consideration for cases of AIG\-NCII, despite it likely accounting for more than half of generative AI usage\(Bouchaud,[2026](https://arxiv.org/html/2607.18263#bib.bib152); Center for Countering Digital Hate,[2026](https://arxiv.org/html/2607.18263#bib.bib53); Security Hero,[2023](https://arxiv.org/html/2607.18263#bib.bib151)\)\.
2. 2\.We analyze how this lack of engagement has resulted in the proliferation of “authenticity” tools\. We demonstrate how these tools are insufficient for AIG\-NCII and, in specific deployment contexts, could exacerbate subject\-centric dignity harms\.
3. 3\.Finally, we offer a series of recommendations for researchers and practitioners to realign technical interventions to account for AIG\-NCII\. We stress that all researchers who work in this domain must partner with domain experts in online sexual violence and utilize threat models that consider subject\-centric dignity harms in technical interventions\.

## 2The absence of AIG\-NCII in harm reduction research

A “deepfake” is a colloquial term used to describe AI\-generated or altered content\(Dielet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib57)\)\. This includes deceptive political videos such as that of using the likeness of President Biden to tell voters not to vote in the New Hampshire primary during the 2024 election\(Bond,[2024](https://arxiv.org/html/2607.18263#bib.bib52)\), or fraud content such as using the likeness of Elon Musk to offer an investment opportunity that led to billions of dollars in monetary losses\(Newet al\.,[2024](https://arxiv.org/html/2607.18263#bib.bib55)\)\. A body of technical research has been developed to address deepfakes\. As we show in our analysis, the vast majority of this work has been conducted without acknowledging harms of a sexual nature\. While not all deepfake content is categorized as AIG\-NCII, it still exists as a dominant form of deepfake, with reports estimating that up to 98% of deepfake videos are pornographic in nature\(Security Hero,[2023](https://arxiv.org/html/2607.18263#bib.bib151)\)\. In fact, the term “deepfake” originates from a direct reference to sexual harm, with the term being derived from the username of a Reddit user who shared custom\-made AI\-generated videos depicting actresses performing sexual acts\(Burkell and Gosse,[2019](https://arxiv.org/html/2607.18263#bib.bib40)\)\. Since then, however, the word has entered the global lexicon to refer to a broad range of AI\-generated content, including AI\-generated misinformation, scams, and political images and videos\.

### 2\.1AIG\-NCII as a distinct harm category

We use the acronym AIG\-NCII \(AI\-generated non\-consensual intimate imagery\) to refer to the phenomenon of sexualized deepfake content of a specific individual that exists without their consent\. This can refer to synthetic content that is created with generative AI technology to “nudify” or “undress” a subject without their explicit agreement\(Van der Nagel,[2020](https://arxiv.org/html/2607.18263#bib.bib121)\)111It is important to note here that AIG\-NCII does not require the subject to be fully nude\(Batoolet al\.,[2024](https://arxiv.org/html/2607.18263#bib.bib127)\)\. Recent examples of AIG\-NCII have included content that has attempted to remove hijabs from women\(Tenbarge,[2026](https://arxiv.org/html/2607.18263#bib.bib143)\)\.\. Other terms that describe this same phenomenon include “AI\-NCII”, “AI\-generated image\-based sexual abuse”\(Henryet al\.,[2026](https://arxiv.org/html/2607.18263#bib.bib31)\), synthetic nonconsensual explicit imagery \(SNCEI\)\(Weiet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib79)\), and “deepfake pornography”\(Furizalet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib140)\)\. AIG\-NCII as a concept is derived from traditional non\-consensual intimate imagery \(NCII\)222Non\-consensual intimate imagery \(NCII\) is also colloquially known as “revenge pornography”\.\. AIG\-NCII is highly gendered, with the vast majority of those impacted being women and girls\(Security Hero,[2023](https://arxiv.org/html/2607.18263#bib.bib151); Bouchaud,[2026](https://arxiv.org/html/2607.18263#bib.bib152)\)\. At a societal level, AIG\-NCII represents attempts to silence, de\-platform, and de\-legitimize agency both online and offline\(Maddocks,[2020](https://arxiv.org/html/2607.18263#bib.bib67)\)

From GANs to diffusion models\.The technical barriers to creating AIG\-NCII have lowered precipitously over the last decade\. Initially, from approximately 2017 to 2022, face\-swapping was the primary mechanism used to create AIG\-NCII, by superimposing a face onto another person’s body\. This was enabled by autoencoder architectures, popularized by open\-source repositories like DeepFaceLab\(Perovet al\.,[2020](https://arxiv.org/html/2607.18263#bib.bib11)\)\. The infamous Deepnude application, which “undressed” women in images, relied directly on the Pix2Pix conditional GAN architecture\(Isolaet al\.,[2017](https://arxiv.org/html/2607.18263#bib.bib142); Wanget al\.,[2018](https://arxiv.org/html/2607.18263#bib.bib150)\)\. The current phase of AIG\-NCII creation is driven by diffusion\-based synthesis, enabled by the release of open\-weight latent diffusion models such as Stable Diffusion\(Rombachet al\.,[2022](https://arxiv.org/html/2607.18263#bib.bib23); Schuhmannet al\.,[2022](https://arxiv.org/html/2607.18263#bib.bib146)\)\. Unlike face\-swapping, diffusion models allow for the generation of sexualized imagery via text\-to\-image prompting\. Techniques such as Dreambooth\(Ruizet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib22)\)and Low\-Rank Adaptation \(LoRA\)\(Huet al\.,[2022](https://arxiv.org/html/2607.18263#bib.bib21)\)are now standard tools in AIG\-NCII communities to fine\-tune a specific individual’s likeness with few reference photos\.

Academic research directly contributes to AIG\-NCII\.Hanet al\.\([2025](https://arxiv.org/html/2607.18263#bib.bib47)\)found in an analysis of the online community for Mr\. DeepFakes, a prominent marketplace for deepfake content, that there was a “significant sharing of academic work” on its website, with direct references to GitHub repositories for deepfake tools that cite 43 academic papers\. Many of these deepfake tools are forked from open\-source models directly from research papers, and many of these applications are simply wrappers around open\-source research code\(Hanet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib47); Gibsonet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib45)\)\. Additionally, the use of nude bodies as training data, often collected without consent, gives rise to these capabilities\(Cintaqiaet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib49)\)\.

Limitations of law and policy\.Despite legal prohibitions in the U\.S\. and abroad, enforcement remains insufficient, costly, and reactive, often placing the burden of discovery on the victim\(Sen\. Durbin,[2024](https://arxiv.org/html/2607.18263#bib.bib70); Qiweiet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib77); Congress\.gov,[2025](https://arxiv.org/html/2607.18263#bib.bib69)\)\. For example, while the U\.K\. Online Safety Act criminalizes sharing AIG\-NCII, abuse has merely shifted to non\-compliant platforms and encrypted apps like Telegram\([52](https://arxiv.org/html/2607.18263#bib.bib65)\)\. Similarly, South Korea responded to major scandals\(Gibson,[2019](https://arxiv.org/html/2607.18263#bib.bib63); Bicker,[2020](https://arxiv.org/html/2607.18263#bib.bib61)\)by criminalizing possession\(Yim,[2024](https://arxiv.org/html/2607.18263#bib.bib60)\), yet deepfake generation tools remain accessible\. Moderation efforts on individual platforms face a similar “whac\-a\-mole” dynamic\(Dinget al\.,[2026](https://arxiv.org/html/2607.18263#bib.bib56)\)\. When CivitAI banned real\-person likeness models\(CivitAI,[2025](https://arxiv.org/html/2607.18263#bib.bib59); Maiberg,[2025a](https://arxiv.org/html/2607.18263#bib.bib80); Wagner and Cetinic,[2025](https://arxiv.org/html/2607.18263#bib.bib78)\), the models simply migrated to HuggingFace\(Maiberg,[2025b](https://arxiv.org/html/2607.18263#bib.bib58)\)\.

### 2\.2Landscape analysis of existing literature

To quantify the misalignment between existing research concerns and the reality of AIG\-NCII, we conducted an analysis of technical defense papers published between 2020 and 2025\. Our analysis reveals that the literature addressing deepfake harms almost entirely ignores AIG\-NCII\.

Methodology\.Our goal was to locate works that aimed to address harms regarding deepfakes or synthetic media more broadly\. We queried Google Scholar for papers containing a specific set of keywords, using the following query:\(“detection” OR “detector” OR “forensics” OR “recognition” OR “watermark”\) AND \(“deepfake” OR “synthetic image” OR “fake image” OR “diffusion”\)\. This initial search criteria yielded 965 papers\. Next, we filtered our results to papers published only at the top\-tier venues:“CVPR” OR “ICCV” OR “ECCV” OR “NeurIPS” OR “ICML” OR “ICLR”\. We excluded workshop papers and included arXiv preprints that were returned in our search\. Papers with more than 80 citations from other venues were also included\. This brought our resulting dataset to 379 papers\. Finally, we filtered the results down to the top 100 most cited papers, and manually excluded papers where diffusion models were utilized for unrelated computer vision tasks such as detecting tumors, detecting cars, or detecting cracks in steel\. We also excluded one retracted paper\. This process left a final dataset of 39 papers that we qualitatively analyzed\. We examined each paper for engagement with AIG\-NCII, with particular attention to the usage of the following terms:“non\-consensual intimate imagery” “NCII”, “revenge porn”,“sexual violence” “porn”, “nudity”, “undress”,“obscene”\. See Table 2 in the Appendix for the final list of 39 papers\.

Results\.We categorized the 39 papers into three tiers of engagement with AIG\-NCII\.

1. 1\.No mention \(34 papers\):The paper frames the problem exclusively as misinformation, fraud, or technical artifact detection\.
2. 2\.Mention only \(5 papers\):The authors reference AIG\-NCII terms in passing, such as in the introduction or broad impact statement, but the proposed technical method remains generic\.
3. 3\.Technical implementation \(0 papers\):The authors design an intervention with a threat model specific to AIG\-NCII\.

Our analysis affirms existing research stating that deepfake research is overwhelmingly motivated by harms related to trust, fraud, and political misinformation\(Rini,[2020](https://arxiv.org/html/2607.18263#bib.bib44); Harris,[2021](https://arxiv.org/html/2607.18263#bib.bib141)\)\.

![Refer to caption](https://arxiv.org/html/2607.18263v1/0_figures/comparison.png)Figure 1:Proportion of AI\-generated content that is classified as AIG\-NCII\(Bouchaud,[2026](https://arxiv.org/html/2607.18263#bib.bib152)\)compared to the proportion of technical defense papers that merely mention AIG\-NCII\. No papers found in our landscape analysis meaningfully engaged with AIG\-NCII\.

## 3Review of authenticity\-based interventions

In order to deal with harms pertaining to truth, our analysis found that the AI/ML community has coalesced around a series of tools designed to distinguishsyntheticfromauthenticmedia\. We categorize these technical interventions into three primary paradigms: detection, provenance, and watermarking\. While they differ significantly in their implementation, they share one foundational assumption, thattruth\-verification, or authenticity, is the primary proxy for safety\.In this section, we review these three paradigms and note the assumptions that underpin their design\.

### 3\.1Detection

Current AI detection attempts to approximate a classification problem, wherein a model learns a decision boundary between the distribution of authentic media and synthetic media\. The primary assumption of this paradigm is that a robust decision boundary exists, and that successfully determining the boundary between authentic and synthetic is a sufficient condition for addressing harm\.

While earlier literature focused on identifying GAN\-specific artifacts in the media asset\(Franket al\.,[2020](https://arxiv.org/html/2607.18263#bib.bib37)\), the proliferation of diffusion models has required a fundamental shift in the feature extraction process for identifying synthetic content\. More recent detection methods target the unique fingerprints introduced by the iterative de\-noising process of diffusion models\. For example, DIRE\(Wanget al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib36)\)utilizes the observation that diffusion\-generated synthetic images show lower reconstruction error inverted through a pre\-trained diffusion process as compared to authentic images\. Similarly,Corviet al\.\([2023](https://arxiv.org/html/2607.18263#bib.bib39)\)identified and analyzed distinct spectral traces left by the Gaussian noise scheduling that are inherent to latent diffusion models\. To address the rapid evolution of generator architectures \(e\.g\., Stable Diffusion\(Rombachet al\.,[2022](https://arxiv.org/html/2607.18263#bib.bib23)\)and FLUX\(Black Forest Labs,[2024](https://arxiv.org/html/2607.18263#bib.bib149)\)\), research has increasingly moved utilizing feature spaces from foundational vision\-language models like CLIP\(Radfordet al\.,[2021](https://arxiv.org/html/2607.18263#bib.bib86)\)to identify synthetic semantic patterns that generalize across architectures\(Ojhaet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib35)\)\.

### 3\.2Provenance

Provenance methods, as defined by theCoalition for Content Provenance and Authenticity \(C2PA\) \([2026](https://arxiv.org/html/2607.18263#bib.bib137)\)technical specification, attempt to establish a history of modifications for a media asset\. Unlike methods in detection, this paradigm for protection relies on cryptographically verifiable information that can be used to verify that an asset is free from tampering\. At each point of asset creation or modification, a digital signature is bound to a hash of the pixel data of the asset, along with a manifest containing metadata assertions, such as content ownership and timestamp\(Rosenthol,[2022](https://arxiv.org/html/2607.18263#bib.bib32)\)\. The assumption guiding provenance methods is that having a verifiable chain\-of\-custody resolves trust in where a piece of media originated, and whether it was edited along the way\. In other words, interventions following the provenance paradigm are concerned with tracking the lineage from a source asset to any of its modified outputs, such that the lineage itself is an indicator of authenticity\.

### 3\.3Watermarking

Content watermarking techniques aim to embed signals invisible to the human eye directly onto the media content at generation, to signal that content as synthetic\. Unlike metadata, which can be easily stripped and removed from the asset, watermarks aim to be more robust against being removed, even with cropping, filters, and other image transformations\. Approaches within this paradigm include latent watermarking\(Fernandezet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib34)\)and sampling\-based watermarking\(Wenet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib33)\)\. Industry implementations, such as Google’s SynthID\(DeepMind,[2025](https://arxiv.org/html/2607.18263#bib.bib85)\), embeds signals onto the media that are detectable only when paired with a specialized detector model or decoding mechanism\. Watermarking is often used to help with post\-hoc detection\.

## 4Consequences of ignoring AIG\-NCII

The intervention paradigms around detection, provenance, and watermarking are meant to restore the viewer’s ability to discern what is authentic\. For AIG\-NCII cases, however, the subject suffers from violation of their dignity regardless of authenticity\. In this section, we discuss possible consequences overlooking these dignity violations\.

### 4\.1Neglect of subject\-centric dignity harms

![Refer to caption](https://arxiv.org/html/2607.18263v1/0_figures/harm-types.png)Figure 2:Examples of deepfakes that cause harm to the viewer, subject, or both parties\.The exclusion of AIG\-NCII in deepfake harm reduction research has resulted in an ecosystem that solely focuses on epistemic harms, or harms relating to truth and authenticity\. However, asHarris \([2021](https://arxiv.org/html/2607.18263#bib.bib141)\)notes, these concerns often overlook the severity of harms to subjects\. FollowingChesney and Citron \([2019](https://arxiv.org/html/2607.18263#bib.bib41)\)andCitron \([2018](https://arxiv.org/html/2607.18263#bib.bib42)\)and drawing from the UN Universal Declaration of Human Rights\(United Nations,[1948](https://arxiv.org/html/2607.18263#bib.bib153)\), we argue that AIG\-NCII violates the fundamental human right of dignity for deepfaked subjects, which is distinct from epistemic harms to viewers\. We build upon the framework byOlteanuet al\.\([2025](https://arxiv.org/html/2607.18263#bib.bib48)\)to distinguish between harms to the viewer and the subject:

Viewer\-centric harmsencompass both individual deception and broader societal epistemic degradation\. On an individual level, this includes fraud, such as scams mimicking a friend’s likeness\. However, asRini \([2020](https://arxiv.org/html/2607.18263#bib.bib44)\)argues, this actually presents a societal harm due to the erosion of authenticity and the destabilization of shared reality\(Chesney and Citron,[2019](https://arxiv.org/html/2607.18263#bib.bib41)\)\. If one was led to believe false information, they may have suffered a viewer\-centric epistemic harm\.

Subject\-centric harmsoccur when a person’s likeness is utilized without consent, regardless of whether the viewer is deceived\. Drawing onCitron \([2018](https://arxiv.org/html/2607.18263#bib.bib42)\)’s idea of sexual privacy, the harm arises from the undignified presentation of the subject and the resulting loss of autonomy\. At a high level, a dignity harm occurs when one’s likeness is used in ways they did notconsentto\. We highlight two instances of this harm that map onto specific AI/ML techniques\. The first is non\-consensual identity preservation, in which a model retains a recognizable likeness of a specific person without their consent\. Defenses against this mode may aim to prevent identity retention—for example, by disrupting a model’s ability to learn or reproduce a specific subject’s features\(Van Leet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib123)\)\. The second is non\-consensual modification, in which the subject’s likeness is intentionally preserved, but the image is altered in ways they did not consent to, such as removing clothing or face\-swapping with pornographic videos\.

Viewer\-centric and subject\-centric harms are orthogonal, as shown in Figure[2](https://arxiv.org/html/2607.18263#S4.F2)\. While these harms often overlap—for instance, a political deepfake may deceive constituents \(viewer\-centric\) while damaging the politician’s reputation \(subject\-centric\)—the mechanisms required to address them differ\. The lack of engagement with subject\-centric harms has led to a lapse in defense design, where merely labeling content as “synthetic” does not mitigate subject\-centric harms\. The misalignment between technical capabilities and safety needs is so acute that even platform governance bodies have begun questioning use of these tools\. The Oversight Board for Meta recently noted that “labeling manipulated content is not appropriate in this instance because the harms stem from the sharing and viewing of these images—and not solely from misleading people about their authenticity”\(Oversight Board,[2024](https://arxiv.org/html/2607.18263#bib.bib138)\)\. Indeed, authenticity\-based interventions do not mitigate harms to subjects, especially given that prominent online forums that host AIG\-NCII already routinely label the content as fake\(Hanet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib47)\)\.

### 4\.2Authentic≠\\neqsafe

The current trajectory of deepfake defense research optimizes for authenticity metrics\. When applied to AIG\-NCII, this creates a fundamental category error, treating authenticity as a proxy for safety\. This substitution fails because, unlike political misinformation where falsehood is the primary harm, sexual violence is defined by the absence of consent, which is violated regardless of whether an image is authentic or synthetic\. As illustrated in Table[1](https://arxiv.org/html/2607.18263#S4.T1), the axis of artificiality \(what existing interventions measure\) is orthogonal to the axis of consent \(what determines safety\)\. Relying on authenticity\-based tools creates a blunt instrument that cannot address the case of AIG\-NCII, because it conflates non\-consensual synthetic imagery with that of consensual synthetic imagery\. At the same time, the ecosystem risks building over\-censoring tools that stifle legitimate expression while failing to address traditional non\-consensual imagery simply because it is “authentic”\. We argue that until technical interventions can account for this orthogonality, evidence of authenticity should not be treated as sufficient evidence of safety\.

Table 1:The orthogonality of harm\. The current detection paradigm can only distinguish between synthetic and authentic media\. However, safety falls along an orthogonal axis that is determined by consent\.
### 4\.3Potential misuse of authenticity tools

The implicit assumption guiding current research agendas is that safety can be achieved once the most accurate authenticity model is developed\. We challenge this assumption\. In the following contexts, we demonstrate how authenticity interventions, when built without consideration for a subject\-centric threat model, may actually exacerbate harms to subject dignity\.

#### When used by online platforms\.

Current regulations, such as the EU AI Act\(ArtificialIntelligenceAct\.eu,[2025](https://arxiv.org/html/2607.18263#bib.bib14)\), as well as platform policies on Meta\(Meta Platforms,[2024](https://arxiv.org/html/2607.18263#bib.bib15)\)and TikTok\(TikTok,[2026](https://arxiv.org/html/2607.18263#bib.bib16)\), prioritize the labeling of synthetic media that is identified\. While effective for viewer\-centric epistemic harms, this model of public labeling could backfire for AIG\-NCII\. A label stating that an image is made with AI does not address the fact that the image is shared on online platforms\. While some platforms explicitly ban nudity, AIG\-NCII also includes cases where the subject is not fully nude, which is ignored by these regulations\(Batoolet al\.,[2024](https://arxiv.org/html/2607.18263#bib.bib127)\)\. When labeling content is prioritized over its removal, this could create a perverse outcome where the abuser may be protected from moderation consequences so long as they are transparent about the synthetic, AI\-generated nature of the image\.

#### When used by abusers\.

Research in online sexual violence indicates that perpetrators are often driven by an assertion of power over a victim, rather than by sexual gratification on its own\(Henry and Beard,[2024](https://arxiv.org/html/2607.18263#bib.bib19); Henry and Powell,[2016](https://arxiv.org/html/2607.18263#bib.bib18)\)\. In fact,Mariniet al\.\([2024](https://arxiv.org/html/2607.18263#bib.bib17)\)has shown that people are less aroused when they find that an image is identified as AI\-generated\. As noted byMassanari \([2017](https://arxiv.org/html/2607.18263#bib.bib66)\), these communities are not passive consumers but active participants who “demonstrate technological prowess” in aggregating disparate pieces of content to target and verify the identity of victims\. These same communities may use authenticity\-based tools to locate content in order to further harass and dox victim\-survivors\. In this context, authenticity verification may become a tool that allows users to sort through authentic and synthetic imagery for abuse\.

#### When used against NCII victim\-survivors\.

Finally, the utility of authenticity labels is not consistent for victim\-survivors, fluctuating depending on the harm that the subject is dealing with\. For victim\-survivors of traditional NCII, plausible deniability may actually offer a safety mechanism to protect their dignity\. In this context, ambiguity offers a form of protective cover for victim\-survivors\. If a detection system definitively labels traditional NCII as “authentic”, it inadvertently acts as a verification tool for abusers, confirming the victim’s exposure to the public\. By removing this uncertainty, technical interventions may out victims who could otherwise have maintained some degree of social safety by casting doubt on the image’s veracity\.

## 5Recommendations

Addressing AIG\-NCII is an extremely difficult task, and it is one that faces significant barriers both ethically and technically\. Given the sensitive nature of this abuse, we cannot always propose definitive solutions\. In this section, we raise recommendations to best address the problem\. Ultimately, we believe there is a need for both technical and social pathways for addressing the proliferation of AIG\-NCII, and this requires coordination between AI/ML researchers and practitioners, social scientists, policy makers, sexual violence prevention experts, and victim\-survivor advocates\.

R1\. Decouple epistemic and dignity harms\.As we have shown, care must be taken when applying authenticity markers\. If a system identifies potential AIG\-NCII, it should not apply a public label, as this keeps content visible, and may exacerbate harm\. Instead, detection should serve as a backend flag that triggers precautionary handling \(suppression or triage\) treating AIG\-NCII like traditional NCII when content depicts real people in sexualized contexts\. We urge the research community to study the specific trade\-offs of labeling AIG\-NCII before deploying detection tools\. Future work should identify how to verify content in a way that protects the privacy of victims while empowering deepfake subjects to defend themselves\. Furthermore, we must weigh the asymmetrical harm of errors in labeling\. While temporarily restricting consensual content is often reversible, failing to intervene in situations of sexual violence imposes irreversible harms\. Future work must rigorously evaluate how transparency standards designed for misinformation may inadvertently endanger privacy\.

R2\. Elevate subject\-centric dignity harms\.The violation of dignity must be elevated to the same tier as political misinformation\. Researchers should mirror the shift in computer security towards analyzing intimate partner violence \(IPV\), where the adversary is personally known to the victim\-survivor\(Chatterjeeet al\.,[2018](https://arxiv.org/html/2607.18263#bib.bib135); Havronet al\.,[2019](https://arxiv.org/html/2607.18263#bib.bib134); Freedet al\.,[2018](https://arxiv.org/html/2607.18263#bib.bib10)\)\. In AIG\-NCII, the adversary is often an actor equipped with limited reference photos and parameter\-efficient fine\-tuning techniques\. Under this framework, the goal shifts from maximizing detection accuracy to minimizing identity preservation\. Defenses are successful if they prevent the reproduction of a specific identity or make it more challenging\. Furthermore, we must reject the assumption that publicly available data is synonymous with consensual data\. Privacy is violated by the migration of information outside its intended context\(Nissenbaum,[2004](https://arxiv.org/html/2607.18263#bib.bib132)\)\. While modeling general human features may be necessary in some contexts, the non\-consensual ingestion of nude or sexualized bodies crosses an ethical red line\(Stark,[2018](https://arxiv.org/html/2607.18263#bib.bib131); Scheuermanet al\.,[2021](https://arxiv.org/html/2607.18263#bib.bib130)\)\. Researchers should abandon the unethical curation of datasets from sensitive domains\(Cintaqiaet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib49)\)\.

R3\. Restrict high\-risk research assets\.We call for a re\-evaluation of open\-release norms for architectures explicitly optimized for high\-fidelity identity retention and inpainting\. While open science is a core value of the field, an emerging consensus in AI Safety literature recognizes that the risks of open\-sourcing highly capable, dual\-use models often outweigh the benefits\(Widderet al\.,[2022](https://arxiv.org/html/2607.18263#bib.bib83); Solaiman,[2023](https://arxiv.org/html/2607.18263#bib.bib145); Segeret al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib144)\)\. Models and fine\-tuning techniques that demonstrate state\-of\-the\-art performance \(such as “cloning” the likeness of a person from a few photographs\) should be subject to gated or researcher\-only access\. By placing additional friction on the tools of creation, we can reduce the size of the threat downstream\.

R4\. Consider proactive prevention\.The field should move beyond post\-hoc detection and engage seriously with adversarial defense mechanisms against inpainting, one of the primary techniques for nudification\. Adversarial immunization introduces human\-imperceptible perturbations that disrupt internal representations in generative models, preventing style mimicry or inpainting\(Jeonet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib119); Guoet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib129); Van Leet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib123); Shanet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib128)\)\. While these defenses can be brittle against transformations\(Höniget al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib2)\), recent work may be closing this gap by making perturbations harder to remove\(Kimet al\.,[2026](https://arxiv.org/html/2607.18263#bib.bib9)\)and dramatically lower the per\-image protection cost\(Ozdenet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib8)\)\. Even imperfect defenses raise the cost for lay\-abusers and disrupt the generative pipelines\.

R5\. Develop safety\-aligned metrics\.Success must be measured by harm reduction, not merely detection accuracy\. We need metrics that verify whether systems actually reduce the prevalence of abuse\. This requires creative evaluation strategies\. For instance,Cretuet al\.\([2025](https://arxiv.org/html/2607.18263#bib.bib133)\)utilized ethical proxies to assess child\-safety filters without generating actual CSAM\. This may offer inspiration for evaluating AIG\-NCII interventions without producing harmful content\.

R6\. Integrate AIG\-NCII into AI Safety\.Subject\-centric harms must be elevated to a core concern within the definition of AI Safety\. Scholars have increasingly critiqued the mainstream discourse for focusing on existential risks while excluding present\-day harms\(Gyevnar and Kasirzadeh,[2025](https://arxiv.org/html/2607.18263#bib.bib76); Ahmedet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib73); Hazraet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib62)\)\. We further note that the dominant model\-side safety techniques \(e\.g\., concept erasure and NSFW filtering\(Gandikotaet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib3); Schramowskiet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib4)\)\) are insufficient for AIG\-NCII because they assume a cooperative model operator\. The AIG\-NCII ecosystem is defined by open models and non\-compliant platforms\(Gibsonet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib45)\)\. AIG\-NCII therefore requires safety research that does not depend on operator goodwill\. If the field is to meet its goal of protecting human welfare, the scope of AI Safety must expand to include the immediate violence of AIG\-NCII\. Conference organizers should recognize the scale of this abuse and allocate the same rigor and resources currently afforded to sub\-fields such as algorithmic bias and toxic content generation\(Weidingeret al\.,[2021](https://arxiv.org/html/2607.18263#bib.bib125)\)\.

R7\. Establish ethical partnerships and guardrails\.Research into AIG\-NCII poses unique ethical and psychological challenges and is not suitable for all researchers\. Guardrails must be implemented to mitigate secondary traumatic stress for those who review sensitive content\(Williamsonet al\.,[2020](https://arxiv.org/html/2607.18263#bib.bib75)\)\. Crucially, technical researchers must avoid speculatively deriving threat models in isolation\. Work must be empirically grounded and co\-designed with domain experts in online sexual violence and victim advocacy\(Costanza\-Chock,[2020](https://arxiv.org/html/2607.18263#bib.bib124)\)\. These experts should be integrated as partners during the initial design phase, rather than merely consulted for post\-hoc validation\.

R8\. Account for victim\-survivor plurality\.Interventions must respect the spectrum of survivor needs\. As noted byMcGlynn and Westmarland \([2019](https://arxiv.org/html/2607.18263#bib.bib120)\), victim\-survivors may prioritize drastically different outcomes, ranging from criminal prosecution to content removal\. A one\-size\-fits\-all technical solution cannot serve all these needs\. Researchers should leverage frameworks of restorative justice to design flexible tools that respect survivor agency, rather than imposing a monolithic technical “solution”\(Schoenebecket al\.,[2021](https://arxiv.org/html/2607.18263#bib.bib122)\)\.

R9\. Acknowledge that social harms require social interventions for remediation\.While existing interventions are focused on detecting and identifying the synthetic nature of deepfakes, this does not actually resolve any of the dignity\-based harms that have been inflicted on victim\-survivors\. Experts in interpersonal violence \(IPV\) and sexual violence reduction should be consulted when developing possible interventions to remediate the harms of AIG\-NCII\.

## 6Alternative Views

In this section, we address the primary objections to our arguments that AI/ML research must account for AIG\-NCII\.

AV1: Addressing social harms is the domain of law and policy, not AI/ML\.One objection to this paper’s argument is that AI/ML researchers are only responsible for optimizing technical capabilities, while the regulation of those capabilities belongs to policymakers\. From this perspective, AIG\-NCII is fundamentally a legal problem that arises from under\-resourced legal systems and platforms that fail to enforce abuse\. How generative AI tools are used is out of scope for AI/ML researchers\.

Response:We agree that law and policy are essential, but we reject the premise that research into technical protections is therefore irrelevant\. First, the timing mismatch between AI development and legal enforcement makes relying solely on the law unfeasible\. The legal system operates on timescales of years, and generative AI capabilities evolve in weeks\. In the example of the UK Online Safety Act, by the time the legislation passed, the dominant mechanism for AIG\-NCII had already shifted from face\-swapping to diffusion synthesis\. The decision to work on defense methods against AIG\-NCII is as much a research choice as it is to work on other technical innovations such as LoRA\. Much like how the fight against Child Sexual Abuse Material \(CSAM\) has involved legal, technical, and advocate coordination, we argue that technical protections can work in tandem with legal efforts against AIG\-NCII, which requires a multi\-pronged approach\.

AV2: Prevention of AIG\-NCII is technically intractable\.Even if researchers accept responsibility, proposed interventions simply do not work\. Adversarial defenses such as inpainting perturbation are brittle and easily defeated by compression or model updates\(Sunet al\.,[2023](https://arxiv.org/html/2607.18263#bib.bib88); Guoet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib129); Goodfellowet al\.,[2014](https://arxiv.org/html/2607.18263#bib.bib148); Athalyeet al\.,[2018](https://arxiv.org/html/2607.18263#bib.bib147); Höniget al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib2)\)\. Furthermore, the sheer scale of the internet makes protecting every user’s photo impossible, and sophisticated abusers can easily bypass protections using slightly altered inputs\. Therefore, proposing “better defenses” offers false hope to victim\-survivors\.

Response:We concede that no technical intervention will ever be a perfect shield\. However, the goal of technical defense is not perfection, but friction\. Currently, an abuser can generate AIG\-NCII in minutes with little to no cost or technical skill\. Adversarial defenses, even if imperfect, raise the cost of abuse by requiring knowledge to bypass\. This friction may disrupt casual abuser \(teenagers, ex\-partners\) who account for a significant volume of harassment but lack the sophistication to break these defenses\. Additionally, while current defenses are brittle, this should be a call for further research, not abandonment\.

AV3: The cultural and institutional incentives of AI/ML research makes engagement with AIG\-NCII unrealistic\.Even if researchers accept the framing of subject\-centric harms, there are structural barriers to working on such topics\. Lack of institutional support, concerns about the ability to publish, and reviewer discomfort with sexualized topics makes it difficult for researchers to engage with AIG\-NCII\.

Response:We acknowledge these incentive structures, but argue that the field cannot claim AIG\-NCII is too peripheral to engage with when it is a direct downstream product of mainstream research decisions\. The open datasets used to train widely\-deployed diffusion models contained non\-consensual images of human bodies, including CSAM\(Thiel,[2023](https://arxiv.org/html/2607.18263#bib.bib7)\)\. The techniques now used for AIG\-NCII \(e\.g\., face\-swapping, inpainting, fine\-tuning on specific individuals\) were each developed and celebrated at top venues\. Having produced the capabilities, the field also bears responsibility for the mitigating their negative societal impacts\. The field has shifted before\. Algorithmic fairness was a fringe topic a decade ago\. It took several key papers to point to the problem\(Buolamwini and Gebru,[2018](https://arxiv.org/html/2607.18263#bib.bib6); Angwinet al\.,[2022](https://arxiv.org/html/2607.18263#bib.bib5)\)and efforts to form dedicated conferences \(FAccT and AIES, both in 2018\)\. The transition required individual researchers willing to legitimize the topic, conference organizers willing to allocate space, and senior scholars willing to advise students working on it\. None of these moves required the field tofirstbe comfortable\. In fact, comfort followed the work, not the other way around\. AIG\-NCII can follow a similar trajectory if researchers refuse to accept “taboo” as a reason for inattention\.

## 7Conclusion

In this position paper, we show that there is a fundamental misalignment between the current technical interventions for deepfakes that address viewer\-centric epistemic harms and the prevailing reality of subject\-centric dignity harms in the form of AIG\-NCII, which account for the majority of generative AI usage\(Center for Countering Digital Hate,[2026](https://arxiv.org/html/2607.18263#bib.bib53); Bouchaud,[2026](https://arxiv.org/html/2607.18263#bib.bib152); Security Hero,[2023](https://arxiv.org/html/2607.18263#bib.bib151); Gibsonet al\.,[2025](https://arxiv.org/html/2607.18263#bib.bib45)\)\. We urge the AI/ML community to realign its priorities to address these harms, else we risk exacerbating harms to victims of AIG\-NCII\. At the same time, we offer our call to action in the form of recommendations with a necessary constraint\. Research into protections against AIG\-NCII should only be undertaken when adequate safety protocols are established, including mitigating harm for researchers and establishing substantive partnerships with victim\-advocates and sexual violence prevention experts\. Ultimately, the research community bears a responsibility to ensure that our definitions of AI safety protect not only truth, but also the dignity of people\.

## Acknowledgments

We thank Rosanna Bellini for conversations that helped refine the framing of this work\. This material is based on works supported by the National Science Foundation under Grant 2311102\.

## References

- A\. Aghasanli, D\. Kangin, and P\. Angelov \(2023\)Interpretable\-through\-prototypes deepfake detection for diffusion models\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 467–474\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.26.20.1.1.1)\.
- M\. Ahmadi, A\. Norouzi, N\. Karimi, S\. Samavi, and A\. Emami \(2020\)ReDMark: framework for residual diffusion watermarking based on deep networks\.Expert Systems with Applications146,pp\. 113157\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.10.4.1.1.1)\.
- S\. Ahmed, K\. Jaźwińska, A\. Ahlawat, A\. Winecoff, and M\. Wang \(2023\)Building the epistemic community of ai safety\.Available at SSRN 4641526\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p7.1)\.
- J\. Angwin, J\. Larson, S\. Mattu, and L\. Kirchner \(2022\)Machine bias\.InEthics of data and analytics,pp\. 254–264\.Cited by:[§6](https://arxiv.org/html/2607.18263#S6.p7.1)\.
- ArtificialIntelligenceAct\.eu \(2025\)Article 50: transparency obligations for providers and deployers of certain ai systems\.Note:European Union Artificial Intelligence Act transparency obligations \(Article 50\)[https://artificialintelligenceact\.eu/article/50/](https://artificialintelligenceact.eu/article/50/)External Links:[Link](https://artificialintelligenceact.eu/article/50/)Cited by:[§4\.3](https://arxiv.org/html/2607.18263#S4.SS3.SSS0.Px1.p1.1)\.
- A\. Athalye, N\. Carlini, and D\. Wagner \(2018\)Obfuscated gradients give a false sense of security: circumventing defenses to adversarial examples\.InInternational conference on machine learning,pp\. 274–283\.Cited by:[§6](https://arxiv.org/html/2607.18263#S6.p4.1)\.
- Q\. Bammey \(2023\)Synthbuster: towards detection of diffusion model generated images\.IEEE Open Journal of Signal Processing5,pp\. 1–9\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.16.10.1.1.1)\.
- A\. Batool, M\. Naseem, and K\. Toyama \(2024\)Expanding concepts of non\-consensual image\-disclosure abuse: a study of ncida in pakistan\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,pp\. 1–17\.Cited by:[§4\.3](https://arxiv.org/html/2607.18263#S4.SS3.SSS0.Px1.p1.1),[footnote 1](https://arxiv.org/html/2607.18263#footnote1)\.
- L\. Bicker \(2020\)Cho Ju\-bin: South Korea chatroom sex abuse suspect named after outcry\.\(en\-GB\)\.External Links:[Link](https://www.bbc.com/news/world-asia-52030219)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- Black Forest Labs \(2024\)Introducing FLUX\.1 tools\.Note:[https://bfl\.ai/blog/24\-11\-21\-tools](https://bfl.ai/blog/24-11-21-tools)Accessed: 2026\-01\-28Cited by:[§3\.1](https://arxiv.org/html/2607.18263#S3.SS1.p2.1)\.
- S\. Bond \(2024\)How AI deepfakes polluted elections in 2024\.NPR\.Cited by:[§2](https://arxiv.org/html/2607.18263#S2.p1.1)\.
- P\. Bouchaud \(2026\)Grok Unleashed\.Technical reportAI Forensics\.Cited by:[item 1](https://arxiv.org/html/2607.18263#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2607.18263#S1.p1.1),[Figure 1](https://arxiv.org/html/2607.18263#S2.F1),[Figure 1](https://arxiv.org/html/2607.18263#S2.F1.3.2),[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p1.1),[§7](https://arxiv.org/html/2607.18263#S7.p1.1)\.
- J\. Buolamwini and T\. Gebru \(2018\)Gender shades: intersectional accuracy disparities in commercial gender classification\.InConference on fairness, accountability and transparency,pp\. 77–91\.Cited by:[§6](https://arxiv.org/html/2607.18263#S6.p7.1)\.
- J\. Burkell and C\. Gosse \(2019\)Nothing new here: emphasizing the social and cultural context of deepfakes\.First Monday\.Cited by:[§2](https://arxiv.org/html/2607.18263#S2.p1.1)\.
- Center for Countering Digital Hate \(2026\)Grok floods X with sexualized images of women and children\.Technical reportCenter for Countering Digital Hate\.Cited by:[item 1](https://arxiv.org/html/2607.18263#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2607.18263#S1.p1.1),[§7](https://arxiv.org/html/2607.18263#S7.p1.1)\.
- R\. Chatterjee, P\. Doerfler, H\. Orgad, S\. Havron, J\. Palmer, D\. Freed, K\. Levy, N\. Dell, D\. McCoy, and T\. Ristenpart \(2018\)The spyware used in intimate partner violence\.In2018 IEEE Symposium on Security and Privacy \(SP\),pp\. 441–458\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p3.1)\.
- B\. Chen, J\. Zeng, J\. Yang, and R\. Yang \(2024\)Drct: diffusion reconstruction contrastive training towards universal detection of diffusion generated images\.InForty\-first International Conference on Machine Learning,Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.17.11.1.1.1)\.
- B\. Chesney and D\. Citron \(2019\)Deep fakes: a looming challenge for privacy, democracy, and national security\.Calif\. L\. Rev\.107,pp\. 1753\.Cited by:[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p2.1)\.
- P\. Cintaqia, A\. Arya, E\. M\. Redmiles, D\. Kumar, A\. McDonald, and L\. Qin \(2025\)Stop the nonconsensual use of nude images in research\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,Vol\.8,pp\. 628–629\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p3.1),[§5](https://arxiv.org/html/2607.18263#S5.p3.1)\.
- D\. K\. Citron \(2018\)Sexual privacy\.Yale LJ128,pp\. 1870\.Cited by:[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p3.1)\.
- CivitAI \(2025\)Policy Update: Removal of Real\-Person Likeness Content \| Civitai\.\(en\)\.External Links:[Link](https://civitai.com/articles/15022/policy-update-removal-of-real-person-likeness-content)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- Coalition for Content Provenance and Authenticity \(C2PA\) \(2026\)C2PA technical specification, version 2\.3\.Note:Technical standard for Content Credentials \(provenance and authenticity metadata\)[https://spec\.c2pa\.org/specifications/specifications/2\.3/index\.html](https://spec.c2pa.org/specifications/specifications/2.3/index.html)External Links:[Link](https://spec.c2pa.org/specifications/specifications/2.3/index.html)Cited by:[§3\.2](https://arxiv.org/html/2607.18263#S3.SS2.p1.1)\.
- K\. Conger, D\. Freedman, and S\. A\. Thompson \(2026\)Musk’s Chatbot Flooded X With Millions of Sexualized Images in Days, New Estimates Show\.The New York Times\.External Links:ISSN 0362\-4331Cited by:[§1](https://arxiv.org/html/2607.18263#S1.p1.1)\.
- Congress\.gov \(2025\)S\.146 \- 119th Congress \(2025\-2026\): TAKE IT DOWN Act \| Congress\.gov \| Library of Congress\.External Links:[Link](https://www.congress.gov/bill/119th-congress/senate-bill/146)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- R\. Corvi, D\. Cozzolino, G\. Zingarini, G\. Poggi, K\. Nagano, and L\. Verdoliva \(2023\)On the detection of synthetic images generated by diffusion models\.InICASSP 2023\-2023 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 1–5\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.9.3.1.1.1),[§3\.1](https://arxiv.org/html/2607.18263#S3.SS1.p2.1)\.
- S\. Costanza\-Chock \(2020\)Design justice: community\-led practices to build the worlds we need\.The MIT Press\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p8.1)\.
- A\. Cretu, K\. Kireev, A\. Abdalla, W\. Obinna, R\. Meier, S\. A\. Bargal, E\. M\. Redmiles, and C\. Troncoso \(2025\)Evaluating concept filtering defenses against child sexual abuse material generation by text\-to\-image models\.arXiv preprint arXiv:2512\.05707\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p6.1)\.
- N\. Damer, M\. Fang, P\. Siebke, J\. N\. Kolf, M\. Huber, and F\. Boutros \(2023\)Mordiff: recognition vulnerability and attack detectability of face morphing attacks created by diffusion autoencoders\.arXiv preprint arXiv:2302\.01843\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.23.17.1.1.1)\.
- DeepMind \(2025\)SynthID: a tool to watermark and identify content generated through ai\.Note:[https://deepmind\.google/models/synthid/](https://deepmind.google/models/synthid/)Accessed: 2026\-01\-XXCited by:[§3\.3](https://arxiv.org/html/2607.18263#S3.SS3.p1.1)\.
- A\. Diel, T\. Lalgi, F\. S\. Mellis, A\. Teufel, and A\. Bäuerle \(2025\)The harm of deepfakes: a scoping review of deepfakes’ negative effects on human mind and behavior\.AI & SOCIETY\.External Links:ISSN 0951\-5666, 1435\-5655,[Document](https://dx.doi.org/10.1007/s00146-025-02774-0)Cited by:[§2](https://arxiv.org/html/2607.18263#S2.p1.1)\.
- M\. L\. Ding, H\. Suresh, and S\. Venkatasubramanian \(2026\)How to stop playing whack\-a\-mole: mapping the ecosystem of technologies facilitating ai\-generated non\-consensual intimate images\.arXiv preprint arXiv:2602\.04759\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- P\. Fernandez, G\. Couairon, H\. Jégou, M\. Douze, and T\. Furon \(2023\)The stable signature: rooting watermarks in latent diffusion models\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 22466–22477\.Cited by:[§3\.3](https://arxiv.org/html/2607.18263#S3.SS3.p1.1)\.
- J\. Frank, T\. Eisenhofer, L\. Schönherr, A\. Fischer, D\. Kolossa, and T\. Holz \(2020\)Leveraging frequency analysis for deep fake image recognition\.InInternational conference on machine learning,pp\. 3247–3258\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.7.1.1.1.1),[§3\.1](https://arxiv.org/html/2607.18263#S3.SS1.p2.1)\.
- D\. Freed, J\. Palmer, D\. Minchala, K\. Levy, T\. Ristenpart, and N\. Dell \(2018\)“A stalker’s paradise” how intimate partner abusers exploit technology\.InProceedings of the 2018 CHI conference on human factors in computing systems,pp\. 1–13\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p3.1)\.
- Furizal, A\. Ma’arif, H\. Maghfiroh, I\. Suwarno, D\. Prayogi, Kariyamin, S\. Lonang, and A\. Sharkawy \(2025\)Social, legal, and ethical implications of AI\-Generated deepfake pornography on digital platforms: A systematic literature review\.Social Sciences & Humanities Open12,pp\. 101882\.External Links:ISSN 2590\-2911,[Document](https://dx.doi.org/10.1016/j.ssaho.2025.101882)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p1.1)\.
- R\. Gandikota, J\. Materzynska, J\. Fiotto\-Kaufman, and D\. Bau \(2023\)Erasing concepts from diffusion models\.InProceedings of the IEEE/CVF international conference on computer vision,pp\. 2426–2436\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p7.1)\.
- C\. Gibson, D\. Olszewski, N\. G\. Brigham, A\. Crowder, K\. R\. Butler, P\. Traynor, E\. M\. Redmiles, and T\. Kohno \(2025\)Analyzing the\{\\\{ai\}\\\}nudification application ecosystem\.In34th USENIX Security Symposium \(USENIX Security 25\),pp\. 1–20\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p3.1),[§5](https://arxiv.org/html/2607.18263#S5.p7.1),[§7](https://arxiv.org/html/2607.18263#S7.p1.1)\.
- J\. Gibson \(2019\)Korea Wakes up to the Deadly Consequences of Spy Cams and Cyberbullying – The Diplomat\.External Links:[Link](https://thediplomat.com/2019/12/korea-wakes-up-to-the-deadly-consequences-of-spy-cams-and-cyberbullying/)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- I\. J\. Goodfellow, J\. Shlens, and C\. Szegedy \(2014\)Explaining and harnessing adversarial examples\.arXiv preprint arXiv:1412\.6572\.Cited by:[§6](https://arxiv.org/html/2607.18263#S6.p4.1)\.
- L\. Guarnera, O\. Giudice, and S\. Battiato \(2024\)Level up the deepfake detection: a method to effectively discriminate images generated by gan architectures and diffusion models\.InIntelligent Systems Conference,pp\. 615–625\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.4.4.2.1.1)\.
- Y\. Guo, Z\. Qu, W\. Lu, and X\. Luo \(2025\)Anti\-inpainting: a proactive defense against malicious diffusion\-based inpainters under unknown conditions\.arXiv preprint arXiv:2505\.13023\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p5.1),[§6](https://arxiv.org/html/2607.18263#S6.p4.1)\.
- B\. Gyevnar and A\. Kasirzadeh \(2025\)AI safety for everyone\.Nature Machine Intelligence,pp\. 1–12\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p7.1)\.
- Y\. Hamid, S\. Elyassami, Y\. Gulzar, V\. R\. Balasaraswathi, T\. Habuza, and S\. Wani \(2023\)An improvised cnn model for fake image detection\.International Journal of Information Technology15\(1\),pp\. 5–15\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.18.12.1.1.1)\.
- C\. Han, A\. Li, D\. Kumar, and Z\. Durumeric \(2025\)Characterizing the\{\\\{mrdeepfakes\}\\\}sexual deepfake marketplace\.In34th USENIX Security Symposium \(USENIX Security 25\),pp\. 5169–5188\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p3.1),[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p4.1)\.
- K\. R\. Harris \(2021\)Video on demand: what deepfakes do and how they harm\.Synthese199\(5\-6\),pp\. 13373–13391\.External Links:ISSN 0039\-7857, 1573\-0964,[Document](https://dx.doi.org/10.1007/s11229-021-03379-y)Cited by:[§2\.2](https://arxiv.org/html/2607.18263#S2.SS2.p4.1),[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p1.1)\.
- S\. Havron, D\. Freed, R\. Chatterjee, D\. McCoy, N\. Dell, and T\. Ristenpart \(2019\)Clinical computer security for victims of intimate partner violence\.In28th USENIX Security Symposium \(USENIX Security 19\),pp\. 105–122\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p3.1)\.
- S\. Hazra, B\. P\. Majumder, and T\. Chakrabarty \(2025\)AI safety should prioritize the future of work\.arXiv preprint arXiv:2504\.13959\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p7.1)\.
- N\. Henry and G\. Beard \(2024\)Image\-based sexual abuse perpetration: a scoping review\.Trauma, Violence, & Abuse25\(5\),pp\. 3981–3998\.Cited by:[§4\.3](https://arxiv.org/html/2607.18263#S4.SS3.SSS0.Px2.p1.1)\.
- N\. Henry and A\. Powell \(2016\)Sexual violence in the digital age: the scope and limits of criminal law\.Social & legal studies25\(4\),pp\. 397–418\.Cited by:[§4\.3](https://arxiv.org/html/2607.18263#S4.SS3.SSS0.Px2.p1.1)\.
- N\. Henry, R\. Umbach, R\. Shelby, G\. Beard, and L\. M\. Given \(2026\)‘It’s still abuse’: community attitudes and perceptions on ai\-generated image\-based sexual abuse\.Information, Communication & Society,pp\. 1–21\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p1.1)\.
- R\. Hönig, J\. Rando, N\. Carlini, and F\. Tramèr \(2025\)Adversarial perturbations cannot reliably protect artists from generative ai\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 70223–70263\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p5.1),[§6](https://arxiv.org/html/2607.18263#S6.p4.1)\.
- \[52\]\(2024\)How Telegram Became a Playground for Criminals, Extremists and Terrorists \- The New York Times\.External Links:[Link](https://www.nytimes.com/2024/09/07/technology/telegram-crime-terrorism.html)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- C\. Hsu, Y\. Zhuang, and C\. Lee \(2020\)Deep fake image detection based on pairwise learning\.Applied Sciences10\(1\),pp\. 370\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.1.1.2.1.1)\.
- E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, W\. Chen,et al\.\(2022\)Lora: low\-rank adaptation of large language models\.\.ICLR1\(2\),pp\. 3\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p2.1)\.
- P\. Isola, J\. Zhu, T\. Zhou, and A\. A\. Efros \(2017\)Image\-to\-image translation with conditional adversarial networks\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 1125–1134\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p2.1)\.
- M\. Ivanovska and V\. Štruc \(2023\)Face morphing attack detection with denoising diffusion probabilistic models\.arXiv preprint arXiv:2306\.15733\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.40.34.1.1.1)\.
- J\. Jeon, W\. J\. Kim, S\. Ha, S\. Son, and S\. Yoon \(2025\)AdvPaint: protecting images from inpainting manipulation via adversarial attention disruption\.arXiv preprint arXiv:2503\.10081\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p5.1)\.
- H\. Kang, S\. Wen, Z\. Wen, J\. Ye, W\. Li, P\. Feng, B\. Zhou, B\. Wang, D\. Lin, L\. Zhang,et al\.\(2025\)Legion: learning to ground and explain for synthetic image detection\.arXiv preprint arXiv:2503\.15264\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.37.31.1.1.1)\.
- M\. Keita, W\. Hamidouche, H\. Bougueffa Eutamene, A\. Taleb\-Ahmed, D\. Camacho, and A\. Hadid \(2025\)Bi\-lora: a vision\-language approach for synthetic image detection\.Expert Systems42\(2\),pp\. e13829\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.36.30.1.1.1)\.
- J\. Kim, Y\. Nam, M\. Kim, S\. Kim, and J\. Jeong \(2026\)BlurGuard: a simple approach for robustifying image protection against ai\-powered editing\.Advances in Neural Information Processing Systems38,pp\. 28664–28706\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p5.1)\.
- L\. Lei, K\. Gai, J\. Yu, and L\. Zhu \(2024\)Diffusetrace: a transparent and flexible watermarking scheme for latent diffusion model\.arXiv preprint arXiv:2405\.02696\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.29.23.1.1.1)\.
- Q\. Liao, Y\. Li, X\. Wang, B\. Kong, B\. Zhu, S\. Lyu, Y\. Yin, Q\. Song, and X\. Wu \(2021\)Imperceptible adversarial examples for fake image detection\.arXiv preprint arXiv:2106\.01615\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.5.2.1.1)\.
- H\. Liu, Z\. Tan, C\. Tan, Y\. Wei, J\. Wang, and Y\. Zhao \(2024\)Forgery\-aware adaptive transformer for generalizable synthetic image detection\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 10770–10780\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.13.7.1.1.1)\.
- Y\. Liu, Z\. Li, M\. Backes, Y\. Shen, and Y\. Zhang \(2023\)Watermarking diffusion model\.arXiv preprint arXiv:2305\.12502\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.20.14.1.1.1)\.
- Y\. Luo, J\. Du, K\. Yan, and S\. Ding \(2024\)LaREˆ 2: latent reconstruction error based method for diffusion\-generated image detection\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 17006–17015\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.19.13.1.1.1)\.
- R\. Ma, J\. Duan, F\. Kong, X\. Shi, and K\. Xu \(2023\)Exposing the fake: effective diffusion\-generated images detection\.arXiv preprint arXiv:2307\.06272\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.21.15.1.1.1)\.
- S\. Maddocks \(2020\)‘A deepfake porn plot intended to silence me’: exploring continuities between pornographic and ‘political’deep fakes\.Porn Studies7\(4\),pp\. 415–423\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p1.1)\.
- E\. Maiberg \(2025a\)A16z\-Backed AI Site Civitai Is Mostly Porn, Despite Claiming Otherwise\.\(en\)\.External Links:[Link](https://www.404media.co/a16z-backed-ai-site-civitai-is-mostly-porn-despite-claiming-otherwise/)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- E\. Maiberg \(2025b\)Hugging Face Is Hosting 5,000 Nonconsensual AI Models of Real People\.\(en\)\.External Links:[Link](https://www.404media.co/hugging-face-is-hosting-5-000-nonconsensual-ai-models-of-real-people/)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- M\. Marini, A\. Ansani, A\. Demichelis, G\. Mancini, F\. Paglieri, and M\. Viola \(2024\)Real is the new sexy: the influence of perceived realness on self\-reported arousal to sexual visual stimuli\.Cognition and Emotion38\(3\),pp\. 348–360\.Cited by:[§4\.3](https://arxiv.org/html/2607.18263#S4.SS3.SSS0.Px2.p1.1)\.
- A\. Massanari \(2017\)\# gamergate and the fappening: how reddit’s algorithm, governance, and culture support toxic technocultures\.New media & society19\(3\),pp\. 329–346\.Cited by:[§4\.3](https://arxiv.org/html/2607.18263#S4.SS3.SSS0.Px2.p1.1)\.
- C\. McGlynn and N\. Westmarland \(2019\)Kaleidoscopic justice: sexual violence and victim\-survivors’ perceptions of justice\.Social & Legal Studies28\(2\),pp\. 179–201\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p9.1)\.
- Inc\. Meta Platforms \(2024\)Our approach to labeling ai\-generated content and manipulated media\.Note:Meta blog post on labeling AI content across Facebook, Instagram, and Threads[https://about\.fb\.com/news/2024/04/metas\-approach\-to\-labeling\-ai\-generated\-content\-and\-manipulated\-media/](https://about.fb.com/news/2024/04/metas-approach-to-labeling-ai-generated-content-and-manipulated-media/)External Links:[Link](https://about.fb.com/news/2024/04/metas-approach-to-labeling-ai-generated-content-and-manipulated-media/)Cited by:[§4\.3](https://arxiv.org/html/2607.18263#S4.SS3.SSS0.Px1.p1.1)\.
- B\. New, L\. Salazar, M\. Lozano, and S\. Fralicks \(2024\)Deepfakes of Elon Musk are contributing to billions of dollars in fraud losses in the U\.S\. \- CBS Texas\.Note:https://www\.cbsnews\.com/texas/news/deepfakes\-ai\-fraud\-elon\-musk/Cited by:[§2](https://arxiv.org/html/2607.18263#S2.p1.1)\.
- H\. Nissenbaum \(2004\)Privacy as contextual integrity\.Wash\. L\. Rev\.79,pp\. 119\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p3.1)\.
- U\. Ojha, Y\. Li, and Y\. J\. Lee \(2023\)Towards universal fake image detectors that generalize across generative models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 24480–24489\.Cited by:[§3\.1](https://arxiv.org/html/2607.18263#S3.SS1.p2.1)\.
- A\. Olteanu, S\. Barocas, S\. L\. Blodgett, L\. Egede, A\. DeVrio, and M\. Cheng \(2025\)Ai automatons: ai systems intended to imitate humans\.arXiv preprint arXiv:2503\.02250\.Cited by:[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p1.1)\.
- Oversight Board \(2024\)New decision addresses meta’s rules on non\-consensual deepfake intimate images\.Note:Oversight Board press release on deepfake intimate image policy and Meta’s rules[https://www\.oversightboard\.com/news/new\-decision\-addresses\-metas\-rules\-on\-non\-consensual\-deepfake\-intimate\-images/](https://www.oversightboard.com/news/new-decision-addresses-metas-rules-on-non-consensual-deepfake-intimate-images/)External Links:[Link](https://www.oversightboard.com/news/new-decision-addresses-metas-rules-on-non-consensual-deepfake-intimate-images/)Cited by:[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p4.1)\.
- T\. C\. Ozden, O\. Kara, O\. Akcin, K\. Zaman, S\. Srivastava, S\. P\. Chinchali, and J\. M\. Rehg \(2025\)DiffVax: optimization\-free image immunization against diffusion\-based editing\.InThe Fourteenth International Conference on Learning Representations,Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p5.1)\.
- Y\. Patel, S\. Tanwar, P\. Bhattacharya, R\. Gupta, T\. Alsuwian, I\. E\. Davidson, and T\. F\. Mazibuko \(2023\)An improved dense cnn architecture for deepfake image detection\.IEEE Access11,pp\. 22081–22095\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.3.3.2.1.1)\.
- I\. Perov, D\. Gao, N\. Chervoniy, K\. Liu, S\. Marangonda, C\. Umé, C\. S\. Facenheim, L\. RP, J\. Jiang, S\. Zhang,et al\.\(2020\)DeepFaceLab: integrated, flexible and extensible face\-swapping framework\.arXiv preprint arXiv:2005\.05535\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p2.1)\.
- L\. Qiwei, S\. Zhang, S\. P\. Pratt, A\. T\. Kasper, E\. Gilbert, and S\. Schoenebeck \(2025\)A law of one’s own: the inefficacy of the dmca for non\-consensual intimate media\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,pp\. 1–17\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark,et al\.\(2021\)Learning transferable visual models from natural language supervision\.InInternational conference on machine learning,pp\. 8748–8763\.Cited by:[§3\.1](https://arxiv.org/html/2607.18263#S3.SS1.p2.1)\.
- M\. A\. Rahman, B\. Paul, N\. H\. Sarker, Z\. I\. A\. Hakim, and S\. A\. Fattah \(2023\)Artifact: a large\-scale dataset with artificial and factual images for generalizable and robust synthetic image detection\.In2023 IEEE International Conference on Image Processing \(ICIP\),pp\. 2200–2204\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.22.16.1.1.1)\.
- A\. Raza, K\. Munir, and M\. Almutairi \(2022\)A novel deep learning approach for deepfake image detection\.Applied Sciences12\(19\),pp\. 9820\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.2.2.2.1.1)\.
- J\. Ricker, S\. Damm, T\. Holz, and A\. Fischer \(2022\)Towards the detection of diffusion model deepfakes\.arXiv preprint arXiv:2210\.14571\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.12.6.1.1.1)\.
- J\. Ricker, D\. Lukovnikov, and A\. Fischer \(2024\)Aeroblade: training\-free detection of latent diffusion images using autoencoder reconstruction error\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 9130–9140\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.15.9.1.1.1)\.
- R\. Rini \(2020\)Deepfakes and the epistemic backstop\.Cited by:[§2\.2](https://arxiv.org/html/2607.18263#S2.SS2.p4.1),[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p2.1)\.
- R\. Rombach, A\. Blattmann, D\. Lorenz, P\. Esser, and B\. Ommer \(2022\)High\-resolution image synthesis with latent diffusion models\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 10684–10695\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p2.1),[§3\.1](https://arxiv.org/html/2607.18263#S3.SS1.p2.1)\.
- L\. Rosenthol \(2022\)C2PA: the world’s first industry standard for content provenance \(conference presentation\)\.InApplications of Digital Image Processing XLV,Vol\.12226,pp\. 122260P\.Cited by:[§3\.2](https://arxiv.org/html/2607.18263#S3.SS2.p1.1)\.
- N\. Ruiz, Y\. Li, V\. Jampani, Y\. Pritch, M\. Rubinstein, and K\. Aberman \(2023\)Dreambooth: fine tuning text\-to\-image diffusion models for subject\-driven generation\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 22500–22510\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p2.1)\.
- M\. K\. Scheuerman, A\. Hanna, and R\. Denton \(2021\)Do datasets have politics? disciplinary values in computer vision dataset development\.Proceedings of the ACM on Human\-Computer Interaction5\(CSCW2\),pp\. 1–37\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p3.1)\.
- S\. Schoenebeck, O\. L\. Haimson, and L\. Nakamura \(2021\)Drawing from justice theories to support targets of online harassment\.new media & society23\(5\),pp\. 1278–1300\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p9.1)\.
- P\. Schramowski, M\. Brack, B\. Deiseroth, and K\. Kersting \(2023\)Safe latent diffusion: mitigating inappropriate degeneration in diffusion models\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 22522–22531\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p7.1)\.
- C\. Schuhmann, R\. Beaumont, R\. Vencu, C\. Gordon, R\. Wightman, M\. Cherti, T\. Coombes, A\. Katta, C\. Mullis, M\. Wortsman,et al\.\(2022\)Laion\-5b: an open large\-scale dataset for training next generation image\-text models\.Advances in neural information processing systems35,pp\. 25278–25294\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p2.1)\.
- Security Hero \(2023\)2023 state of deepfakes\.Technical reportSecurity Hero\.Cited by:[item 1](https://arxiv.org/html/2607.18263#S1.I1.i1.p1.1),[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p1.1),[§2](https://arxiv.org/html/2607.18263#S2.p1.1),[§7](https://arxiv.org/html/2607.18263#S7.p1.1)\.
- E\. Seger, N\. Dreksler, R\. Moulange, E\. Dardaman, J\. Schuett, K\. Wei, C\. Winter, M\. Arnold, S\. Ó\. hÉigeartaigh, A\. Korinek,et al\.\(2023\)Open\-sourcing highly capable foundation models: an evaluation of risks, benefits, and alternative methods for pursuing open\-source objectives\.arXiv preprint arXiv:2311\.09227\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p4.1)\.
- R\. J\. \[\. Sen\. Durbin \(2024\)S\.3696 \- 118th Congress \(2023\-2024\): DEFIANCE Act of 2024\.legislation\(eng\)\.Note:Archive Location: 2024\-01\-30External Links:[Link](https://www.congress.gov/bill/118th-congress/senate-bill/3696)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- S\. Shan, J\. Cryan, E\. Wenger, H\. Zheng, R\. Hanocka, and B\. Y\. Zhao \(2023\)Glaze: protecting artists from style mimicry by\{\\\{text\-to\-image\}\\\}models\.In32nd USENIX Security Symposium \(USENIX Security 23\),pp\. 2187–2204\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p5.1)\.
- S\. Sinitsa and O\. Fried \(2024\)Deep image fingerprint: towards low budget synthetic image detection and model lineage analysis\.InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision,pp\. 4067–4076\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.30.24.1.1.1)\.
- I\. Solaiman \(2023\)The gradient of generative ai release: methods and considerations\.InProceedings of the 2023 ACM conference on fairness, accountability, and transparency,pp\. 111–122\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p4.1)\.
- H\. Song, S\. Huang, Y\. Dong, and W\. Tu \(2023\)Robustness and generalizability of deepfake detection: a study with diffusion models\.arXiv preprint arXiv:2309\.02218\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.27.21.1.1.1)\.
- X\. Song, X\. Guo, J\. Zhang, Q\. Li, L\. Bai, X\. Liu, G\. Zhai, and X\. Liu \(2024\)On learning multi\-modal forgery representation for diffusion generated video detection\.Advances in Neural Information Processing Systems37,pp\. 122054–122077\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.31.25.1.1.1)\.
- L\. Stark \(2018\)Facial recognition, emotion and race in animated social media\.First Monday\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p3.1)\.
- J\. Sun, J\. Wang, W\. Nie, Z\. Yu, Z\. Mao, and C\. Xiao \(2023\)A critical revisit of adversarial robustness in 3d point cloud recognition with diffusion\-driven purification\.InInternational Conference on Machine Learning,pp\. 33100–33114\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.39.33.1.1.1),[§6](https://arxiv.org/html/2607.18263#S6.p4.1)\.
- K\. Sun, S\. Chen, T\. Yao, H\. Liu, X\. Sun, S\. Ding, and R\. Ji \(2024\)Diffusionfake: enhancing generalization in deepfake detection via guided stable diffusion\.Advances in Neural Information Processing Systems37,pp\. 101474–101497\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.28.22.1.1.1)\.
- K\. Tenbarge \(2026\)Grok is being used to mock and strip women in hijabs and sarees\.Note:WIRED\[Online; accessed 27\-Jan\-2026\]External Links:[Link](https://www.wired.com/story/grok-is-being-used-to-mock-and-strip-women-in-hijabs-and-sarees/)Cited by:[footnote 1](https://arxiv.org/html/2607.18263#footnote1)\.
- D\. Thiel \(2023\)Identifying and eliminating csam in generative ml training data and models\.Stanford Internet Observatory, Cyber Policy Center, December23\(3\),pp\. 131\.Cited by:[§6](https://arxiv.org/html/2607.18263#S6.p7.1)\.
- TikTok \(2026\)About ai\-generated content\.Note:TikTok support page on AI\-generated content definitions and labeling requirements[https://www\.tiktok\.com/tns\-inapp/pages/ai\-generated\-content](https://www.tiktok.com/tns-inapp/pages/ai-generated-content)External Links:[Link](https://www.tiktok.com/tns-inapp/pages/ai-generated-content)Cited by:[§4\.3](https://arxiv.org/html/2607.18263#S4.SS3.SSS0.Px1.p1.1)\.
- United Nations \(1948\)Universal declaration of human rights\.Cited by:[§1](https://arxiv.org/html/2607.18263#S1.p2.1),[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p1.1)\.
- E\. Van der Nagel \(2020\)Verifying images: deepfakes, control, and consent\.Porn Studies7\(4\),pp\. 424–429\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p1.1)\.
- T\. Van Le, H\. Phung, T\. H\. Nguyen, Q\. Dao, N\. N\. Tran, and A\. Tran \(2023\)Anti\-dreambooth: protecting users from personalized text\-to\-image synthesis\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 2116–2127\.Cited by:[§4\.1](https://arxiv.org/html/2607.18263#S4.SS1.p3.1),[§5](https://arxiv.org/html/2607.18263#S5.p5.1)\.
- L\. Wagner and E\. Cetinic \(2025\)Perpetuating misogyny with generative ai: how model personalization normalizes gendered harm\.arXiv preprint arXiv:2505\.04600\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- T\. Wang, M\. Liu, J\. Zhu, A\. Tao, J\. Kautz, and B\. Catanzaro \(2018\)High\-resolution image synthesis and semantic manipulation with conditional gans\.InProceedings of the IEEE conference on computer vision and pattern recognition,pp\. 8798–8807\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p2.1)\.
- Z\. Wang, J\. Bao, W\. Zhou, W\. Wang, H\. Hu, H\. Chen, and H\. Li \(2023\)Dire for diffusion\-generated image detection\.InProceedings of the IEEE/CVF International Conference on Computer Vision,pp\. 22445–22455\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.8.2.1.1.1),[§3\.1](https://arxiv.org/html/2607.18263#S3.SS1.p2.1)\.
- M\. Wei, C\. Yeung, F\. Roesner, and T\. Kohno \(2025\)” We’re utterly ill\-prepared to deal with something like this”: teachers’ perspectives on student generation of synthetic nonconsensual explicit imagery\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,pp\. 1–18\.Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p1.1)\.
- L\. Weidinger, J\. Mellor, M\. Rauh, C\. Griffin, J\. Uesato, P\. Huang, M\. Cheng, M\. Glaese, B\. Balle, A\. Kasirzadeh,et al\.\(2021\)Ethical and social risks of harm from language models\.arXiv preprint arXiv:2112\.04359\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p7.1)\.
- S\. Wen, J\. Ye, P\. Feng, H\. Kang, Z\. Wen, Y\. Chen, J\. Wu, W\. Wu, C\. He, and W\. Li \(2025\)Spot the fake: large multimodal model\-based synthetic image detection with artifact explanation\.arXiv preprint arXiv:2503\.14905\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.34.28.1.1.1)\.
- Y\. Wen, J\. Kirchenbauer, J\. Geiping, and T\. Goldstein \(2023\)Tree\-ring watermarks: fingerprints for diffusion images that are invisible and robust\.arXiv preprint arXiv:2305\.20030\.Cited by:[§3\.3](https://arxiv.org/html/2607.18263#S3.SS3.p1.1)\.
- D\. G\. Widder, D\. Nafus, L\. Dabbish, and J\. Herbsleb \(2022\)Limits and possibilities for “ethical ai” in open source: a study of deepfakes\.InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency,pp\. 2035–2046\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p4.1)\.
- E\. Williamson, A\. Gregory, H\. Abrahams, N\. Aghtaie, S\. Walker, and M\. Hester \(2020\)Secondary trauma: emotional safety in sensitive research\.Journal of Academic Ethics18\(1\),pp\. 55–70\.Cited by:[§5](https://arxiv.org/html/2607.18263#S5.p8.1)\.
- H\. Wu, J\. Zhou, and S\. Zhang \(2025\)Generalizable synthetic image detection via language\-guided contrastive learning\.IEEE Transactions on Artificial Intelligence\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.25.19.1.1.1)\.
- Z\. Yang, K\. Zeng, K\. Chen, H\. Fang, W\. Zhang, and N\. Yu \(2024\)Gaussian shading: provable performance\-lossless image watermarking for diffusion models\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 12162–12171\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.14.8.1.1.1)\.
- H\. Yim \(2024\)South Korea to criminalize watching or possessing sexually explicit deepfakes\.External Links:[Link](https://www.reuters.com/world/asia-pacific/south-korea-criminalise-watching-or-possessing-sexually-explicit-deepfakes-2024-09-26/)Cited by:[§2\.1](https://arxiv.org/html/2607.18263#S2.SS1.p4.1)\.
- Z\. Yu, J\. Ni, Y\. Lin, H\. Deng, and B\. Li \(2024\)Diffforensics: leveraging diffusion prior to image forgery detection and localization\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,pp\. 12765–12774\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.24.18.1.1.1)\.
- B\. Zhang, S\. Li, G\. Feng, Z\. Qian, and X\. Zhang \(2022\)Patch diffusion: a general module for face manipulation detection\.InProceedings of the AAAI conference on artificial intelligence,Vol\.36,pp\. 3243–3251\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.33.27.1.1.1)\.
- L\. Zhang, X\. Liu, A\. V\. Martin, C\. X\. Bearfield, Y\. Brun, and H\. Guan \(2024\)Attack\-resilient image watermarking using stable diffusion\.Advances in Neural Information Processing Systems37,pp\. 38480–38507\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.35.29.1.1.1)\.
- Y\. Zhang and X\. Xu \(2023\)Diffusion noise feature: accurate and fast generated image detection\.arXiv preprint arXiv:2312\.02625\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.38.32.1.1.1)\.
- Y\. Zhao, T\. Pang, C\. Du, X\. Yang, N\. Cheung, and M\. Lin \(2023\)A recipe for watermarking diffusion models\.arXiv preprint arXiv:2303\.10137\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.11.5.1.1.1)\.
- Y\. Zhu, X\. Wang, H\. Chen, R\. Salloum, and C\. J\. Kuo \(2022\)A\-pixelhop: a green, robust and explainable fake\-image detector\.InICASSP 2022\-2022 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 8947–8951\.Cited by:[Table 2](https://arxiv.org/html/2607.18263#A0.T2.5.32.26.1.1.1)\.

Table 2:39 papers surveyed, listed in order of number of citation each paper received in January 2026\.

Similar Articles

AI safety and alignment

Reddit r/artificial

The article discusses concerns about AI safety and alignment as AI becomes more intelligent and integrated into society, referencing Anthropic's call for a pause to address potential catastrophic risks.

Agent Safety Is Action Alignment

arXiv cs.AI

This paper argues that applying content-safety refusal methods to AI agents is a category error—agentic harm lies in authority misuse rather than output—and proposes action alignment enforced outside the model via least privilege.

Why I am against GenAI and everything it stands for

Lobsters Hottest

The author argues that generative AI is harmful, citing the use of stolen training data, its role in spreading misinformation, and its embodiment of exploitative capitalism, while distinguishing it from traditional machine learning.