SAI
← All ICML 2026 orals

Position: AI/ML Deepfake Research is Misaligned with AI Generated Non-Consensual Intimate Imagery (AIG-NCII)

Li Qiwei, Wells Lucas Santo, Sarita Schoenebeck, Eric Gilbert

OralPosition TrackReplication not startedPaper PDFOpenReview

Position: AI/ML Deepfake Research is Misaligned with AI Generated Non-Consensual Intimate Imagery (AIG-NCII)

SAI paper + code review · Referee report

Summary

This position paper argues that the AI/ML deepfake defense literature is structurally misaligned with the empirical reality of AI-generated non-consensual intimate imagery (AIG-NCII). The conceptual move is compact and useful: existing detection, provenance, and watermarking work treats authenticity as a proxy for safety, but for AIG-NCII the harm is defined by absence of consent and is orthogonal to whether the image is synthetic. The authors introduce a viewer-centric vs subject-centric taxonomy of harms and a synthetic-vs-authentic × safe-vs-harmful matrix to show that the field currently only distinguishes along one of the two axes that matter. They support the misalignment claim with a landscape analysis of 39 highly-cited deepfake defense papers (2020–2025), of which they report 34 make no mention of AIG-NCII, 5 mention it only in passing, and 0 target it with a specific threat model. Section 4 argues that authenticity tools can, in some deployment contexts, exacerbate subject harms (backend-visibility risks; confirming exposure of traditional NCII victims; enabling abuser sorting), and the recommendations (R1–R9) urge decoupling epistemic from dignity harms, restricting high-risk research assets, investing in proactive perturbation defenses, and integrating subject-centric harm into AI Safety.

The framing is timely and the taxonomy is a genuine contribution to how the ML community should scope its threat models. The main conceptual limitation is that the paper's headline empirical claims — that AIG-NCII accounts for the majority of generative AI usage, and that the deepfake defense literature ignores it — are stronger than the evidence rigorously supports. The cited sources measure specific systems (Grok), specific modalities (deepfake videos), and specific ecosystems (nudify apps), not all generative AI. The 39-paper corpus is filtered by keywords that pre-select for the authenticity-oriented sub-field, so the 34/5/0 finding partly bakes in its conclusion; the same manuscript later cites a growing anti-personalization literature that is excluded from the denominator. Coding protocol, inter-rater reliability, and raw evidence for the tier assignments are also not documented. These are addressable through scope-narrowing, clearer coding artifacts, and a companion count for the anti-personalization literature; the conceptual argument survives them.

Strengths

  • Conceptual contribution. The paper's central move — separating viewer-centric epistemic harm from subject-centric dignity harm, and separating the artificiality axis from the consent axis — is crisp and gives the community a taxonomy it currently lacks. The authentic != safe framing is a clean articulation of a category error the field is making.
  • Deployment-context analysis. Section 4.3 does something rare in position papers: it walks through how the same authenticity tool can help, do nothing, or actively harm depending on who wields it (platforms, abusers, traditional NCII survivors). The Oversight Board quote and the "plausible deniability as protective cover" argument are non-obvious and change how one should read regulatory labeling proposals.
  • Directly linking academic capabilities to abuse pipelines. The paper is unusually specific about how open-source research code, LoRA/DreamBooth-style personalization, and open-weight diffusion models are being forked into AIG-NCII tooling (Han et al. 2025; Gibson et al. 2025). This grounds R3's argument for reconsidering release norms.
  • Actionable recommendations. R1–R9 span technical (proactive perturbations, safety-aligned metrics), infrastructural (gated release), and organizational (partnerships with sexual-violence experts, secondary-trauma guardrails) directions, which is the correct decomposition for a problem that will not be solved with a single classifier.
  • Serious alternative-views section. AV1–AV3 anticipate the strongest objections — that policy should own this, that defenses are intractable, that institutional incentives block engagement — and give reasoned rebuttals, including a candid concession that no defense will be perfect.

Weaknesses

  • Central empirical claim overreaches its sources. The contribution list and conclusion both state that AIG-NCII accounts for the majority of generative AI usage. The cited evidence — 98% of deepfake videos being pornographic, the Grok undressing statistic, the anomalous 4.4M-image Grok week — supports the majority of deepfake or nudify-style usage, not the majority of all generative AI usage (which is dominated by text and code). The introduction hedges as "a primary driver of generative AI usage"; the flat assertions elsewhere overstate what has been shown.
  • Landscape survey is scoped in a way that partly presupposes the conclusion. The search query intersects (detection|detector|forensics|recognition|watermark) with deepfake/synthetic-image terms, which by construction filters to authenticity-oriented interventions and excludes the very anti-personalization literature that the same paper cites in R4 (Anti-DreamBooth, Glaze, PhotoGuard, Diffvax, Blurguard, AdvPaint, Anti-Inpainting). The claim "deepfake research is overwhelmingly motivated by trust/fraud/misinformation" therefore holds for the authenticity sub-field the survey samples, but is generalized to the whole field.
  • Inclusion rule is self-contradictory. "Filtered ... only at top-tier venues" is immediately followed by two exception pathways (arXiv preprints; any venue with >80 citations). The actual predicate is a disjunction, not the "only" the sentence claims. Table 2 confirms this: papers appear from Applied Sciences, IEEE Access, WACV, IEEE OJSP, IEEE ICIP, Springer IJSA, Intelligent Systems Conf., and Expert Systems, all outside the enumerated top-tier set — and all five checkmarked "mentions AIG-NCII" papers come from those non-top venues.
  • Table 1 mixes two senses of "safe" in the very cell that is supposed to define the consent axis. "Consensual pornography" is safe because a real subject consented; "artistic nudity using AI" (of no identifiable person) is out of scope of the consent axis rather than a positive case of consent-based safety. Using both under one "Safe" column blurs the axis the paper is trying to establish. Relatedly, the prose calls the horizontal axis "consent" while the table labels it "Safe/Harmful", which conflates the axis with its conclusion.
  • Marini et al. cuts against the argument it is asked to support. Section 4.3 argues abusers are motivated by power rather than gratification, then cites Marini et al. (2024) that people are less aroused when they know an image is AI-generated. That finding implies abusers should prefer believed-authentic material, which weakens the same paragraph's claim that abusers will use authenticity tools to sort for and target AIG-NCII.
  • "Prioritize labeling over removal" overreads what EU AI Act / Meta / TikTok actually do. All three regimes maintain removal obligations for NCII/CSAM alongside labeling. The premise for the "perverse outcome ... abuser protected so long as they are transparent" argument needs a specific pointer to where these frameworks decline to require removal of AIG-NCII, not just that they require labeling of synthetic media.
  • Recommendation R5 conflates two very different measurement targets. The text asks for metrics that verify systems "actually reduce the prevalence of abuse" (a causal/ecological quantity) but exemplifies this with Cretu et al.'s (2025) component-level filter evaluation using ethical proxies. These are quite different measurement problems; the recommendation should say which one it wants (or both, and why).
  • Uncited quantitative rhetoric on casual abusers. The AV2 rebuttal rests on "casual abuser (teenagers, ex-partners) who account for a significant volume of harassment" — a factual claim presented without a source in a paper that otherwise cites carefully.
  • Cross-section citation reuse creates apparent inconsistencies. Van Le et al. (2023) is described in Section 4.1 as a defense against non-consensual identity preservation (fine-tuning-time), and in R4 as a defense against "style mimicry or inpainting" — related but distinct threat models. Similarly, Guo et al. (2025) appears both in R4 as an emerging defense and in AV2 as evidence of defense brittleness.
  • Fairness analogy elides the capability-to-harm disanalogy. The AV3 rebuttal draws a direct parallel to algorithmic fairness, but the paper's own Section 2 documents that academic deepfake code is forked directly into abuse tooling, a pipeline that fairness research did not create. The analogy is motivating but should acknowledge this asymmetry, since it is the very reason R3 (gated release) and R7 (guardrails) are needed.
  • Minor precision issues. "Deepnude relied directly on Pix2Pix" collapses pix2pix vs pix2pixHD; DeepNude (inpainting) is presented in the paragraph about face-swap-as-primary-mechanism, which are architecturally distinct pathways; "responsibility for the mitigating their negative societal impacts" is a garbled clause in the AV3 rebuttal; and "23,000 being of children" is ambiguous between "sexualized images of children" and "images of children" (context suggests the former but this should be stated).

Reproducibility & code

The paper releases no separate code repository; the reproducibility surface is the landscape-analysis methodology in Section 2.2 and Appendix Table 2. Reading these together, the core arithmetic is auditable (39 rows, 5 checkmarks, so 34 no-mention and 12.8% mention share follow directly), but several methodological artifacts a reader would need to independently reproduce the pipeline are missing.

  • Google Scholar snapshot not archived. The exact search date, tooling (scrape vs manual paging), and the raw 965-hit list are not reported, so the 965 → 379 funnel cannot be re-derived without trusting the count. Scholar totals are known to drift by day and by locale.
  • Inclusion rule is under-specified and contradicts Table 2. "Only" top-tier venues is followed by arXiv and >80-citation exceptions, and Table 2 in fact contains many non-top venues (Applied Sciences, IEEE Access, WACV, IEEE OJSP, IEEE ICIP, Intelligent Systems Conf., Expert Systems, Springer IJSA). The interaction between the workshop-exclusion rule, the arXiv inclusion, and the >80-citation exception is not written down as a single predicate.
  • Citation counts driving the top-100 cut are not published. The 379 → 100 → 39 reduction relies on citation ranks and manual exclusions (unrelated CV tasks; one retracted paper) that are not itemized. A supplementary table listing the top-100 candidates with their citation counts and inclusion/exclusion label would make this step auditable.
  • Coding protocol for the 34/5/0 tiers not documented. The manuscript does not report whether coding was full-text or abstract-only, single-coder or multi-coder, or with what inter-rater reliability. The definitional boundary that most matters — "mention only" (5) vs "technical implementation specific to AIG-NCII" (0) — is stated only in prose, and it does the paper's largest rhetorical work.
  • Per-paper evidence snippets not shipped. Table 2 shows only checkmarks. A supplementary spreadsheet listing, for each of the 39 papers, the sentence that triggered a mention (or its absence) and the tier assignment would let readers audit the boundary calls that produce the 34/5/0 headline.
  • Figure 1 uses external and internal denominators that are not commensurate. The 51.7% AIG-NCII share comes from Bouchaud (2026), a report whose sample scope and denominator this paper does not restate; the 12.8% share is over the 39-paper corpus. Both bars can be believed on their own, but juxtaposing them as if they were a like-for-like comparison overstates precision.

None of the above threatens the paper's qualitative thesis, but each item stands between the current draft and independent verification of its headline numbers.

Recommended Changes

Essential

  • Narrow the "majority of generative AI usage" claim. In the abstract, contribution list, and conclusion, restrict the denominator to what the cited evidence actually measures — deepfake usage, nudify-app usage, or specific systems — rather than all generative AI usage. Restore the "likely" hedge in the conclusion or replace with a scope-restricted phrasing (this addresses the overreach flagged in Weaknesses and in the two overclaim comments).
  • Broaden or explicitly restrict the landscape survey's scope. Either add a companion count for the anti-personalization / immunization sub-field (Anti-DreamBooth, Glaze, PhotoGuard, Diffvax, Blurguard, AdvPaint, Anti-Inpainting, etc.), or explicitly note that the survey samples only authenticity-oriented interventions by construction of the query. This addresses the "survey scope partly presupposes the conclusion" weakness.
  • Rewrite the inclusion rule as a single predicate. Drop "only", state the disjunction (top-tier venue OR arXiv OR >80 citations OR ...), specify how workshop exclusion interacts, and give the exact search snapshot date. This addresses the self-contradictory inclusion rule and the mismatch with Table 2 venues.
  • Document the coding protocol and release per-paper evidence. Report whether coding was full-text, single- or multi-coder, and any inter-rater reliability; publish a supplementary spreadsheet giving, for each of the 39 papers, the triggering sentence (or its absence) and the tier rationale — especially the boundary between "mention only" and "technical implementation".
  • Fix Table 1's mixed senses of "safe". Replace "Self-expression of artistic nudity using AI" with a Synthetic-Safe exemplar that involves subject consent (e.g., AI-generated intimate imagery of oneself), so the consent axis is genuinely the same across both "Safe" cells. Consider relabeling the columns as "Consensual / Non-consensual" to match the prose.
  • Reconcile the Marini et al. citation with the "power not gratification" argument. Either drop the arousal finding, or extend the argument to explain why abusers would still seek authenticity-labeled content in the presence of reduced arousal — for example, by distinguishing consumption gratification from targeting/harassment motives.
  • Support the "casual abusers account for a significant volume" claim with a citation, or reframe as a plausibility argument rather than an empirical premise.

Suggested

  • Tighten the regulatory framing. Replace "prioritize labeling of synthetic media" with a claim that is defensible against EU AI Act / Meta / TikTok obligations — e.g., "treat labeling as sufficient in lieu of removal in some contexts" — and cite where each framework declines to require removal of AIG-NCII.
  • Sharpen R5. State whether the recommendation is for system-level prevalence-reduction metrics, safe component-level metrics, or both, and clarify what role Cretu et al. (2025) plays as inspiration.
  • Add a sentence to the fairness analogy in AV3 acknowledging that AIG-NCII differs from fairness in that academic capabilities are forked directly into abuse pipelines, so the trajectory requires the safeguards proposed in R3 and R7.
  • Clarify Van Le et al. and Guo et al. citations. State what each defense actually targets (fine-tuning-time identity retention vs inpainting), and if a work both proposes a defense and characterizes its brittleness, say so at each citation.
  • Fix minor precision issues. Distinguish pix2pix vs pix2pixHD, or drop the specific-architecture claim about DeepNude; separate the face-swap and undress/inpainting pathways in Section 2.1; repair the "responsibility for the mitigating their negative" clause; and disambiguate "23,000 being of children" (sexualized images of children).
  • Annotate Figure 1's denominators so readers cannot mistake the 51.7% vs 12.8% juxtaposition for a like-for-like measurement.