Full Report
Researchers tested top image editing models on Hugging Face and found they could easily create explicit deepfakes—and 1,000 image editing prompts show how people use the software.
Analysis Summary
# Incident Report: Widespread Generation of Nonconsensual Deepfakes on Hugging Face
## Executive Summary
Researchers from the nonprofit AI Forensics discovered that top image-editing models hosted on Hugging Face lack sufficient safeguards, allowing users to easily generate nonconsensual explicit deepfakes. By deploying "honey-pot" Spaces, researchers confirmed that a significant portion of user demand on the platform is directed toward creating sexually explicit imagery.
## Incident Details
- **Discovery Date:** July 2026 (Report publication)
- **Incident Date:** Ongoing
- **Affected Organization:** Hugging Face
- **Sector:** Technology / Artificial Intelligence
- **Geography:** Global (Platform-wide)
## Timeline of Events
### Initial Access
- **Date/Time:** Investigative period leading up to July 2026.
- **Vector:** Exploitation of open-access AI "Spaces."
- **Details:** Users leveraged existing, legitimate image-to-image models hosted on the platform to bypass weak or non-existent content filters.
### Lateral Movement
- **Details:** Not applicable in a traditional network sense; however, the behavior proliferated across multiple different AI "Spaces" (models) within the Hugging Face ecosystem.
### Data Exfiltration/Impact
- **Impact:** Generation of Nonconsensual Intimate Imagery (NCII). Researchers documented 1,000+ prompts from users intentionally seeking to "nudify" or create explicit content of individuals.
### Detection & Response
- **Detection:** Identified by European nonprofit AI Forensics through systematic testing of nine top image-editing Spaces.
- **Response actions taken:** Researchers published a formal report highlighting the "nudifier" problem and the failure of existing platform safety mechanisms.
## Attack Methodology
- **Initial Access:** Public web interface of Hugging Face Spaces.
- **Persistence:** High; models remained available for public use despite violating implicit safety standards.
- **Defense Evasion:** Use of specific image-editing prompts designed to bypass rudimentary keyword filters.
- **Impact:** Abuse of Generative AI to create harmful, nonconsensual digital forgeries.
## Impact Assessment
- **Financial:** Potential loss of valuation or investor confidence due to regulatory scrutiny (Hugging Face is valued in the billions).
- **Data Breach:** Exposure of personal likenesses (biometric/visual data) manipulated for explicit purposes.
- **Operational:** Increased overhead for TRUST & Safety teams to moderate thousands of uploaded models.
- **Reputational:** High; the platform is being labeled as a primary source for "nudify" software, contradicting safety pledges.
## Indicators of Compromise
- **Behavioral indicators:** High volume of image-to-image prompts containing keywords related to undressing, specific body parts, or "nudifying" clothes.
- **Platform indicators:** Hosting of models specifically fine-tuned on explicit datasets (e.g., Unstable Diffusion variants).
## Response Actions
- **Containment:** None currently reported as fully effective; researchers found that 7 out of 9 top models remained vulnerable at the time of the report.
- **Eradication:** Removal of "honey-pot" testing environments and recommendations for platform-wide bans on "nudifier" style apps.
## Lessons Learned
- **Key takeaways:** Open-source platforms often prioritize accessibility over safety, leading to the democratization of harmful tools.
- **What could have been done better:** Hugging Face could have implemented more robust server-side image classification to block explicit outputs regardless of the model being used.
## Recommendations
- **Prevention:** Implement mandatory safety filters (NSFW detectors) on all image-generating Spaces.
- **Policy:** Explicitly ban "nudify" applications in Terms of Service and use automated scanning to identify and remove violating models.
- **Transparency:** Increase reporting on how many models are removed for NCII violations to deter malicious actors.