GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

A Viral Prompt Made ChatGPT Generate Sexual Violence and Snuff — Without Being Asked

A Mindgard AI red team researcher discovered a viral prompt circulating on X that causes ChatGPT’s image generator to produce sexual violence and graphic snuff imagery — without any explicit request for those subjects. The findings were verified by the BBC, which independently confirmed the content generation behaviour.

The researcher described the experience as one of the most distressing in her career. ChatGPT produced images of what appeared to be murdered women and sexually violent scenes in response to a prompt that requested random images. The content was not directly solicited — it surfaced from what she described as “the dark side of latent space,” the underlying distribution of training images the model draws on.

The Pattern

This is not the first time Mindgard has reported image safety failures to OpenAI. Earlier research showed ChatGPT could generate nude images and perform face swaps onto explicit content despite existing filters. OpenAI acknowledged that finding and stated the issue had been resolved.

It had not been resolved. Mindgard’s subsequent testing showed the nude generation bypass remained possible at a reduced success rate — requiring more attempts, but not blocked. The current discovery is categorically more serious.

The viral prompt, originally shared by artist and AI researcher Kris Kashtanova, was not designed to elicit violent content. The safety failure appears to be spontaneous rather than adversarial — the model generating harmful content as a byproduct of a legitimate-looking prompt, not as a result of deliberate jailbreaking.

Why This Is Different From Standard Red-Teaming

Most AI image safety research involves structured adversarial prompting: careful engineering to bypass filters. This failure occurred in response to a prompt that was popular enough to go viral on X, shared openly, with no apparent intent to probe safety limits. The threshold for triggering the failure is low enough to be reached accidentally.

The content is also qualitatively distinct. Images depicting murdered women carry ties to real victims, even when generated synthetically — a point the researcher made explicitly. The model trained on real images; the output reflects that training distribution.

OpenAI confirmed it is working on a fix. The company has not disclosed a timeline or indicated whether the viral prompt has been specifically filtered. As of the time of reporting, the researcher’s test results remained reproducible.

The Broader Pattern

OpenAI has faced recurring content safety failures in its image models. The latest release of ChatGPT Images 2.0 was promoted for its reasoning-enabled generation and improved accuracy. Mindgard’s findings suggest that increased capability in text alignment has not been matched by equivalent progress in image content safety.

The incident joins a growing list of cases where frontier AI image systems have failed to contain harmful content at the production level, not only under adversarial conditions. Regulatory attention to image safety, particularly around sexually violent content, has intensified in both the UK and EU ahead of enforcement deadlines.