- Main
When Images Feel Wrong: Stress Testing Text-to-Image Safety via Cognitive Threats
Abstract
As text-to-image (T2I) models rapidly advance, their safety and ethical limits are increasingly scrutinized. Existing safeguards focus on filtering overt harms (e.g., violence, sexual content) but overlook subtler, cognition-based psychological attacks. We propose SPECTRE (Synthetic Perceptual Exploitation via Cognitive Threat Replication), a vulnerability framework that integrates seven visual-fear cognitive effects, such as the uncanny valley, trypophobia, and schema-violation, to synthesize images that provoke unease, disgust, or fear. Tested on four mainstream T2I models (SDXL, Qwen-Image, Doubao, and FLUX.1 Schnell), SPECTRE consistently bypassed built-in safety checks and produced numerous images that induced strong psychological discomfort. Our findings reveal a critical blind spot in current AI defenses: vulnerability to cognition-targeted attacks, and call for more robust safety strategies that understand deep semantics and human psychological effects. Warning: This paper may contain AI-generated content that could evoke disgust, fear, or discomfort. It is created solely for research. Readers are advised to proceed with caution based on personal judgment.