🤖 AI Summary
This study investigates large language models’ (LLMs) sensitivity to psychoactive drug–metaphor-based personification prompts (e.g., LSD, cocaine, alcohol, cannabis), focusing on their disruptive effects on reasoning stability and output format compliance. Using deterministic decoding and rigorously controlled experiments on the ARC-Challenge benchmark, we systematically quantify— for the first time—the impact of “psychoactive framing interventions” on LLM reliability. Results show that all four prompts significantly degrade accuracy from a baseline of 45% to 10–30% (p < 0.05, Fisher’s exact test), primarily due to persona-induced enforcement of the rigid output format “Answer: <LETTER>”. We introduce the novel concept of “few-shot consumable” attacks, demonstrating that a single-sentence, lightweight prompt—without weight modification—can substantially compromise model behavior. This work provides both theoretical insight and empirical evidence for evaluating LLM robustness and prompt-level security.
📝 Abstract
Large language models (LLMs) are sensitive to the personas imposed on them at inference time, yet prompt-level "drug" interventions have never been benchmarked rigorously. We present the first controlled study of psychoactive framings on GPT-5-mini using ARC-Challenge. Four single-sentence prompts -- LSD, cocaine, alcohol, and cannabis -- are compared against a sober control across 100 validation items per condition, with deterministic decoding, full logging, Wilson confidence intervals, and Fisher exact tests. Control accuracy is 0.45; alcohol collapses to 0.10 (p = 3.2e-8), cocaine to 0.21 (p = 4.9e-4), LSD to 0.19 (p = 1.3e-4), and cannabis to 0.30 (p = 0.041), largely because persona prompts disrupt the mandated "Answer: <LETTER>" template. Persona text therefore behaves like a "few-shot consumable" that can destroy reliability without touching model weights. All experimental code, raw results, and analysis scripts are available at https://github.com/lexdoudkin/llms-on-drugs.