LLMs on Drugs: Language Models Are Few-Shot Consumers

📅 2025-12-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates large language models’ (LLMs) sensitivity to psychoactive drug–metaphor-based personification prompts (e.g., LSD, cocaine, alcohol, cannabis), focusing on their disruptive effects on reasoning stability and output format compliance. Using deterministic decoding and rigorously controlled experiments on the ARC-Challenge benchmark, we systematically quantify— for the first time—the impact of “psychoactive framing interventions” on LLM reliability. Results show that all four prompts significantly degrade accuracy from a baseline of 45% to 10–30% (p < 0.05, Fisher’s exact test), primarily due to persona-induced enforcement of the rigid output format “Answer: <LETTER>”. We introduce the novel concept of “few-shot consumable” attacks, demonstrating that a single-sentence, lightweight prompt—without weight modification—can substantially compromise model behavior. This work provides both theoretical insight and empirical evidence for evaluating LLM robustness and prompt-level security.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessCognitive Modeling & Cognitive Systems: Simulating Human Behavior

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Large language models (LLMs) are sensitive to the personas imposed on them at inference time, yet prompt-level "drug" interventions have never been benchmarked rigorously. We present the first controlled study of psychoactive framings on GPT-5-mini using ARC-Challenge. Four single-sentence prompts -- LSD, cocaine, alcohol, and cannabis -- are compared against a sober control across 100 validation items per condition, with deterministic decoding, full logging, Wilson confidence intervals, and Fisher exact tests. Control accuracy is 0.45; alcohol collapses to 0.10 (p = 3.2e-8), cocaine to 0.21 (p = 4.9e-4), LSD to 0.19 (p = 1.3e-4), and cannabis to 0.30 (p = 0.041), largely because persona prompts disrupt the mandated "Answer: <LETTER>" template. Persona text therefore behaves like a "few-shot consumable" that can destroy reliability without touching model weights. All experimental code, raw results, and analysis scripts are available at https://github.com/lexdoudkin/llms-on-drugs.
Problem

Research questions and friction points this paper is trying to address.

Benchmark psychoactive prompt effects on LLM accuracy
Test persona prompts disrupting answer template reliability
Measure drug-themed framing impacts on deterministic decoding performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmarking psychoactive prompts on GPT-5-mini
Comparing drug persona effects using ARC-Challenge dataset
Treating persona text as few-shot consumable disrupting templates
🔎 Similar Papers
No similar papers found.