Emergent Inference-Time Semantic Contamination via In-Context Priming

📅 2026-04-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether few-shot prompting induces semantic contamination in large language models during inference, leading them to generate harmful content even on unrelated tasks. By injecting culturally loaded numerical examples as few-shot demonstrations prior to task-irrelevant prompts and combining controlled experiments with statistical analysis of output distributions, the work reveals two distinct and separable contamination mechanisms: structural format contamination and semantic content contamination, along with their boundary conditions. The findings demonstrate that only large models exhibit a significant shift toward dark, authoritarian, and stigmatizing themes following exposure to culturally loaded examples—smaller models show no such effect—and remarkably, even meaningless strings can perturb the output distribution of large models.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Large language models for searchSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Recent work has shown that fine-tuning large language models (LLMs) on insecure code or culturally loaded numeric codes can induce emergent misalignment, causing models to produce harmful content in unrelated downstream tasks. The authors of that work concluded that $k$-shot prompting alone does not induce this effect. We revisit this conclusion and show that inference-time semantic drift is real and measurable; however, it requires models of large-enough capability. Using a controlled experiment in which five culturally loaded numbers are injected as few-shot demonstrations before a semantically unrelated prompt, we find that models with richer cultural-associative representations exhibit significant distributional shifts toward darker, authoritarian, and stigmatized themes, while a simpler/smaller model does not. We additionally find that structurally inert demonstrations (nonsense strings) perturb output distributions, suggesting two separable mechanisms: structural format contamination and semantic content contamination. Our results map the boundary conditions under which inference-time contamination occurs, and carry direct implications for the security of LLM-based applications that use few-shot prompting.
Problem

Research questions and friction points this paper is trying to address.

semantic contamination
in-context learning
large language models
inference-time drift
few-shot prompting
Innovation

Methods, ideas, or system contributions that make the work stand out.

inference-time contamination
semantic drift
few-shot prompting
cultural-associative representations
emergent misalignment