Semantic Primes as Explanans for Emotion in Large Language Models

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current emotion interpretation frameworks for large language models often suffer from circularity and arbitrariness. This work addresses this limitation by introducing semantic primes from Natural Semantic Metalanguage (NSM) as a foundational, non-circular basis for interpretable emotion representation. The proposed approach is evaluated through neural representation probing, directional interventions, and semantic equivalence tests across multiple models, including Llama-1B, Gemma-2B/9B, and OLMo-7B. Experimental results demonstrate that NSM primes are highly recoverable within model representations; emotion interventions grounded in these primes yield approximately threefold gains in control strength and twofold improvements in selectivity; furthermore, models treat prime-based paraphrases and original emotion terms as semantically equivalent.
📝 Abstract
Progresses have been made on understanding emotion mechanisms of large language models (LLMs). However, how to explain emotion in LLMs, or even what constitutes good explanations, are less clear. Emotion representations, components, circuits are widely recoverable, but as explanations of a model's own computation they are circular; the emotion space dimensions tend to be arbitrary and non-terminating. A pressing question to ask is whether a more primitive set of internal variables does the work: the semantic primes of the Natural Semantic Metalanguage (NSM). Across four instruction-tuned LLMs (Llama-1B, Gemma-2B, Gemma-9B, OLMo-7B), experiments show that the NSM primes are (1) recoverable internal elements; and (2) on the reference model, intervening with a prime based direction controls emotion about three times as strongly, and twice as selectively, as the best appraisal based direction; and (3) the model treats a prime based explication as interchangeable with the corresponding emotion. These evidences suggest that NSM primes seem to be better explanans for emotion in LLMs than many alternative options according to scientific explanations criteria.
Problem

Research questions and friction points this paper is trying to address.

emotion
explanation
large language models
semantic primes
Natural Semantic Metalanguage
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic primes
Natural Semantic Metalanguage
emotion explanation
large language models
model interpretability