🤖 AI Summary
This work addresses the limitation of existing large language models in personality control, which predominantly rely on static traits such as the Big Five while neglecting dynamic cognitive processes. For the first time, it integrates Jungian eight-function cognitive typology into an activation intervention framework, constructing a role-playing narrative dataset and evaluation protocol to extract and modulate cognitive-function-specific intervention vectors in the Llama-3.1-8B model. The study reveals that personality-related information is concentrated in intermediate model layers, and the intervention vectors exhibit a geometric structure wherein multidimensional directions are not linearly composable. Experiments demonstrate successful monotonic and controllable modulation across all eight cognitive functions, uncovering the structured yet nonlinear nature of personality representations in activation space.
📝 Abstract
Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled as a set of cognitive processes using the eight Jungian Cognitive Functions. To this end, we introduce a framework comprising a Jungian evaluation protocol and a dataset of over 2,100 role-playing character narrations.
Activation steering vector extraction and evaluation experiments on Llama-3.1-8B demonstrate effective monotonic control over all eight cognitive functions through activation steering. Beyond controllability, our analysis reveals that: 1. personality information is concentrated in middle transformer layers; 2. steering vectors exhibit structured geometric relationships consistent with distinctions between rational and irrational functions; 3. effective multi-dimensional steering directions cannot be recovered as linear combinations of single-function directions. These findings provide new insights into the representation of personality in LLM activation space and establish a framework for studying interpretable, effective, and multi-dimensional personality control.