🤖 AI Summary
This study addresses the limitation that personality control in large language models (LLMs) relies on linear assumptions, which leads to feature interference and compositional bias. To overcome this, the work proposes modeling LLM personality representations as Riemannian manifolds within activation spaces, thereby transcending Euclidean geometric constraints. Specifically, it introduces a curvature-aware framework based on Ollivier-Ricci curvature estimation alongside a geodesic interpolation algorithm to achieve precise personality steering, and constructs a behavior-based BST benchmark suite for evaluation. The proposed approach significantly enhances multi-feature compositional consistency, demonstrates that geodesic distance more accurately predicts behavioral similarity, and enables more coherent transitions across intermediate personality states.
📝 Abstract
Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model's activation space and apply Euclidean operations---addition, scaling, and linear interpolation---under the linear representation hypothesis. However, these methods themselves report systematic failures: non-orthogonal trait dimensions, asymmetric ceiling and resistance effects, and significant deviations in multi-trait composition, suggesting that the linear isotropic assumption does not hold. We propose PersonaManifold, a framework that models persona representations as points on a curved, low-dimensional Riemannian submanifold in activation space. We estimate the manifold's intrinsic geometry---local metric tensors, geodesic distances, and Ollivier-Ricci curvature---and introduce geodesic steering, which interpolates between personas along manifold geodesics rather than Euclidean straight lines. We also propose the Behavioral Similarity Triplet (BST) benchmark, which automatically generates situational questions grounded in six established psychological constructs and defines persona similarity through behavioral responses rather than self-report questionnaires. Experiments on three open-source LLMs show that persona activations form a manifold with heterogeneous curvature, geodesic distance predicts behavioral similarity more accurately than Euclidean alternatives with independent contributions from anisotropy and curvature, and geodesic steering produces more coherent intermediate personas on both our BST benchmark and external evaluations, with the advantage concentrated in high-deviation regions where the manifold deviates most from flatness.