Persona Dosing: Calibrated Activation Steering for Graded Trait Control

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the absence of behavioral scales for precisely modulating the expression intensity of personality traits in large language models by proposing PersonaDose. This method integrates a descriptive conditioning controller (FLAS) with flow-time calibration, decoupling the controller’s learning range from query precision to enable graded control of personality traits based on target intensities without requiring paired target-intensity data. Experimental results demonstrate that PersonaDose significantly improves the expression accuracy of core personality traits on models such as Llama, achieving an average positioning error of only 4.7–6.2 points and outperforming activation addition baselines.
πŸ“ Abstract
An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a requested mean intensity. PersonaDose specializes a shared, description-conditioned FLAS controller on persona responses, then calibrates its flow time against measured trait expression. Training responses are not paired with requested target intensities. Across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition. Calibration-selected settings retain an expression advantage on held-out questions, although the coherence floor does not hold for every trait there. Across seven trained traits, calibrated requests yield mean targeting errors of 4.7-6.2 points over 14-22 calibration-reachable targets out of 28 per model. These results separate the behavioral range learned by a controller from the accuracy of requests within that range.
Problem

Research questions and friction points this paper is trying to address.

persona dosing
activation steering
trait control
language model
calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Persona Dosing
Activation Steering
Calibrated Control
FLAS Controller
Trait Expression
πŸ’Ό Related Jobs
No related jobs found.
Z
Zehao Jin
Georgia Institute of Technology
J
Junran Wang
Georgia Institute of Technology
R
Ruixuan Deng
Georgia Institute of Technology
Jiahao Chen
Jiahao Chen
Zhejiang University
AI SecurityTrustworthy AIGenAI SecurityGenAI Privacy
J
Jingyuan Zhang
Georgia Institute of Technology
Y
Yuxuan Zhang
University of British Columbia
Xinjie Shen
Xinjie Shen
PhD Student, Georgia Tech
Human-AI CollaborationNetworksHCI