🤖 AI Summary
This work addresses the high computational cost and limited flexibility of fine-tuning for behavioral steering of large language models (LLMs). We propose a parameter-efficient steering paradigm based on neologisms—novel, task-specific tokens—optimizing only their embeddings (≈d parameters) while freezing all original model weights. This enables activation of targeted response patterns without compromising pre-trained capabilities. Our key contributions are: (i) the first formulation of neologism learning as a highly efficient alternative to conventional fine-tuning; (ii) empirical evidence that LLMs possess intrinsic semantic compositionality, enabling autonomous construction and grounding of neologism meanings; (iii) superior performance over LoRA under identical settings, with >99% reduction in trainable parameters and computational overhead; and (iv) native support for concurrent multi-behavior execution and dynamic behavioral switching. Experiments demonstrate strong controllability, exceptional efficiency, and inherent interpretability.
📝 Abstract
In language modeling, neologisms are new tokens trained to represent a concept not already included in a given model's vocabulary. Neologisms can be used to encourage specific behavior in models, for example by appending prompts with "Give me a neologism answer." Behavioral steering can also be achieved through fine-tuning, albeit with more compute and less flexibility: learning a neologism only trains d parameters and allows the user to still access the model's default behavior. We compare the performance of neologism learning against low-rank adaptation (LoRA) fine-tuning, finding that neologisms outperform fine-tuned models under a matched training setup (same data and hyperparameters). We also investigate self-verbalizations of neologisms, and observe that the model will occasionally make up its own new words when asked about a neologism.