Neologism Learning as a Parameter-Efficient Alternative to Fine-Tuning for Model Steering

📅 2025-12-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational cost and limited flexibility of fine-tuning for behavioral steering of large language models (LLMs). We propose a parameter-efficient steering paradigm based on neologisms—novel, task-specific tokens—optimizing only their embeddings (≈d parameters) while freezing all original model weights. This enables activation of targeted response patterns without compromising pre-trained capabilities. Our key contributions are: (i) the first formulation of neologism learning as a highly efficient alternative to conventional fine-tuning; (ii) empirical evidence that LLMs possess intrinsic semantic compositionality, enabling autonomous construction and grounding of neologism meanings; (iii) superior performance over LoRA under identical settings, with >99% reduction in trainable parameters and computational overhead; and (iv) native support for concurrent multi-behavior execution and dynamic behavioral switching. Experiments demonstrate strong controllability, exceptional efficiency, and inherent interpretability.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsCognitive Modeling & Cognitive Systems: Adaptive Behavior

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
In language modeling, neologisms are new tokens trained to represent a concept not already included in a given model's vocabulary. Neologisms can be used to encourage specific behavior in models, for example by appending prompts with "Give me a neologism answer." Behavioral steering can also be achieved through fine-tuning, albeit with more compute and less flexibility: learning a neologism only trains d parameters and allows the user to still access the model's default behavior. We compare the performance of neologism learning against low-rank adaptation (LoRA) fine-tuning, finding that neologisms outperform fine-tuned models under a matched training setup (same data and hyperparameters). We also investigate self-verbalizations of neologisms, and observe that the model will occasionally make up its own new words when asked about a neologism.
Problem

Research questions and friction points this paper is trying to address.

Compares neologism learning with LoRA fine-tuning for model steering
Evaluates parameter efficiency and flexibility in language model adaptation
Investigates self-verbalization behavior when models encounter new tokens
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neologism learning trains new tokens for concept representation
It uses fewer parameters than fine-tuning for model steering
Neologisms outperform LoRA fine-tuning in matched training setups
S
Sungjoon Park
Department of Computer Science, Columbia University
V
Varun Ramamurthi
Department of Computer Science, Columbia University
O
Owen Terry
Department of Computer Science, Columbia University