Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of activation steering in conventional text-to-image models, which relies on linear representation assumptions and struggles to accurately capture nonlinear intermediate states between visual concepts. To overcome this, we propose KANSteer, a framework that pioneers the integration of Kolmogorov-Arnold Networks into the activation space analysis of diffusion Transformers. By transcending linear directional constraints, our approach formulates concept traversal as smooth curve trajectories along one-dimensional coordinates. Experimental results demonstrate that KANSteer effectively fits activation paths deviating from straight lines, enabling human-intuitive nonlinear concept transitions and more precise attribute control. This work establishes a novel paradigm for interpretable generative modeling.
📝 Abstract
Mechanistic interpretability often relies on the Linear Representation Hypothesis (LRH), which assumes that high-level concepts are encoded as linear directions in activation space. Yet a natural visual concept does not necessarily require a linear visual transition: between sunny and stormy lies an intermediate weather state such as a sky with a few white clouds, not simply a weaker storm; between a caterpillar and a butterfly, the progression is not a caterpillar with continuously growing wings. This raises the question of whether such true intermediate states are also represented nonlinearly by the model. Indeed, when we prompt text-to-image models directly for intermediate attributes, their activations rarely fall along the straight direction connecting the endpoints. Therefore, we propose KANSteer, which models concept traversal as a curve passing through its intermediate states. Seeking a representation that is both simple and interpretable, we propose to use Kolmogorov-Arnold Networks (KANs), which provide a one-dimensional coordinate whose learned functions define the trajectory. This allows the steering direction to vary along the concept while preserving an interpretable representation. Across several concepts and text-to-image diffusion transformers, we find that their activation trajectories substantially deviate from straight lines, and that KANSteer provide a closer fit and smoother traversal of intermediate attributes than linear steering.
Problem

Research questions and friction points this paper is trying to address.

Linear Representation Hypothesis
Mechanistic Interpretability
Text-to-Image Models
Non-linear Activation Steering
Concept Traversal
Innovation

Methods, ideas, or system contributions that make the work stand out.

Non-Linear Activation Steering
Kolmogorov-Arnold Networks
Mechanistic Interpretability
Text-to-Image Models
Concept Traversal
🔎 Similar Papers
No similar papers found.