🤖 AI Summary
This study addresses the limitation of fixed steering vectors in flow models, which fail to adapt to evolving generative states and consequently induce unintended global alterations. To overcome this, we propose an adaptive steering field method that generalizes static vectors into dynamic fields, progressively re-estimating steering directions based on noise states along the generation trajectory. This approach enables model-agnostic and inversion-free adaptive estimation while supporting both concept induction and suppression, preserving local structures without requiring spatial masks. Experimental results demonstrate that our method establishes a new state-of-the-art on safety steering benchmarks and achieves superior semantic fidelity in unsupervised image editing tasks.
📝 Abstract
As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected activations. While functional, a fixed and example-agnostic vector applied uniformly along the entire trajectory cannot adapt to the changing state of the generation and often causes unintended global changes. We introduce Steering Fields, a generalization of steering vectors that adaptively re-estimates the steering direction at each step of the generative process. Steering Fields operate on the noisy states of flow models, expose a continuous trade-off between steering strength and content preservation, and are compositional, enabling the simultaneous induction and inhibition of concepts, setting a new state of the art on safety steering benchmarks. Despite using no explicit spatial masks or object priors, the trajectory-adaptive estimation naturally preserves local structure, in a manner reminiscent of image editing. In fact, Steering Fields can serve as a structure-preserving image-editing technique that achieves state-of-the-art semantic fidelity (CLIP, VQAScore), while remaining model-agnostic and inversion-free.