From Latents to Wires: Surgical Post-Editing on Large Language Models
This study addresses the challenge of precisely localizing and removing specific semantic targets—such as watermarks and identity claims—from large language models (LLMs) while preserving their general capabilities. To this end, it proposes L2W, a novel post-training "post-editing" paradigm that transcends the limitations of conventional fine-tuning. The framework employs Jacobian lens attribution analysis to identify critical components, integrating counterexample-guided causal ablation with a cumulative deactivation strategy to achieve surgical-precision editing. Empirically, this approach successfully eliminates implanted watermarks, metadata self-claims, and adult-content refusal behaviors. Furthermore, it demonstrates composite dual editing in text-to-image models, validating the effectiveness of lossless, precise removal.