🤖 AI Summary
This study addresses the vulnerability of open-weight models to fine-tuning attacks, demonstrating that merely removing harmful information is insufficient and that existing mutual information invariants are constrained by the fastest parameterization path. To overcome these limitations, this work proposes a function-preserving reparameterization theory, validated through weight-data and label-representation mutual information analysis, gradient geometry reconstruction, and controlled experiments. The authors prove that preserving representation-level independence maintains the full Jacobian matrix and construct counterexamples exhibiting zero mutual information yet one-step recovery, revealing that recovery time depends on training order and parameterization schemes. Ultimately, this research argues that safety certification should constrain attack dynamics rather than relying solely on mutual information measured at the time of model release.
📝 Abstract
Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.