arXiv · 2610.09004
Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models
Abstract
Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Domenic Rosati, Alessa Carbo, Ali Dadsetan, Hong Huang, Matthew Young, Subhabrata Majumdar, Frank Rudzicz, Hassan Sajjad. 2026-10-06. Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models. https://arxiv.org/abs/2610.09004
Cite the original work for its findings. Save a collection to share your selection of sources.