arXiv · 2605.04992
You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
Abstract
Public repositories now distribute thousands of specialized Low-Rank Adaptation (LoRA) modules, but a third-party adapter routinely arrives without the refusal behavior the same adapter would have had if it had been trained responsibly. Restoring it by fine-tuning on safety data induces the opposite failure: the domain expertise the adapter was published to provide degrades. A practitioner who downloads a finished adapter holds neither the domain corpus nor a safety corpus, so neither remedy is available. We propose Neural Weight Translation (NeWTral), which maps unsafe domain adapters to their safety-aligned counterparts entirely within parameter space while preserving their expertise. NeWTral is a non-linear translator pre-trained on unsafe-to-safe adapter pairs, with a Mixture of Experts router that blends a high-fidelity surgical translator and an aggressive alignment expert per tensor. Across four architectural families (Llama, Mistral, Qwen, Gemma) at scales up to 72B parameters and eight professional domains, NeWTral cuts the average Attack Success Rate from 70% to 13% while retaining 88% knowledge fidelity. The effect reproduces on JailbreakBench, on 12 of 14 adapters downloaded from a public model hub, and under automated jailbreak attacks. Curing is a single pass over the weights, needing no original training data and no retraining.
Explore related subjects
Keep this discovery
Marco Arazzi, Vignesh Kumar Kembu, Antonino Nocera, Stjepan Picek, Saraga Sakthidharan. 2026-05-06. You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation. https://arxiv.org/abs/2605.04992
Cite the original work for its findings. Save a collection to share your selection of sources.