You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation
Public repositories now distribute thousands of specialized Low-Rank Adaptation (LoRA) modules, but a third-party adapter routinely arrives without the refusal behavior the same adapter would have had if it had been trained responsibly. Restoring it by fine-tuning on safety data induces the opposite failure: the domain expertise the adapter was published to provide degrades. A practitioner who downloads a finished adapter holds neither the domain corpus nor a safety corpus, so neither remedy is available. We propose Neural Weight Translation (NeWTral), which maps unsafe domain adapters to their safety-aligned counterparts entirely within parameter space while preserving their expertise. NeWTral is a non-linear translator pre-trained on unsafe-to-safe adapter pairs, with a Mixture of Experts router that blends a high-fidelity surgical translator and an aggressive alignment expert per tensor. Across four architectural families (Llama, Mistral, Qwen, Gemma) at scales up to 72B parameters and eight professional domains, NeWTral cuts the average Attack Success Rate from 70% to 13% while retaining 88% knowledge fidelity. The effect reproduces on JailbreakBench, on 12 of 14 adapters downloaded from a public model hub, and under automated jailbreak attacks. Curing is a single pass over the weights, needing no original training data and no retraining.