arXiv · 2609.32116
Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking
Abstract
Best-of-$N$ jailbreaking (BoN) bypasses safeguards of aligned models by drawing $N$ independent augmentations of an unsafe prompt and sampling $M$ completions of each. Previous works have shown that the attack success rate (ASR) seems to follow a power-law in $N$, which we challenge. The exponent drifts with $N$, with an exponential crossover which is a finite-size artifact of the adversarial dataset. Little work has been done to explore the entire two-budget ($N, M$) attack surface as well as its dependence on the generation temperature $T$. We introduce a simple barrier model where each prompt has a baseline safety level and each augmentation a random thermally activated barrier. Then four numbers, each backed by an interpretable safety mechanism, determine the entire ($N, M$) attack surface. They extrapolate predictions from $N \leq 100$ to $N = 10^4$, collapse five distinct models on the same scaling function and predict ASR at different temperatures from the one they were fitted at.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Marco Biroli. 2026-09-26. Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking. https://arxiv.org/abs/2609.32116
Cite the original work for its findings. Save a collection to share your selection of sources.