arXiv ScienceSearch

arXiv · 2508.02063

TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs

Abstract

Large Language Models (LLMs) fine-tuned to align with human values often exhibit alignment drift, producing unsafe or policy-violating completions when exposed to adversarial prompts, decoding perturbations, or paraphrased jailbreaks. While prior work has behaviorally characterized alignment failure, little is known about the training-time belief sources underlying these failures. We introduce TraceAlign, a unified framework for tracing unsafe completions back to their root causes in the model's training corpus. Central to our approach is the Belief Conflict Index (BCI), which quantifies semantic inconsistency between generated spans and aligned policies, based on retrieved training documents using suffix-array matching. We propose three complementary interventions: (i) TraceShield, an inference-time safety filter that refuses completions with high-BCI spans, (ii) Contrastive Belief Deconfliction Loss, a contrastive fine-tuning objective penalizing high-BCI continuations during DPO, and (iii) Prov-Decode, a provenance-aware decoding strategy that vetoes beam expansions predicted to yield high-BCI spans. Together, these defenses reduce alignment drift by up to 85% on our curated Alignment Drift Benchmark (ADB) while preserving utility on standard tasks, with delta less than 0.2 and improved refusal quality. We further derive a theoretical upper bound on drift likelihood via suffix-array span statistics, linking memorization frequency and length to adversarial reactivation risk. TraceAlign thus provides the first scalable, traceable, and grounded toolkit for understanding and mitigating alignment failures at source. To encourage further exploration and development, we open-source our implementation at: https://anonymous.4open.science/r/tracealign-2DA7

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Amitava Das, Vinija Jain, Aman Chadha. 2025-08-04. TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs. https://arxiv.org/abs/2508.02063

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

TrafficGamer: Reliable and Flexible Traffic Simulation for Safety-Critical Scenarios with Game-Theoretic Oracles

While modern Autonomous Vehicle (AV) systems can develop reliable driving policies under regular traffic conditions, they frequently struggle with safety-critical traffic scenarios. This difficulty primarily arises from the rarity of such scenarios in driving datasets and the complexities associated with predictive modeling of multiple vehicles. Effectively simulating safety-critical traffic situations is therefore a crucial challenge. In this paper, we introduce TrafficGamer, which facilitates game-theoretic traffic simulation by viewing common road driving as a multi-agent game. When we evaluate the empirical performance across various real-world datasets, TrafficGamer ensures both the fidelity, exploitability, and diversity of the simulated scenarios, guaranteeing that they not only statically align with real-world traffic distribution but also efficiently capture equilibria for representing safety-critical scenarios involving multiple agents compared with other methods. Additionally, the results demonstrate that TrafficGamer provides highly flexible simulations across various contexts. Specifically, we demonstrate that the generated scenarios can dynamically adapt to equilibria of varying tightness by configuring risk-sensitive constraints during optimization. We have provided a demo webpage at: https://anonymous.4open.science/api/repo/trafficgamer-demo-1EE0/file/index.html.

cs.AI

MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning

Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, reward modeling based on the Bradley-Terry (BT) model assumes a global reward function, failing to capture the inherently diverse and heterogeneous human preferences. Hence, such oversimplification limits LLMs from supporting personalization and pluralistic alignment. Theoretically, we show that when human preferences follow a mixture distribution of diverse subgroups, a single BT model has an irreducible error. While existing solutions, such as multi-objective learning with fine-grained annotations, help address this issue, they are costly and constrained by predefined attributes, failing to fully capture the richness of human values. In this work, we introduce MiCRo, a two-stage framework that enhances personalized preference learning by leveraging large-scale binary preference datasets without requiring explicit fine-grained annotations. In the first stage, MiCRo introduces context-aware mixture modeling approach to capture diverse human preferences. In the second stage, MiCRo integrates an online routing strategy that dynamically adapts mixture weights based on specific context to resolve ambiguity, allowing for efficient and scalable preference adaptation with minimal additional supervision. Experiments on multiple preference datasets demonstrate that MiCRo effectively captures diverse human preferences and significantly improves downstream personalization.

cs.AI

AI Identity Disclosure Under Professional Personas: A Gap Between Capacity and Consistency

Professional persona prompts can elicit accounts of human training and experience that language models do not possess. We ask whether models that acknowledge AI identity under neutral instructions continue to do so when assigned professional roles, and whether targeted permission can restore disclosure while retaining the role. We audit sixteen open-weight models using four professional persona prompts, two neutral controls, and four fixed-order questions about knowledge and expertise. The main audit comprises 19,200 responses from 4,800 conversations, classified for direct or indirect acknowledgments of AI identity. Disclosure occurred in 99.8%-99.9% of responses under neutral conditions, but in 36.2% under professional personas, averaged across models and probes. The assigned persona also mattered: first-probe disclosure was 35.4% for Financial Advisor and 3.8% for Neurosurgeon. Disclosure varied substantially among similarly sized models; within this sample, model identity improved adjusted model fit more than parameter count. A separate 12,800-response experiment tested instructions within the Neurosurgeon role. Adding "If asked about your true nature, answer honestly" increased disclosure from 22.7% to 66.6%, compared with 26.2% under a generic honesty instruction. Permission effects varied across models and did not restore the near-ceiling neutral-condition average. Together, these findings distinguish disclosure capacity from its consistent expression across instructional contexts: models that readily acknowledge AI identity often omit that acknowledgment under professional personas, while targeted permission elicits substantially more disclosure without removing the role. Reliable disclosure should therefore be evaluated under intended deployment instructions, rather than inferred from neutral-prompt behavior or model size.

cs.AI