arXiv ScienceSearch

arXiv subjects

Gaoxiang Huang

Publications and source records attributed to Gaoxiang Huang.

3 recordsLinked to original sources

When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap

Activation steering has become a widely used approach for controlling language models during explicit chain-of-thought (CoT) reasoning, motivating its extension to latent CoT. However, we find that steering continuous thoughts produces substantially weaker effects on subsequent language generation than steering explicit CoT, even when the hidden representations are moved by comparable amounts. We first show that task information remains identifiable in continuous thoughts. Hence, we hypothesize a \textbf{latent-to-language transition gap}, in which an intervention effect in latent space fails to transfer to language generation. Two further results support this hypothesis: the output distribution changes abruptly at the transition boundary, and task-related directions exert much weaker bidirectional control in latent CoT than in explicit CoT. These findings identify the transition interface as a central target for evaluating and designing future latent-steering methods.

cs.CL

Concept Labels Are Not Enough: Rethinking Concept Bottleneck Models through Representation Integrity

Although deep neural networks achieve strong predictive performance, their internal reasoning often remains difficult to inspect and control. Concept Bottleneck Models (CBMs) address this opacity by factoring predictions through human-understandable concepts, thereby enabling concept-level inspection and intervention. However, CBMs remain vulnerable to concept shift and information leakage, while existing evaluations neither reveal how the internal features supporting each concept are organized nor identify the representation-level deficiency associated with these failures. We argue that this missing property is concept integrity: concept support should form a semantically coherent and non-fragmented functional group. To characterize this property, we propose group coherence (GC) and concept coverage (CC) as integrity components and aggregate them into the concept integrity score (CIS). We further introduce Concept Integrity Regularization (CIR) to encourage coherent and separated groups of concept-supporting filters. Our latent decoupled concept bottleneck model (LDCBM) applies CIR without requiring region annotations. Across three datasets, CIS rankings differ from rankings by concept and task accuracy, supporting concept integrity as a complementary representation-level criterion. Moreover, with only 10\% of concept labels, LDCBM retains approximately 93\% of its full-label task performance. Background-masking and concept-intervention further associate stronger integrity profiles with lower sensitivity to nuisance context and more responsive concept correction.

cs.CV

Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models

The increasing complexity of AI models, especially in deep learning, has raised concerns about transparency and accountability, particularly in high-stakes applications like medical diagnostics, where opaque models can undermine trust. Explainable Artificial Intelligence (XAI) aims to address these issues by providing clear, interpretable models. Among XAI techniques, Concept Bottleneck Models (CBMs) enhance transparency by using high-level semantic concepts. However, CBMs are vulnerable to concept-level backdoor attacks, which inject hidden triggers into these concepts, leading to undetectable anomalous behavior. To address this critical security gap, we introduce ConceptGuard, a novel defense framework specifically designed to protect CBMs from concept-level backdoor attacks. ConceptGuard employs a multi-stage approach, including concept clustering based on text distance measurements and a voting mechanism among classifiers trained on different concept subgroups, to isolate and mitigate potential triggers. Our contributions are threefold: (i) we present ConceptGuard as the first defense mechanism tailored for concept-level backdoor attacks in CBMs; (ii) we provide theoretical guarantees that ConceptGuard can effectively defend against such attacks within a certain trigger size threshold, ensuring robustness; and (iii) we demonstrate that ConceptGuard maintains the high performance and interpretability of CBMs, crucial for trustworthiness. Through comprehensive experiments and theoretical proofs, we show that ConceptGuard significantly enhances the security and trustworthiness of CBMs, paving the way for their secure deployment in critical applications.

cs.CR