arXiv ScienceSearch

arXiv subjects

Zirui He

Publications and source records attributed to Zirui He.

13 recordsLinked to original sources

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity

Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framework for mitigating shortcut learning in reward model training. Unlike static shortcut heuristics, DynaCF measures shortcut sensitivity online during optimization by applying semantics-preserving counterfactual perturbations and tracking the resulting margin shifts and preference flips under the current model. Samples with higher shortcut sensitivity are dynamically downweighted in the Bradley-Terry objective, encouraging the model to rely less on superficial patterns and more on task-relevant preference signals. Extensive experiments show that DynaCF consistently improves robustness in preference modeling.

cs.LG

Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation

Activation verbalization explains hidden representations in natural language, but existing methods are mostly limited to self-explanation, where each model explains only its own activations. We introduce Universal Activation Verbalizer (UAV), a framework that uses a shared decoder to explain activations from heterogeneous donor models. UAV learns a lightweight adapter that converts donor activations into soft tokens in decoder's embedding space, and further supports adapter-only transfer by reusing a frozen decoder-side LoRA while training only a new adapter for another donor. Across classification, fact retrieval, and gist summarization, UAV remains competitive with strong self-explanation baselines while enabling cross-model verbalization across model families and scales. Ablations show that decoder-side tuning mainly improves task behavior, whereas the adapter provides the activation-grounded factual and semantic information needed for faithful explanations. Code and data are available at https://github.com/hy-zhao23/ActExp.

cs.CL

Hybrid functional calculation of electrical activity and complexing mechanism of Cu-related defects

Copper is a detrimental impurity in silicon with high diffusivity and a high tendency to precipitate. Interaction between Cu and other defects is essential for understanding the nature of Cu precipitation in silicon. Despite extensive experimental investigations of Cu-related defects in silicon, a comprehensive understanding remains elusive due to limitations of techniques in resolving defect configurations, as well as inconsistencies between theoretical and experimental results regarding transition levels. Moreover, the underlying formation mechanism of the well-known $\mathrm{Cu_{PL}}$ line is still unclear. In this work, configurations, formation energies, and transition levels of Cu-related defects in silicon are calculated using the HSE06 functional and finite-size correction. Defects involved in this study include $\mathrm{Cu_i}$, $\mathrm{Cu_{Si}}$, Cu-B, Cu-P, and Cu-H. A $\mathrm{Cu_{i4}V}$ model is proposed to explain the discrepancies between theory and experiment about $\mathrm{Cu_{PL}}$ defect. Our calculations may provide insight into the electrically active defects and the early states of Cu precipitation in silicon.

cond-mat.mtrl-sci

RealRoute: Dynamic Query Routing System via Retrieve-then-Verify Paradigm

Despite the success of Retrieval-Augmented Generation (RAG) in grounding LLMs with external knowledge, its application over heterogeneous sources (e.g., private databases, global corpora, and APIs) remains a significant challenge. Existing approaches typically employ an LLM-as-a-Router to dispatch decomposed sub-queries to specific sources in a predictive manner. However, this "LLM-as-a-Router" strategy relies heavily on the semantic meaning of different data sources, often leading to routing errors when source boundaries are ambiguous. In this work, we introduce RealRoute System, a framework that shifts the paradigm from predictive routing to a robust Retrieve-then-Verify mechanism. RealRoute ensures \textit{evidence completeness through parallel, source-agnostic retrieval, followed by a dynamic verifier that cross-checks the results and synthesizes a factually grounded answer}. Our demonstration allows users to visualize the real-time "re-routing" process and inspect the verification chain across multiple knowledge silos. Experiments show that RealRoute significantly outperforms predictive baselines in the multi-hop Rag reasoning task. The RealRoute system is released as an open-source toolkit with a user-friendly web interface. The code is available at the URL: https://github.com/Joseph1951210/RealRoute.

cs.IR

FinAnchor: Aligned Multi-Model Representations for Financial Prediction

Financial prediction from long documents involves significant challenges, as actionable signals are often sparse and obscured by noise, and the optimal LLM for generating embeddings varies across tasks and time periods. In this paper, we propose FinAnchor(Financial Anchored Representations), a lightweight framework that integrates embeddings from multiple LLMs without fine-tuning the underlying models. FinAnchor addresses the incompatibility of feature spaces by selecting an anchor embedding space and learning linear mappings to align representations from other models into this anchor. These aligned features are then aggregated to form a unified representation for downstream prediction. Across multiple financial NLP tasks, FinAnchor consistently outperforms strong single-model baselines and standard ensemble methods, demonstrating the effectiveness of anchoring heterogeneous representations for robust financial prediction.

cs.CL

Exceptionally high carrier mobility in hexagonal diamond

Hexagonal diamond (h-diamond), or Lonsdaleite, is a promising wide-bandgap semiconductor known for its high thermal conductivity and hardness. Based on \textit{ab initio} calculations, we demonstrate its exceptionally high carrier mobilities. At room temperature, the hole mobilities along the $\perp c$ and $\parallel c$ directions are 6000 and 6024 cm$^{2}$V$^{-1}$s$^{-1}$, respectively, while the corresponding electron mobilities reach 12339 and 28473 cm$^{2}$V$^{-1}$s$^{-1}$. These values are significantly superior to those of most known semiconductors, including cubic diamond. The small effective masses in h-diamond are comparable to those in the cubic phase, which cannot explain its substantially higher mobilities. Instead, two underlying mechanisms are uncovered. First, selection rules enforced by the symmetry of h-diamond significantly suppress scattering, particularly for transverse acoustic phonons, which predominate in the cubic phase around room temperature. Secondly, the spatial mismatch between the electronic wavefunctions and phonon-induced scattering potentials leads to real-space electron-phonon decoupling, which manifests as the suppression of out-of-plane polarised longitudinal acoustic scattering for holes, and a systematic weakening of acoustic scattering for electrons.

cond-mat.mtrl-sci

Rep2Text: Decoding Full Text from a Single LLM Token Representation

Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In this work, we investigate a fundamental question: to what extent can the original input text be recovered from a single last-token representation in an LLM? To this end, we propose Rep2Text, a novel framework for decoding text from last-token representations. Rep2Text employs a trainable adapter that maps a target model's last-token representation into the token embedding space of a decoding language model, which then autoregressively reconstructs the input text. Experiments across various model combinations (Llama-3.1-8B, Gemma-7B, Mistral-7B-v0.1, Llama-3.2-3B, etc.) show that, on average, roughly half of the tokens in 16-token sequences can be recovered from this compressed representation while preserving strong semantic coherence. Further analysis reveals a clear information bottleneck effect: as sequence length increases, token-level recovery declines, while semantic information remains relatively well preserved. We also find that scaling effects are less pronounced in inversion tasks. Finally, our framework demonstrates robust generalization to out-of-distribution clinical data.

cs.CL

First-principles Prediction of Carrier Mobility in Semiconductor Nanowires Based on the Spatially Dependent Boltzmann Transport Equation

Carrier mobility in bulk semiconductors is typically governed by electron-phonon (e-ph) scattering. In nanostructures, spatial confinement can lead to significant surface scattering, lowering mobility and breaking the spatial homogeneity assumption of conventional models. In this work, a fully ab initio framework based on the spatially dependent Boltzmann transport equation for one-dimensional nanowires is developed. We apply it to Si and GaN assuming diffusive surface scattering, and reveal the mobility-diameter relation: $\mu_\mathrm{1D} = \mu_\mathrm{bulk} \left[1-\left(d/d_0\right)^{-\beta}\right]$. The parameter $d_0$, comparable to the carrier mean free path, defines a boundary layer exhibiting a considerable mobility gradient, and also quantifies the competition between e-ph and surface scattering together with $\beta$. We further discuss the effects of orientation, cross-sectional shape, and temperature. Moreover, experimental data are generally lower than our predictions, possibly due to structural imperfections, systematic errors from measurements, etc. Therefore, our theoretical method can provide an intrinsic benchmark toward optimized experimental realizations.

cond-mat.mtrl-sci

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memorized solutions appear as genuine reasoning. Existing detection methods largely rely on surface overlap, completion behavior, or final-output likelihood, and often degrade when inputs are simply rephrased. In this paper, we propose LogitTrace(Layerwise Logit Trajectories), a framework for analyzing memorization-like decision dynamics through intermediate logit trajectories. Instead of judging memorization only from the final answer, LogitTrace examines how model preferences emerge and stabilize across layers. We find that contaminated examples tend to show earlier commitment, while clean examples exhibit more gradual evidence accumulation. These trajectory signals allow a lightweight classifier to separate contaminated and clean examples across multiple models and input variants. Controlled LoRA injection experiments further show that repeated exposure to target samples induces similar trajectory patterns. Overall, our results suggest that LogitTrace provides evidence beyond surface overlap and final-output confidence, offering a useful lens for studying memorization-like behavior in LLMs.

cs.CL

SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models

Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but controlling their behavior reliably remains challenging, especially in open-ended generation settings. This paper introduces a novel supervised steering approach that operates in sparse, interpretable representation spaces. We employ sparse autoencoders (SAEs) to obtain sparse latent representations that aim to disentangle semantic attributes from model activations. Then we train linear classifiers to identify a small subspace of task-relevant dimensions in latent representations. Finally, we learn supervised steering vectors constrained to this subspace, optimized to align with target behaviors. Experiments across sentiment, truthfulness, and political polarity steering tasks with multiple LLMs demonstrate that our supervised steering vectors achieve higher success rates with minimal degradation in generation quality compared to existing methods. Further analysis reveals that a notably small subspace is sufficient for effective steering, enabling more targeted and interpretable interventions. Our implementation is publicly available at https://github.com/Ineedanamehere/SAE-SSV.

cs.CL

SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection

Predicting earnings surprises from financial documents, such as earnings conference calls, regulatory filings, and financial news, has become increasingly important in financial economics. However, these financial documents present significant analytical challenges, typically containing over 5,000 words with substantial redundancy and industry-specific terminology that creates obstacles for language models. In this work, we propose the SAE-FiRE (Sparse Autoencoder for Financial Representation Enhancement) framework to address these limitations by extracting key information while eliminating redundancy. SAE-FiRE employs Sparse Autoencoders (SAEs) to decompose dense neural representations from large language models into interpretable sparse components, then applies statistical feature selection methods, including ANOVA F-tests and tree-based importance scoring, to identify the top-k most discriminative dimensions for classification. By systematically filtering out noise that might otherwise lead to overfitting, we enable more robust and generalizable predictions. Experimental results across three financial datasets demonstrate that SAE-FiRE significantly outperforms baseline approaches.

q-fin.CP

SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models

The ability of large language models (LLMs) to follow instructions is crucial for their practical applications, yet the underlying mechanisms remain poorly understood. This paper presents a novel framework that leverages sparse autoencoders (SAE) to interpret how instruction following works in these models. We demonstrate how the features we identify can effectively steer model outputs to align with given instructions. Through analysis of SAE latent activations, we identify specific latents responsible for instruction following behavior. Our findings reveal that instruction following capabilities are encoded by a distinct set of instruction-relevant SAE latents. These latents both show semantic proximity to relevant instructions and demonstrate causal effects on model behavior. Our research highlights several crucial factors for achieving effective steering performance: precise feature identification, the role of final layer, and optimal instruction positioning. Additionally, we demonstrate that our methodology scales effectively across SAEs and LLMs of varying sizes.

cs.LG

Mitigating Shortcuts in Language Models with Soft Label Encoding

Recent research has shown that large language models rely on spurious correlations in the data for natural language understanding (NLU) tasks. In this work, we aim to answer the following research question: Can we reduce spurious correlations by modifying the ground truth labels of the training data? Specifically, we propose a simple yet effective debiasing framework, named Soft Label Encoding (SoftLE). We first train a teacher model with hard labels to determine each sample's degree of relying on shortcuts. We then add one dummy class to encode the shortcut degree, which is used to smooth other dimensions in the ground truth label to generate soft labels. This new ground truth label is used to train a more robust student model. Extensive experiments on two NLU benchmark tasks demonstrate that SoftLE significantly improves out-of-distribution generalization while maintaining satisfactory in-distribution accuracy.

cs.CL