arXiv ScienceSearch

arXiv subjects

Haijiang Liu

Publications and source records attributed to Haijiang Liu.

8 recordsLinked to original sources

StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions

Large language models often answer the same multiple-choice question inconsistently when it is posed under support-oriented and elimination-oriented framings. We investigate whether these discrepancies arise from different internal representations induced by the two framings. We introduce a dual-framing protocol with minimally varied prompts that use either support- or elimination-oriented framing while keeping the evaluation target fixed. To probe the internal computation, we append an untrained special token, [STATE], and treat its residual-stream activation as an intervention interface. Across both models, the two framings induce separable [STATE] activations concentrated in intermediate layers. Swapping these activations between paired prompts systematically changes predictions and improves cross-framing agreement, providing intervention-based evidence that the activations are behaviorally relevant. Beyond instance-level substitution, mean-difference steering directions derived from the dual-framing contrast exhibit more bounded layer-wise responses than matched contrastive activation addition directions under the evaluated protocol.

cs.CL

Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis

As Large Language Models (LLMs) increasingly influence high-stakes decision-making across global contexts, ensuring their alignment with diverse cultural values has become a critical governance challenge. This study presents a Multi-Layered Auditing Platform for Responsible AI that systematically evaluates cross-cultural value alignment in China-origin and Western-origin LLMs through four integrated methodologies: Ethical Dilemma Corpus for assessing temporal stability, Diversity-Enhanced Framework (DEF) for quantifying cultural fidelity, First-Token Probability Alignment for distributional accuracy, and Multi-stAge Reasoning frameworK (MARK) for interpretable decision-making. Our comparative analysis of 20+ leading models, such as Qwen, GPT-4o, Claude, LLaMA, and DeepSeek, reveals universal challenges-fundamental instability in value systems, systematic under-representation of younger demographics, and non-linear relationships between model scale and alignment quality-alongside divergent regional development trajectories. While China-origin models increasingly emphasize multilingual data integration for context-specific optimization, Western models demonstrate greater architectural experimentation but persistent U.S.-centric biases. Neither paradigm achieves robust cross-cultural generalization. We establish that Mistral-series architectures significantly outperform LLaMA3-series in cross-cultural alignment, and that Full-Parameter Fine-Tuning on diverse datasets surpasses Reinforcement Learning from Human Feedback in preserving cultural variation...

cs.CY

Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive Reasoning

Introducing MARK, the Multi-stAge Reasoning frameworK for cultural value survey response simulation, designed to enhance the accuracy, steerability, and interpretability of large language models in this task. The system is inspired by the type dynamics theory in the MBTI psychological framework for personality research. It effectively predicts and utilizes human demographic information for simulation: life-situational stress analysis, group-level personality prediction, and self-weighted cognitive imitation. Experiments on the World Values Survey show that MARK outperforms existing baselines by 10% accuracy and reduces the divergence between model predictions and human preferences. This highlights the potential of our framework to improve zero-shot personalization and help social scientists interpret model predictions.

cs.CL

Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond

The large-scale development of large language models (LLMs) in medical contexts, such as diagnostic assistance and treatment recommendations, necessitates that these models possess accurate medical knowledge and deliver traceable decision-making processes. Clinical knowledge, encompassing the insights gained from research on the causes, prognosis, diagnosis, and treatment of diseases, has been extensively examined within real-world medical practices. Recently, there has been a notable increase in research efforts aimed at integrating this type of knowledge into LLMs, encompassing not only traditional text and multimodal data integration but also technologies such as knowledge graphs (KGs) and retrieval-augmented generation (RAG). In this paper, we review the various initiatives to embed clinical knowledge into training-based, KG-supported, and RAG-assisted LLMs. We begin by gathering reliable knowledge sources from the medical domain, including databases and datasets. Next, we evaluate implementations for integrating clinical knowledge through specialized datasets and collaborations with external knowledge sources such as KGs and relevant documentation. Furthermore, we discuss the applications of the developed medical LLMs in the industrial sector to assess the disparity between models developed in academic settings and those in industry. We conclude the survey by presenting evaluation systems applicable to relevant tasks and identifying potential challenges facing this field. In this review, we do not aim for completeness, since any ostensibly complete review would soon be outdated. Our goal is to illustrate diversity by selecting representative and accessible items from current research and industry practices, reflecting real-world situations rather than claiming completeness. Thus, we emphasize showcasing diverse approaches.

cs.AI

Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations

Large-scale surveys are essential tools for informing social science research and policy, but running surveys is costly and time-intensive. If we could accurately simulate group-level survey results, this would therefore be very valuable to social science research. Prior work has explored the use of large language models (LLMs) for simulating human behaviors, mostly through prompting. In this paper, we are the first to specialize LLMs for the task of simulating survey response distributions. As a testbed, we use country-level results from two global cultural surveys. We devise a fine-tuning method based on first-token probabilities to minimize divergence between predicted and actual response distributions for a given question. Then, we show that this method substantially outperforms other methods and zero-shot classifiers, even on unseen questions, countries, and a completely unseen survey. While even our best models struggle with the task, especially on unseen questions, our results demonstrate the benefits of specialization for simulation, which may accelerate progress towards sufficiently accurate simulation in the future.

cs.CL

Observation of flat bands due to band hybridization in 3d-electron heavy-fermion compound CaCu3Ru4O12

We report angle-resolved photoemission spectroscopy and first-principles numerical calculations for the band structure evolution of the 3d heavy-fermion compound CaCu3Ru4O12. Below 200 K, we observed an emergent hybridization gap between the Cu 3d electron-like band and the Ru 4d hole-like band and the resulting flat band features near the Fermi energy centered around the Brillouin zone corner. Our results confirm the non-Kondo nature of CaCu3Ru4O12, in which the Cu 3dxy electrons are less correlated and not in the Kondo limit. Comparison between theory and experiment also suggests that other mechanism such as nonlocal interactions or spin fluctuations beyond the local dynamical mean-field theory may be needed in order to give a quantitative explanation of the peculiar properties in this material.

cond-mat.str-el

Large Fermi Surface Expansion through Anisotropic c-f Mixing in the Semimetallic Kondo Lattice System CeBi

Using angle-resolved photoemission spectroscopy (ARPES) and resonant ARPES, we report evidence of strong anisotropic conduction-f electron mixing (c-f mixing) in CeBi by observing a largely expanded Ce-5d pocket at low temperature, with no change in the Bi-6p bands. The Fermi surface (FS) expansion is accompanied by a pronounced spectral weight transfer from the local 4f 0 peak of Ce (corresponding to Ce3+) to the itinerant conduction bands near the Fermi level. Careful analysis suggests that the observed large FS change (with a volume expansion of the electron pocket up to 40%) can most naturally be explained by a small valence change (~ 1%) of Ce, which coexists with a very weak Kondo screening. Our work therefore provides evidence for a FS change driven by real charge fluctuations deep in the Kondo limit, which is made possible by the low carrier density.

cond-mat.str-el

Anomalous doping evolution of nodal dispersion revealed by in-situ ARPES on continuously doped cuprates

We study the systematic doping evolution of nodal dispersions by in-situ angle-resolved photoemission spectroscopy on the continuously doped surface of a high-temperature superconductor Bi$_2$Sr$_2$CaCu$_2$O$_{8+x}$. We reveal that the nodal dispersion has three segments separated by two kinks, located at ~10 meV and roughly 70 meV, respectively. The three segments have different band velocities and different doping dependence. In particular, the velocity of the high-energy segment increases monotonically as the doping level decreases and can even surpass the bare band velocity. We propose that electron fractionalization is a possible cause for this anomalous nodal dispersion and may even play a key role in the understanding of exotic properties of cuprates.

cond-mat.supr-con