arXiv ScienceSearch

arXiv subjects

Keqing Zhang

Publications and source records attributed to Keqing Zhang.

2 recordsLinked to original sources

Do LLMs Have Values? A Quantitative Analysis and Alignment Framework for Values in Large Language Models

As Large Language Models (LLMs) increasingly handle complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. However, current efforts are confounded by a striking behavioral paradox: they fluctuate unpredictably under minor wording changes ("swing"), yet stubbornly ignore explicit instructions to correct ingrained biases ("rigidity"). Resolving this duality is critical for reliable AI alignment. To systematically understand and safely steer these latent subjective preferences, our study is structured around three fundamental questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150,000 queries per model) and 95,000 human survey profiles into a shared sociological space, we empirically confirm that they do. However, they do not mirror human diversity, instead crystallizing into a highly concentrated, idealized value core. Second, how can these values be quantified? We propose the Prior-Environment-Cognition (PEC) framework. This model mathematically defines value expression as the joint outcome of inherent dispositions like parameter weights (Prior), external contexts such as user prompts (Environment), and internal reasoning processes like Chain-of-Thought (Cognition). Finally, how can LLMs' values be aligned toward a desired target? Using PEC diagnostics, we establish an adaptive "Alignment Prescription". Rather than blindly applying resource-intensive training, this method identifies the minimum effective intervention needed for each dimension, ranging from zero-cost prompts to targeted parameter updates. Extensive empirical validation confirms that our approach successfully verifies the presence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.

cs.AI

Wall stabilization of the rigid ballooning $m=1$ mode in a long-thin mirror trap

The prospect of stabilization of the $m=1$ ``rigid'' ballooning mode in an open axially symmetric long-thin trap with the help of a conducting lateral wall surrounding a column of isotropic plasma is studied. It is found that for effective wall stabilization, the beta parameter must exceed $70\%$. The dependence of the critical beta on the mirror ratio, the radial pressure profile, and the axial profile of the vacuum magnet has been studied. It is shown that when a conductive lateral wall is combined with conductive end plates simulating attachment of the end MHD stabilizers to the central cell of an open trap, there are two critical beta values and two stability zones that can merge, making stable the entire range of allowable beta values $0<β<1$.

physics.plasm-ph