arXiv ScienceSearch

arXiv subjects

Jason Hung

Publications and source records attributed to Jason Hung.

6 recordsLinked to original sources

How much of a measured AI preference is the model, and how much is the instrument?

Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.

cs.AI

Benchmarking Open-Weight Foundation Models for Global AI Technical Governance

Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and international organisations. There is, however, growing evidence that such models produce significantly less accurate responses for countries that are underrepresented in their training data-a pattern described in existing literature as geographic bias. Existing studies examining this phenomenon are subject to three methodological limitations that together undermine their findings: (1) reliance on proprietary systems whose weights are not publicly released, which prevents independent replication; (2) evaluation of model knowledge about years that fall after data collection for model training had concluded, leading to geographic ignorance in addition to the natural limits of each model's knowledge; and (3) use of coarse binary response classification that cannot distinguish models' confident fabrication (HF) from their honest acknowledgement of uncertainty. This study addresses all three limitations by benchmarking four open-weight frontier language models against the Global AI Dataset v2 (GAID v2), a verified ground-truth database of 24,453 indicators across 227 countries published on Harvard Dataverse in January 2026. A total of 18 indicators, mapped to the eight thematic dimensions of the IEEE IRAI 2026 framework, are selected from GAID v2, yielding approximately 2,990 country-metric-year observations across six evaluation years (within the period of 2010-2023). Model responses are classified using a five-category scheme that distinguishes (a) verified accuracy (VA), (b) HF, (c) honest refusal (HR), (d) qualitative hedging (QH), and (e) misattribution (MF). Geographic disparities in accuracy are estimated through mixed-effects logistic regression and difference-in-differences (DiD) analysis.

cs.CY

Animal Welfare and Policy Risk Index (AWPRI): Constructing a Cross-National Exposure Measure of Animal Use and Governance Capacity, 24 Countries, 2010-2022

This research paper introduces the Animal Welfare and Policy Risk Index (AWPRI), a composite exposure index covering 24 countries over the period 2010-2022. The AWPRI is constructed from 12 variables organised across three conceptual layers, namely Animal-use exposure (L1), Governance capacity (L2), and AI amplification (L3). Each layer is the mean of its own variables and the composite is the mean of the three layers, so a variable sitting in the three-variable layer carries more weight than one sitting in the five-variable layer. Variables are min-max normalised within each year, so a country's value on any variable, layer, or composite suggests its position among the 24 countries that year. K-means clustering of the layer scores selects two groups (silhouette coefficient = 0.504 at k = 2, against 0.444 at k = 3 and 0.330 at k = 4). Principal component analysis (PCA) of the 12-variable 2022 cross-section requires five components to reach 90.8% of cumulative variance, showing that the index is not reducible to a single dimension. The construction passes 46 robustness checks, with the rank correlation against the index as built staying above 0.98 under z-score normalisation, pooled min-max bounds, layer reweighting, and leave-one-variable-out. China (0.703), Vietnam (0.663), and Thailand (0.612) record the highest composite scores in 2022; Sweden (0.316), Kenya (0.339), and Germany (0.362) the lowest. Tested against the World Animal Protection (WAP)'s Animal Protection Index, the composite returns a null (Spearman \r{ho} = -0.264, p = 0.213, n = 24) while L1 on its own tracks the WAP grade (\r{ho} = -0.485, p = 0.016). The AWPRI and its interactive visualisation are publicly accessible at https://awpri.aiinsocietyhub.com/.

econ.EM

Geographic Blind Spots in AI Control Monitors: A Cross-National Audit of Claude Opus 4.6

Artificial intelligence (AI) control protocols assume that trusted large language model (LLM) monitors reliably assess proposed actions across all deployment contexts. This paper tests that assumption in the geographic dimension. We audit Claude Opus 4.6-the monitor specified in Apart Research's AI Control Hackathon Track 3 benchmark-for systematic gaps in its factual knowledge of the global AI landscape. We develop the AI Control Knowledge Framework (ACKF), a six-dimension thematic scheme, and operationalise it with 17 verified indicators drawn from the Global AI Dataset v2 (GAID v2): 24,453 indicators across 227 countries published on Harvard Dataverse. A five-category response classification scheme distinguishes verifiable fabrication (VF) from honest refusal (HR); logistic regression with country-clustered standard errors combined with difference-in-differences (DiD) estimation quantifies geographic disparities in monitor accuracy across 2,820 country-metric-year observations. Contrary to our initial hypothesis, Claude Opus 4.6 produces higher fabrication rates for Global North queries than for Global South counterparts-a pattern consistent with a partial-knowledge mechanism in which the model attempts answers more frequently for Global North contexts but commits to incorrect values. This fabrication profile constitutes an exploitable vulnerability, where an adversarial AI system could frame harmful actions in governance or public attitude terms to reduce the probability of detection. This study provides the first cross-national, multi-domain audit of an AI control monitor's geographic knowledge gaps, with direct implications for the design of control protocols.

cs.CY

Global AI Bias Audit for Technical Governance

This paper presents the outputs of the exploratory phase of a global audit of Large Language Models (LLMs) project. In this exploratory phase, I used the Global AI Dataset (GAID) Project as a framework to stress-test the Llama-3 8B model and evaluate geographic and socioeconomic biases in technical AI governance awareness. By stress-testing the model with 1,704 queries across 213 countries and eight technical metrics, I identified a significant digital barrier and gap separating the Global North and South. The results indicate that the model was only able to provide number/fact responses in 11.4% of its query answers, where the empirical validity of such responses was yet to be verified. The findings reveal that AI's technical knowledge is heavily concentrated in higher-income regions, while lower-income countries from the Global South are subject to disproportionate systemic information gaps. This disparity between the Global North and South poses concerning risks for global AI safety and inclusive governance, as policymakers in underserved regions may lack reliable data-driven insights or be misled by hallucinated facts. This paper concludes that current AI alignment and training processes reinforce existing geoeconomic and geopolitical asymmetries, and urges the need for more inclusive data representation to ensure AI serves as a truly global resource.

cs.CY

Trajectories and Comparative Analysis of Global Countries Dominating AI Publications, 2000-2025

This study investigates the shifting global dynamics of Artificial Intelligence (AI) research by analysing the trajectories of countries dominating AI publications between 2000 and 2025. Drawing on the comprehensive OpenAlex datasets and employing fractional counting to avoid double attribution in co-authored work, the research maps the relative shares of AI publications across major global players. The analysis reveals a profound restructuring of the international AI research landscape. The US and the European Union (representing EU27), once the undisputed and established leaders, have experienced a notable decline in relative dominance, with their combined share of publications falling from over 57% in 2000 to less than 25% in 2025. In contrast, China has undergone a dramatic ascent, expanding its global share of AI publications from under 5% in 2000 to nearly 36% by 2025, therefore emerging as the single most dominant contributor. Alongside China, India has also risen substantially, consolidating a multipolar Asian research ecosystem. These empirical findings highlight the strategic implications of concentrated research output, particularly China's capacity to shape the future direction of AI innovation and standard-setting. Beyond publication volume, the study further examines research quality by comparing each country's share of high-impact publications against its overall output, and analyses citation impact trajectories across major players. The findings show that in addition to China leading in volume, the country has also recently led in high-impact publications. Such an observation challenges the general assumption that Western powers retain dominance in high-impact AI scholarship.

cs.CY