arXiv Science⌕ Search

arXiv · 2610.10064

Comprehension Audits to Mitigate Risks from Automated AI Research

Abstract

AI is already writing a majority of code for frontier AI labs. This creates a safety risk if there is insufficient human oversight. Existing work proposes minimum comprehension thresholds and unaided checks to mitigate this. To our knowledge, however, there is currently no published frontier-AI assurance regime that requires demonstrated evidence that the responsible humans understand what they are building as a precommitted condition for continuing development or usage. We propose comprehension audits, a novel development-process assurance mechanism in which the responsible people explain R&D contributions to auditors to demonstrate understanding. With independent administration and graded reports, they provide a gate: development of a contribution stops based on a failure to demonstrate human understanding until remediated, with escalating consequences for repeated failures. Our analysis of leading open-source AI projects finds increased output of code with reduced human review commentary rates per line of code, with far lower rates for automated fleet accounts. We advocate for labs to conduct them with embedded independent auditors.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ronald J. Bodkin, Bahrad A. Sokhansanj, Gillian K. Hadfield. 2026-10-07. Comprehension Audits to Mitigate Risks from Automated AI Research. https://arxiv.org/abs/2610.10064

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

What Personal Information Improves LLM-Based Next-Location Prediction?

Large language models (LLMs) are increasingly used for individual next-location prediction, with personal information easily added to prompts alongside mobility history. Yet the incremental predictive value of such information remains unclear. Using linked sociodemographic records and mobility traces from 5,000 Shenzhen residents, this study separates model responsiveness from predictive value. GPT-5 is the primary model, with GPT-5.5 and Claude Opus 4.6 used for replication. In 1,000 paired prediction instances, models rank 100 candidate destinations with and without age, gender, occupation and income while all other inputs are held fixed. Behavioural history raises top-1 accuracy from 5.6% to 18.5% as prior history increases from zero to six days. By contrast, sociodemographic attributes produce no detectable overall gain, although replacing correct attributes with those of another person reduces accuracy by 5.4 percentage points. Candidate construction also matters, removing distance raises accuracy by 7.7 points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These findings identify behavioural history as the clearest source of incremental value and show that personal information should be evaluated under matched, explicitly specified conditions before its privacy and governance costs are justified.

cs.CY↗

Bridging MOOCs, Smart Teaching, and AI-Assisted Learning: A Unified Quantitative Model

MOOCs, Smart Teaching (ST), and AI-assisted learning support different stages of higher education, yet a unified quantitative basis for analyzing their contributions, course design, and resource allocation remains limited. This study develops a mathematical model integrating MOOC-based pre-class learning, ST-based in-class adaptation, and AI-assisted post-class personalization. Learner mastery is represented as a bounded multidimensional state, with a common exponential learning-response function describing how instructional resources reduce remaining knowledge gaps. The stages differ in their allocation rules: predefined course-content emphasis for MOOCs, feedback-driven class-level adaptation for ST, and individualized gap-based allocation for AI-assisted learning. Simulations across 100 independently generated classes demonstrate stable cumulative progression and quantify stage-wise gains, while budget analysis reveals diminishing returns from additional AI support. For a fixed learner and learning mechanism, varying course-content emphasis shows a strong association between course-learner alignment and post-MOOC mastery. An optimal AI allocation is also derived under a fixed budget: supported components reach a common residual mastery gap, while components below this threshold receive no resources. A controlled comparison yields over 14% greater learning gain than proportional allocation. These analyses make the instructional process quantitatively analyzable and provide a basis for examining course-learner fit and coordinating limited learning resources. The model offers an analytical foundation for instructional decisions, with practical application requiring empirical estimation of learner states and calibration of learning-response parameters.

cs.CY↗

Reasoning Enhances Robustness to Prompt Injection in LLM-Based Consensus

Large Language Models (LLMs) are gaining traction as a method to generate consensus statements and aggregate preferences in digital democracy experiments. Yet, participants can introduce critical vulnerabilities in LLM-based systems. Here, we examine the vulnerability and robustness of off-the-shelf consensus-generating LLMs to prompt-injection attacks, which consist of introducing strategically designed texts to amplify particular viewpoints, erase certain opinions, or divert consensus toward unrelated or irrelevant topics. Using data collected from a 2023 experiment conducted in the UK and predefining a majority-rule, we construct attack-free and adversarial variants of prompts containing public policy questions and opinion texts, classify opinion and consensus valences with a fine-tuned BERT model, and estimate LLM--human majority agreement rates. Overall, we find that default LLMs exhibit widespread vulnerability, especially when: (i) disagreement and agreement are finely balanced, (ii) under rational, instruction-like rhetorical strategies, and (iii) for attacks that shift consensus toward positions aligned with GB-unionist conservative manifestos relative to pro-independence left manifestos. A robustness pipeline combining GPT-OSS-SafeGuard injection detection, structured opinion representations, and GSPO-based reinforcement learning substantially reduces directional failures, outperforming state-of-the-art alternatives. While more effective detectors may enable more sophisticated attacks, these findings advance our understanding of both the vulnerabilities and the potential defenses of consensus-generating LLMs in digital democracy applications.

cs.CY↗