arXiv ScienceSearch

arXiv · 2505.02313

What Is AI Safety? What Do We Want It to Be?

Abstract

The field of AI safety seeks to prevent or reduce the harms caused by AI systems. A simple and appealing account of what is distinctive of AI safety as a field holds that this feature is constitutive: a research project falls within the purview of AI safety just in case it aims to prevent or reduce the harms caused by AI systems. Call this appealingly simple account The Safety Conception of AI safety. Despite its simplicity and appeal, we argue that The Safety Conception is in tension with at least two trends in the ways AI safety researchers and organizations think and talk about AI safety: first, a tendency to characterize the goal of AI safety research in terms of catastrophic risks from future systems; second, the increasingly popular idea that AI safety can be thought of as a branch of safety engineering. Adopting the methodology of conceptual engineering, we argue that these trends are unfortunate: when we consider what concept of AI safety it would be best to have, there are compelling reasons to think that The Safety Conception is the answer. Descriptively, The Safety Conception allows us to see how work on topics that have historically been treated as central to the field of AI safety is continuous with work on topics that have historically been treated as more marginal, like bias, misinformation, and privacy. Normatively, taking The Safety Conception seriously means approaching all efforts to prevent or mitigate harms from AI systems based on their merits rather than drawing arbitrary distinctions between them.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jacqueline Harding, Cameron Domenico Kirk-Giannini. 2025-05-05. What Is AI Safety? What Do We Want It to Be?. https://arxiv.org/abs/2505.02313

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Algorithmic Shortlisting in Participatory Budgeting

Participatory budgeting is a democratic innovation that allows citizens to propose and vote on public investment projects. To help organizers manage large volumes of submissions, we design and test privacy-preserving methods for algorithmic shortlisting. These algorithms predict which projects are likely to be funded using only project features and anonymous historical voting data. We demonstrate the limitations of a naive approach that uses a large language model to rank projects based on past success and propose a vote-based pipeline that enables state-of-the-art LLMs to perform on par with classical machine learning. Our findings indicate that user preferences in participatory budgeting are stable enough to allow algorithmic shortlisting to approximate an initial selection of projects effectively.

cs.CY

Human Resilience in the AI Era -- What Machines Can't Replace

AI is changing work and decision making faster than many institutions can adapt their operating practices. We argue that this adaptation gap makes human resilience a core capability for the AI era. We define resilience as the capacity to absorb disruption while preserving effective action and human agency around core purposes. The framework operates at three interacting levels. Psychological resilience keeps a person goal-directed under stress. Social resilience makes trusted support and correction available across a group. Organizational resilience turns detected problems into learning and recovery. We connect established resilience and technostress research with direct AI-in-the-loop experiments. General resilience is trainable, while AI-specific causal evidence is still emerging. Direct AI studies show that assistance can raise productivity and spread expertise. Other experiments show improved expressed empathy and more calibrated reliance. We translate these findings into a practical agenda for AI education, workplace design, governance, and evaluation. The central proposal is socio-technical: structural safeguards define the operating boundary, while resilient people and institutions provide adaptive capacity when conditions change.

cs.CY

Toward a Time-Aware Assessment Framework for the Carbon Cost of AI-Enabled Decarbonization

AI is increasingly used to support decarbonization decisions across the built environment, yet the development, training, and use of AI consume energy and induce CO2e emissions. However, existing assessments often report physical-system savings while omitting AI-side emissions. Moreover, they rarely account for the mismatch between when AI costs occur and when decarbonization benefits materialize, which may be substantial for infrastructure-scale projects. To address these issues, we present a time-aware assessment framework that models avoided emissions and AI-induced emissions as discrete-time streams over a finite time horizon. In demonstrating this process, we seek to show that time-aware assessment can support temporal decision-making, identify cases in which accounting for time value of carbon can change preferred rankings relative to time-invariant totals, and explore how decisions may vary with slightly different governance priorities. Using four representative interventions with intentionally different temporal profiles (multi-project low-carbon concrete design support, AI-assisted construction logistics, agentic HVAC control, and predictive maintenance), we demonstrate how discounting can change preferred rankings relative to time-invariant totals and supports ranking sensitivity analysis, discounted payback screening, and break-even discount-rate analysis. We also provide decision guidelines that support go/no-go screening, timing decisions, and minimum "bang-for-your-buck" thresholds. Ultimately, this work contributes a lightweight framework for deciding whether and when to deploy AI-enabled interventions for decarbonization under explicit time preference.

cs.CY