arXiv ScienceSearch

arXiv subjects

Brian Jabarian

Publications and source records attributed to Brian Jabarian.

9 recordsLinked to original sources

Et Tu, Brute? Economic Misalignment in Personal AI Agents

Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., their email inbox and a structured profile of personal attributes, with the intention of making an optimal, personalized decision for the user. We show that by simply providing this personal context, the agent steers recommendations based on inferred wealth, without being explicitly instructed to do so. In a suite of 325K experiments on 13 agents across three types of economic decisions (flights, health insurance, and graduate programs), we find that 8 models systematically choose more expensive options for wealthier users when requests are identical. This steering continues even when it directly goes against the user's stated objective: when explicitly instructed to find the cheapest option, some agents still act on the wealth profile they have inferred. It also occurs when wealth is inferred from ambient data, such as emails unrelated to the task. And it persists under privacy controls that block specific attributes: blocking financial attributes largely removes the disparity, but blocking other attributes leaves it unchanged and can increase it by up to 40% for insurance, as agents rely on the remaining signals to infer wealth. Larger and more capable models are no better; Claude Opus 4.8 shows the largest effect. We term this misalignment "adversarial delegation", in which the very conditions that make a personal AI agent useful - access to personal information - enable it to act against the user's interests.

cs.AI

The Economics of Recursive Self-Improvement

We model the economics of recursive self-improvement (RSI) and assess its plausibility and impacts. First, we build a sequence of increasingly rich models of AI progress to highlight the feedback loops behind RSI. We represent our models as directed graphs and show that net acceleration in AI capabilities depends on the product of elasticities across each feedback loop. Second, we distinguish between "narrow" and "broad" AI capabilities, capturing the possibility that AI systems improve narrowly at optimizing AI R&D benchmarks without improving at broader economically valuable tasks. Third, we document existing estimates of key parameters and provide a wish list of empirical objects that AI companies can measure and feasibly share publicly. Finally, we calibrate the model with existing data. A back-of-the-envelope calculation suggests that feedback loops are not currently strong enough to generate a self-sustaining acceleration, though they appear to be strengthening. We conclude by assessing the plausibility and implications of such an acceleration.

econ.GN

Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews

We study AI agents as information-collection technologies: automated systems that elicit decision-relevant signals from humans through live interactions. We test how such AI automation impacts information collection and organizational outcomes using a natural field experiment with 70,000 applicants applying for real jobs. Applicants were randomly assigned to be interviewed by either human recruiters or AI voice agents. Afterward, human recruiters evaluate the interviews and make hiring decisions. Applicants interviewed by AI agents are 12% more likely to receive job offers, and these gains translate into higher job starts and worker retention, with no decline in the productivity of hired workers. Analyzing interview transcripts reveals that AI voice agents achieve controlled variance: their interviews are more structured and consistent while remaining responsive to individual applicants, which is associated with more hiring-relevant information collected. Our results suggest that a key advantage of AI automation lies in environments where information collection is delegated across many human workers and repeated such that variance in task execution becomes noise in decision-relevant signals, which AI compresses through adaptive standardization.

econ.GN

GUIDE: Generative Utility Inference and Decision Engine

Measuring the preferences of human users remains a fundamental challenge of AI alignment. Existing elicitation approaches struggle to efficiently discover multidimensional preferences or accurately ground these inferences in domain knowledge. To address this, we introduce GUIDE, an LLM-driven elicitation architecture that infers user preferences through conversations by combining Bayesian adaptive sampling for question selection and symbolic representation learning to initialize domain-specific preference models. GUIDE generalizes adaptive sampling to diverse elicitation questions through an extensible type system of transforms on a parameterized preference state. GUIDE produces domain-specific preference representations through an initialization process using symbolic rule-based learning to capture world knowledge and set priors over preference dimensions grounded in data about decision alternatives. The architecture provides observability and steerability to facilitate deployment and analyze elicitation processes. In silico experiments on investment portfolio optimization demonstrate that GUIDE improves cold-start and minimizes recommendation regret consistently within early elicitation interactions across user personas compared to prior work, LLM-only baselines, and ablated GUIDE versions.

cs.LG

AI Behavioral Science

We outline a foundation for a new field of ``AI Behavioral Science,'' covering three perspectives. First, as AI becomes ubiquitous and is increasingly proprietary and opaque, it becomes vital to develop techniques for assessing AI behavior. We outline how tools developed to assess people's behaviors by social scientists can be used to assess and infer AI's behaviors biases, tendencies, and heuristics. Second, we also discuss how AI can change the ways in which we learn about human behavior. Beyond its computational power, AI offers new techniques for simulating, inferring, and predicting human behaviors that we outline and discuss. Third, as humans and AI are interacting in increasingly complex and intertwined systems, we need to understand the implications for the resulting economic and political outcomes. We outline issues that are increasingly pressing concerning the future of human-AI interactions and potential changes and disruptions that can ensue.

cs.HC

Large Language Models for Behavioral Economics: Internal Validity and Elicitation of Mental Models

In this article, we explore the transformative potential of integrating generative AI, particularly Large Language Models (LLMs), into behavioral and experimental economics to enhance internal validity. By leveraging AI tools, researchers can improve adherence to key exclusion restrictions and in particular ensure the internal validity measures of mental models, which often require human intervention in the incentive mechanism. We present a case study demonstrating how LLMs can enhance experimental design, participant engagement, and the validity of measuring mental models.

cs.HC

Critical Thinking Via Storytelling: Theory and Social Media Experiment

In a stylized voting model, we establish that increasing the share of critical thinkers -- individuals who are aware of the ambivalent nature of a certain issue -- in the population increases the efficiency of surveys (elections) but might increase surveys' bias. In an incentivized online social media experiment on a representative US population (N = 706), we show that different digital storytelling formats -- different designs to present the same set of facts -- affect the intensity at which individuals become critical thinkers. Intermediate-length designs (Facebook posts) are most effective at triggering individuals into critical thinking. Individuals with a high need for cognition mostly drive the differential effects of the treatments.

econ.TH

A Two-Ball Ellsberg Paradox: An Experiment

We conduct an incentivized experiment on a nationally representative US sample \\ (N=708) to test whether people prefer to avoid ambiguity even when it means choosing dominated options. In contrast to the literature, we find that 55\% of subjects prefer a risky act to an ambiguous act that always provides a larger probability of winning. Our experimental design shows that such a preference is not mainly due to a lack of understanding. We conclude that subjects avoid ambiguity \textit{per se} rather than avoiding ambiguity because it may yield a worse outcome. Such behavior cannot be reconciled with existing models of ambiguity aversion in a straightforward manner.

econ.TH

The Moral Burden of Ambiguity Aversion

In their article, "Egalitarianism under Severe Uncertainty", Philosophy and Public Affairs, 46:3, 2018, Thomas Rowe and Alex Voorhoeve develop an original moral decision theory for cases under uncertainty, called "pluralist egalitarianism under uncertainty". In this paper, I firstly sketch their views and arguments. I then elaborate on their moral decision theory by discussing how it applies to choice scenarios in health ethics. Finally, I suggest a new two-stage Ellsberg thought experiment challenging the core of the principle of their theory. In such an experiment pluralist egalitarianism seems to suggest the wrong, morally and rationally speaking, course of action -- no matter whether I consider my thought experiment in a simultaneous or a sequential setting.

econ.TH