arXiv ScienceSearch

arXiv · 2607.23424

Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

Abstract

A red herring, an irrelevant passage added to a problem, corrupts a language model's reasoning and, through it, its final answer, while the form of the response survives untouched. The benchmark, called the Graduate Economic Reasoning Benchmark (GERB), is sixty graduate-level microeconomics problems, each a detailed setup with a verified final answer and a step-by-step reference solution. Each problem has two versions, one with the red herring and one without, and each of those is asked in two ways, one requesting an explanation and one not. This is a within-subject $2\times2$ factorial experimental design. Thirty-eight language models answer all four versions of every problem. The clean problems (the control group) are already hard, with the models answering under sixty percent correctly on average. The red herring lowers the probability of a correct final answer by 12.3 percentage points, about a quarter of the models' mean accuracy of 0.525. The damage is largest on the problems the model rates as easy. Reasoning ability confers no protection, as the red herring's effect does not differ detectably across models with and without reasoning ability. It does change how the failure looks, since a model with no reasoning mode repeats one wrong answer across waves while a reasoning model wavers. The red herring also leads a model to rate a problem as easier than its clean version, while answering it wrong more often. Although open- and closed-weight models reach the same accuracy, the open-weight models reach it at a substantially lower cost per correct final answer. The form of the response is preserved even as its substance fails. The model still produces an explanation (explanation given), the final answer still follows from the reasoning shown (coherence), and, in the aggregate, it remains the same across waves (consistency).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Piyush Akimitsu. 2026-07-31. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam. https://arxiv.org/abs/2607.23424

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment

Generative AI has become an everyday advisor, and the systems people consult are live and interactive, not pre-scripted. We ask whether access to such a system changes behavior under risk. In an incentivized experiment (N = 158), participants made lottery choices with an optional decision aid presented as a conventional pre-written tool, a live one-shot AI, or a live interactive AI they could query, with information format held equivalent across conditions. Risk preferences are elicited via DOSE. We find no evidence that access to a live AI advisor changes risk aversion.

econ.GN

Bricks or Cash? Externalities of Housing Upgrading in High-density Cities

We estimate housing externalities in a high-density city, exploiting the staggered rollout of Singapore's nationwide Main Upgrading Programme for public housing. Controlling for nonrandom neighborhood exposure, we find that upgrading raises treated buildings' prices by 11.5% upon completion and neighboring buildings' resale prices by about 2% within 500 meters, decaying to zero beyond. A model with distance-decaying externalities shows that in dense settings spillovers justify the distortions of in-kind provision; this advantage diminishes and reverses at lower densities. Administrative data on over 2 million residents show that upgrading disproportionately retains older incumbents, suggesting age-specific amenities as an underexplored externality channel.

econ.GN

The Joneses Visit an Economics Lab

Existing literature offers persuasive evidence that individuals care about how their consumption compares to that of peers, and proposes a large variety of explanatory models. The present paper proposes a common framework for many of those models, and compares their ability to predict behavior in a laboratory experiment. We find evidence of Keeping up with the Joneses motivations but also find that conspicuous consumption is enhanced by Veblen motivations arising from peers' ability to observe one's own choice. Among the seven quasi-linear preference models we compare, our data are best explained by a model that contrasts envy and pride (upward vs downward comparisons) using a value function borrowed from Prospect Theory.

econ.GN