arXiv Science⌕ Search

arXiv · 2609.35448

AI-based matching improves refugee employment in a double-blind randomized trial

Abstract

Refugee integration is a central policy challenge for host countries, and where governments initially place refugees shapes their integration trajectories. Yet placement officers often have limited information about where each case is most likely to succeed. Algorithmic refugee matching uses administrative data, machine learning, and constrained optimization to recommend employment-optimized placements in real time as cases arrive, with human placement officers retaining final authority. Between January 2020 and June 2023, the Swiss State Secretariat for Migration randomly assigned about 2,000 refugee cases to receive a canton recommendation either algorithmically optimized for employment or drawn to approximate existing procedures, with placement officers and refugees blinded to assignment. The two arms used identical but separate canton and origin-group quotas, so gains reflect better refugee-canton matching rather than reallocation toward stronger labor markets. The trial began just before the COVID-19 pandemic shifted labor-market conditions. For the pre-registered primary outcome -- the share of months employed during the first three years -- the pooled intention-to-treat (ITT) estimate across the 2020-2023 placement cohorts was +2.2 percentage points (about 10% of the 22.3% control mean; 95% CI [+0.05, +4.33]), rising to +3.9 pp (about 17%; [+1.11, +6.68]) for the post-COVID 2022-2023 cohorts. Effects grew over time: at 36 months, the pooled ITT on the employment rate was +5.2 pp (about 11%; 95% CI [+1.10, +9.25]) -- comparable to the gains from hundreds of hours of intensive language training. Overall, the results provide rare field evidence that AI-based decision support can improve high-stakes public-sector allocation, offering a scalable, low-cost way to raise refugee employment.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kirk Bansak, Jens Hainmueller, Dominik Hangartner, Jeremy Ferwerda, Elisabeth Paulson, Angie Delevoye, Nicholas Adams-Cohen, Ashwin Ramaswami, Selina Kurer, Joelle Pianzola, Michael Hotard. 2026-09-28. AI-based matching improves refugee employment in a double-blind randomized trial. https://arxiv.org/abs/2609.35448

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Should I State or Should I Show? Aligning AI with Human Preferences

The proliferation of AI agents introduces a new principal-agent problem which stems from human principals' difficulty in articulating preferences. We report results from an experiment focused on mitigating this problem via revealed preferences from choice data. Compared to stated preference from human-written prompts, individuals communicate their preferences more effectively through choices, with as few as two sufficing to match prompts' predictive value. This gap largely reflects subjects' difficulty in translating preferences into prompts and is largest among those exhibiting more Allais-type behavioral patterns. Subjects also misperceive the approaches' relative performance, often choosing the less accurate one.

econ.GN↗

The missing price feedback: Why studies overstate local peaks from synchronized home batteries under dynamic pricing

Existing studies consistently find that home batteries and electric vehicles optimized against real-time electricity prices create new load peaks in local distribution grids, because all assets respond to the same price signal in sync. However, all but one of the 18 studies reviewed here treat wholesale prices as exogenous. This paper argues that treating prices as non-responsive biases the result: in reality, charging in low-price hours raises the wholesale price, which dampens the incentive to charge. I test this by simulating households with rooftop solar and home batteries under a spot-based retail tariff, calibrated to German data for 2025 with a model of equilibrium wholesale prices. With exogenous prices, my results confirm the finding from the previous literature: synchronous charging pushes the coincident import peak of the local grid to 83% above its no-battery level. With endogenous prices, the same fleet leaves the peak 6% below the no-battery benchmark. In other words, treating prices as endogenous reduces the coincident peak at high battery penetration nearly by half. This finding is robust across alternative price functions and exogenous price paths, different measures of the coincident peak, 25 scenario variations and three historical years. I conclude that new local peaks remain possible, particularly in grids where battery deployment runs ahead of the national average, but the risk and magnitude are considerably smaller than the existing literature suggests.

econ.GN↗

Representation Risk in Pretrained Image Encoders

Applied researchers increasingly convert images into features with pretrained encoders, then use those features in a downstream prediction model. The encoder is often treated as an implementation detail. We show that it can instead be a consequential source of model uncertainty. We call this uncertainty representation risk: plausible pretrained encoders map the same images into different feature spaces and can yield sharply different out-of-sample conclusions from predictive performance. We compare ten modern and legacy frozen encoders across applications involving house prices, racehorse performance, breast-cancer histology, chest radiographs, continuous facial age, and rice disease. With common dimension control, heads, and group-safe splits, validation selects SigLIP 2 for houses, raising test $R^2$ from 0.396 for ResNet50 to 0.629, and DINOv2 for horses, raising $R^2$ from 0.029 to 0.105. No encoder is best in every task. Candidate procedures are constructed using training data and compared on a separate validation partition. The selected procedure reaches 0.658 for houses and 0.979 accuracy for pneumonia. Fixed-split gains are small for horses and rice, while repeated partitions reveal instability in horse feature union. Continuous age selects SigLIP 2 at 4.786 years MAE. The principal representation gaps persist with neural heads, similarly sized DINOv2 and ViT models, and limited adaptation. These results support a simple workflow: benchmark plausible representations, select on locked validation data, combine only when separate validation evidence justifies the additional cost, and report paired and split-level uncertainty. We implement this workflow in LOOKAGAIN-ML, the software package used to conduct the analyses in this paper.

econ.GN↗