arXiv · 2601.04288
Human-in-the-Loop Testing of AI Agents for Air Traffic Control with a Regulated Assessment Framework
Abstract
We present a rigorous, human-in-the-loop evaluation framework for assessing the performance of AI agents on the task of Air Traffic Control, grounded in a regulator-certified simulator-based curriculum used for training and testing real-world trainee controllers. By leveraging legally regulated assessments and involving expert human instructors in the evaluation process, our framework enables a more authentic and domain-accurate measurement of AI performance. This work addresses a critical gap in the existing literature: the frequent misalignment between academic representations of Air Traffic Control and the complexities of the actual operational environment. It also lays the foundations for effective future human-machine teaming paradigms by aligning machine performance with human assessment targets.
Explore related subjects
Keep this discovery
Ben Carvell, Marc Thomas, Andrew Pace, Christopher Dorney, George De Ath, Richard Everson, Nick Pepper, Adam Keane, Samuel Tomlinson, Richard Cannon. 2026-01-07. Human-in-the-Loop Testing of AI Agents for Air Traffic Control with a Regulated Assessment Framework. https://doi.org/10.2514/6.2026-2558
Cite the original work for its findings. Save a collection to share your selection of sources.