arXiv · 2609.27849
Same Team Label, Different Evidence: A Full-Text Audit of Claim Denominators in Human-AI Teaming Research
Abstract
Human-AI Teaming (HAT) reviews often group studies by labels such as advisor, teammate, or coordinator. Yet the same label can describe one person taking AI advice, several people coordinating around AI, or a workflow that distributes authority and responsibility. Pooling these studies can therefore change the human unit behind a claim. We examine how full-text evidence changes the set of studies behind a claim. We audited 86 full texts purposively selected from a 419-record title/abstract map. We find that full-text reading changed core membership for 40 records: 36 of 74 apparent core candidates moved out, while 4 of 12 boundary candidates moved in. Team vocabulary did not reliably identify the social unit: 14 of 27 human-AI dyads and 20 of 23 multi-human peer teams used team or collaboration terms. Only 20 of 86 papers specified who could see AI output. Four blinded language-model runs unanimously labeled 53 screening cases and 59 arrangements, yet 32% and 34% of those consensus decisions differed from the full-text labels. These results identify claim-denominator drift as a synthesis problem in HAT research. We contribute a full-text audit centered on human arrangements and a claim-pooling checkpoint for deciding when evidence about trust, coordination, performance, efficiency, and accountability can be compared.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hanjing Shi, Kimberly Wang, Sabrina Doherty, Dominic DiFranzo. 2026-08-25. Same Team Label, Different Evidence: A Full-Text Audit of Claim Denominators in Human-AI Teaming Research. https://arxiv.org/abs/2609.27849
Cite the original work for its findings. Save a collection to share your selection of sources.