arXiv · 2609.09070
Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics
Abstract
Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.
Explore related subjects
Keep this discovery
Andy Nkansah, Hanna Plotnitskaya, Stanislau Salavei, Anna Kozlova, Piotr Gibas, Julian Milek, Viktar Harbachou, Aleksey Ropan, Pavel Satalkin. 2026-09-08. Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics. https://arxiv.org/abs/2609.09070
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.