arXiv · 2610.08161
Symphony for Text Generation: Benchmarking Clinical Note Generation
Abstract
Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Daniel Varab, Victor Petrén Bach Hansen, Asbjørn W. Helge, Kevin Pelgrims, Mathias Baltzersen, Adrian Young-San Roessler, Vanessa Klungtvedt, Maximilian Brand, Lasse Krogsbøll, Henrik Cullen, Lars Maaløe. 2026-10-06. Symphony for Text Generation: Benchmarking Clinical Note Generation. https://arxiv.org/abs/2610.08161
Cite the original work for its findings. Save a collection to share your selection of sources.