Symphony for Text Generation: Benchmarking Clinical Note Generation

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of quantitative evaluations regarding the impact of ambient scribing systems on clinical note quality by constructing a multilingual benchmark dataset and evaluation framework. Methodologically, it proposes a multidimensional controlled assessment system that integrates natural language inference metrics with large language model-based pairwise judgments. Leveraging the PDSQI-9 dimensions and API infrastructure, this work systematically compares the performance of Corti against commercial transcription software. The results demonstrate that Corti performs on par with or surpasses existing competitors. Furthermore, by open-sourcing reproducible benchmark data, this research establishes a novel paradigm for the standardized evaluation of clinical ambient scribing systems.
📝 Abstract
Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.
Problem

Research questions and friction points this paper is trying to address.

clinical note generation
ambient documentation systems
benchmarking
multilingual dataset
evaluation framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

Clinical Note Generation
Multilingual Dataset
Ambient Documentation Systems
LLM-as-a-Judge
Evaluation Framework
🔎 Similar Papers
Daniel Varab
Daniel Varab
V
Victor Petrén Bach Hansen
A
Asbjørn W. Helge
K
Kevin Pelgrims
M
Mathias Baltzersen
A
Adrian Young-San Roessler
V
Vanessa Klungtvedt
M
Maximilian Brand
L
Lasse Krogsbøll
H
Henrik Cullen
Lars Maaløe
Lars Maaløe
Co-Founder & CTO @ Corti | Adj. Assoc. Professor of Machine Learning @ DTU
Machine Learning