🤖 AI Summary
This study addresses the lack of quantitative evaluations regarding the impact of ambient scribing systems on clinical note quality by constructing a multilingual benchmark dataset and evaluation framework. Methodologically, it proposes a multidimensional controlled assessment system that integrates natural language inference metrics with large language model-based pairwise judgments. Leveraging the PDSQI-9 dimensions and API infrastructure, this work systematically compares the performance of Corti against commercial transcription software. The results demonstrate that Corti performs on par with or surpasses existing competitors. Furthermore, by open-sourcing reproducible benchmark data, this research establishes a novel paradigm for the standardized evaluation of clinical ambient scribing systems.
📝 Abstract
Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.