Reliability Testing of Medical Model Performance under Distributed Deployment

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of output deviations from offline evaluations caused by execution stack changes during the distributed deployment of medical AI models. To this end, it proposes the first distributed execution-aware benchmark for medical models. By developing an evaluation framework that compares model behaviors across centralized and distributed environments—incorporating techniques such as tensor parallelism and mixed precision—this work extends the assessment dimension from isolated capabilities to deployment consistency, with validation conducted on language, vision, and multimodal models. Experimental results demonstrate that execution variations induce measurable output discrepancies, yielding success rates of only 0.21–0.43 for unimodal and 0.32–0.98 for multimodal models. These findings reveal emerging challenges in evaluating the reliability of medical AI systems under distributed deployment scenarios.
📝 Abstract
Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.
Problem

Research questions and friction points this paper is trying to address.

distributed deployment
medical models
evaluation-deployment mismatch
reliability testing
output consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributed Inference
Evaluation-Deployment Mismatch
Reliability Testing
Medical Model Benchmark
Execution Consistency
🔎 Similar Papers
No similar papers found.