🤖 AI Summary
Preoperative patients often face challenges in obtaining personalized, timely answers due to procedural time constraints and stringent privacy requirements. To address this, we propose LENOHA—a security-first, on-device clinical AI architecture that eliminates generative components entirely. It employs locally deployed sentence encoders (e.g., E5-large-instruct) for semantic classification of patient queries and retrieves precise answers from a physician-validated, static FAQ database, thereby eliminating hallucinations and uncontrolled text generation. The system operates on a single GPU, ensuring end-to-end privacy preservation, ultra-low energy consumption (1.0 mWh/request), and robust performance under low-bandwidth conditions. Evaluated in dental and gastroscopy preoperative settings, LENOHA achieves 0.983 accuracy, 0.996 AUC, only seven misclassifications, and a mean response latency of 0.10 seconds—matching GPT-4o’s performance while offering verifiable reliability and clinical deployability.
📝 Abstract
Patients awaiting invasive procedures often have unanswered pre-procedural questions; however, time-pressured workflows and privacy constraints limit personalized counseling. We present LENOHA (Low Energy, No Hallucination, Leave No One Behind Architecture), a safety-first, local-first system that routes inputs with a high-precision sentence-transformer classifier and returns verbatim answers from a clinician-curated FAQ for clinical queries, eliminating free-text generation in the clinical path. We evaluated two domains (tooth extraction and gastroscopy) using expert-reviewed validation sets (n=400/domain) for thresholding and independent test sets (n=200/domain). Among the four encoders, E5-large-instruct (560M) achieved an overall accuracy of 0.983 (95% CI 0.964-0.991), AUC 0.996, and seven total errors, which were statistically indistinguishable from GPT-4o on this task; Gemini made no errors on this test set. Energy logging shows that the non-generative clinical path consumes ~1.0 mWh per input versus ~168 mWh per small-talk reply from a local 8B SLM, a ~170x difference, while maintaining ~0.10 s latency on a single on-prem GPU. These results indicate that near-frontier discrimination and generation-induced errors are structurally avoided in the clinical path by returning vetted FAQ answers verbatim, supporting privacy, sustainability, and equitable deployment in bandwidth-limited environments.