Multi-Level Evidence Aggregation for Robust Facial Phenotype Retrieval in Rare Genetic Disorder Prioritization

πŸ“… 2026-08-11
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of existing facial phenotyping retrieval methods, which rely solely on single-image matching and fail to leverage aggregated evidence from multiple images of the same patient or multiple cases of the same rare genetic disorder, thereby constraining diagnostic ranking accuracy. To overcome this, the authors propose an inference-time, multi-level evidence fusion framework that operates without retraining the encoder. Building upon the GestaltMatcher-Arc encoder, the approach integrates individual-level image embedding aggregation, disease-level weighted centroid construction, and a hybrid individual-to-centroid scoring strategy, shifting the paradigm from single-image matching to patient- and disease-level evidence integration. Evaluated on GMDB v1.1.4, the method significantly improves Top-1 accuracy, raising it from 38.52% to 48.82% on GMDB-Freq and from 19.38% to 23.79% on GMDB-Rare, with performance reaching up to 60.94% on multi-image subsets.
πŸ“ Abstract
AI-assisted facial phenotyping supports rare genetic disorder prioritization by retrieving visually similar diagnosed cases from facial image reference databases such as the GestaltMatcher Database (GMDB). Existing GestaltMatcher-based retrieval frameworks compare each test image with individual gallery images in a facial phenotype embedding space. However, this pointwise formulation does not fully exploit available evidence, because patients may have multiple images and disorders may be represented by multiple diagnosed gallery patients. We propose an inference-time multi-level evidence aggregation framework that improves facial phenotype retrieval without modifying the underlying GestaltMatcher-Arc encoder. The framework combines embedding-level patient aggregation of multiple images from the same individual, patient-weighted disorder centroids, and hybrid individual-centroid scoring to integrate test-patient observations, disorder-level gallery evidence, and local nearest-neighbor evidence. We evaluated the approach on GMDB v1.1.4 across disorders represented during training (GMDB-Freq), unseen disorders (GMDB-Rare), and multi-image patient subsets, using a unified gallery containing both GMDB-Freq and GMDB-Rare disorders. Multi-level evidence aggregation improved mean per-disorder top-$N$ retrieval accuracy across all evaluation subsets. Top-1 accuracy increased from 38.52% to 48.82% on GMDB-Freq and from 19.38% to 23.79% on GMDB-Rare. On multi-image subsets, top-1 accuracy increased from 46.12% to 60.94% on GMDB-Multi-Freq and from 18.54% to 26.71% on GMDB-Multi-Rare. These findings show that inference-time aggregation can improve next-generation facial phenotype retrieval without retraining the encoder, supporting a shift from isolated single-image matching toward multi-level aggregation of patient and disorder evidence for rare-disorder prioritization.
Problem

Research questions and friction points this paper is trying to address.

facial phenotyping
rare genetic disorder
evidence aggregation
image retrieval
phenotype embedding
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-level evidence aggregation
facial phenotyping
rare genetic disorder prioritization
inference-time fusion
embedding aggregation
A
Alexander Hustinx
Institute for Genomic Statistics and Bioinformatics, University Hospital Bonn, Bonn, Germany
C
Carolin KaffinΓ©
Institute for Genomic Statistics and Bioinformatics, University Hospital Bonn, Bonn, Germany
Behnam Javanmardi
Behnam Javanmardi
Institute for Genomic Statistics and Bioinformatics, University Hospital Bonn, Bonn, Germany
Tzung-Chien Hsieh
Tzung-Chien Hsieh
Institute for Genome Statistics and Bioinformatics, University Bonn
BioinformaticsComputer ScienceGenetics
P
Peter Krawitz
Institute for Genomic Statistics and Bioinformatics, University Hospital Bonn, Bonn, Germany