ClinicalAligner26AM: A Cross-Lingual Aligner for Dataset Translation; Evidences from the MultiClinCorpus Shared Task

๐Ÿ“… 2026-06-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the limited word-level cross-lingual alignment performance of existing neural alignment models in biomedical and clinical domains by proposing a novel approach that integrates multi-granularity signal fusion with optimal transport sharpening. Starting from a ClinicalEncoder26AM-initialized multilingual large-context model, the method constructs a cost matrix by combining sentence-, phrase-, and token-level signals, then employs the Sinkhorn-Knop algorithm to generate soft alignment targets. A lightweight student aligner is trained via knowledge distillation to match these targets in terms of cosine similarity. During inference, source segment scores are projected through the alignment matrix, and the longest high-scoring valid span is decoded. This is the first study to combine multi-granularity fusion with optimal transport sharpening for clinical text alignment, achieving first and second places across all languages and entity types in the MultiClinCorpus shared task, with character-weighted F1 scores consistently exceeding 0.95.
๐Ÿ“ Abstract
Word-level cross-lingual alignment is central to annotation projection, translation auditing, and cross-lingual faithfulness estimation, yet existing neural aligners are rarely adapted to specialized domains. In this paper, we introduce ClinicalAligner26AM, a large-context multilingual aligner model for biomedical and clinical text initialized from ClinicalEncoder26AM. Our training recipe is inspired by AWESoME Align. We build our soft alignment target by sharpening with Sinkhorn-Knop optimal transport a cost matrix established for parallel clinical texts and conversations through the fusion of sentence-level, phrase-level, and token-level signals. We distill this sharpened alignment matrix directly into our student aligner, by encouraging its naive cosine-based token similarity scores to match this target. At inference time, we project source-span scores through the learned token alignment matrix and decode the longest valid high-scoring span in the target text, optionally supported by MultiClinNER predictions summarized in Appendix B. We evaluate CA26AM on the MultiClinCorpus shared task, which projects Spanish clinical entity annotations into six target languages. Our two submitted systems ranked respectively first and second across all languages and entity types, with character-weighted F1 scores above 0.95 in nearly all settings.
Problem

Research questions and friction points this paper is trying to address.

cross-lingual alignment
clinical text
annotation projection
specialized domains
biomedical NLP
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-lingual alignment
clinical text
optimal transport
alignment distillation
annotation projection
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.