🤖 AI Summary
General-purpose language models struggle to parse oncology-specific terminology and de-identification markers, limiting the efficiency of clinical natural language processing. To address this challenge, this study proposes OncoNoteBERT, a domain-specific encoder trained from scratch on a UK outpatient oncology corpus. The method employs a customized WordPiece tokenizer combined with masked language modeling pre-training to optimize domain-specific semantic representations. Experimental results demonstrate that the proposed model achieves a perplexity of 2.83 alongside significantly improved tokenization efficiency. Furthermore, it attains accurate predictions in 12 out of 13 probing tasks, outperforming external baseline models. These findings establish OncoNoteBERT as an efficient, specialized solution for processing oncological text in clinical settings.
📝 Abstract
Real-world outpatient oncology notes contain specialised terminology, tumour staging expressions, treatment names, toxicity descriptions, and institution-specific de-identification markers that may not be represented efficiently by general biomedical or adjacent clinical language models. We developed and evaluated oncology-specific BERT-style encoders using a governed UK outpatient oncology corpus comprising 290,026 notes from 21,564 patients treated for lung and head-and-neck cancer. We compared RadBERT and PathologyBERT with two local strategies: OncoNote-RadBERT, produced by continued masked language model pretraining, and OncoNoteBERT, trained from scratch with an oncology WordPiece tokenizer. Models were evaluated using masked language modelling loss and perplexity on the validation set, tokenizer fragmentation metrics, clinical term tokenisation, masked-token probes, and exploratory representation analysis. Both external encoders fit the oncology corpus poorly in zero-shot evaluation (perplexity 113.04 for RadBERT; 2035.03 for PathologyBERT), while continued pretraining produced the strongest fit (2.10 for OncoNote-RadBERT). OncoNoteBERT achieved perplexity 2.83 but produced the most efficient tokenisation, with lower subword fertility and shorter normalised sequence length. It also returned a clinically acceptable prediction for 12 of 13 masked-token probes, compared with 7 of 13 for OncoNote-RadBERT. This divergence between corpus-level fit and masked-token performance was partly attributable to tokenizer fragmentation rather than learned semantics alone. Both locally developed models represented the institutional placeholder as a single learnable token. These findings show that continued adaptation and bespoke tokenisation provide complementary benefits, and that representation-layer design matters before adjacent-domain encoders are applied to oncology NLP.