SciTBERT: A family of chronologically consistent language models for scientific and technological language processing

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the temporal lookahead bias and domain distribution mismatch inherent in pretrained language models when processing scientific and technological text. To mitigate these issues, we propose a temporally consistent training paradigm that leverages year-bounded corpora of academic papers and patents for pretraining, complemented by a citation-based post-training strategy to eliminate temporal leakage. This approach yields the SciTBERT model family. Furthermore, we introduce PatRepEval, a novel benchmark designed to effectively bridge the representation spaces of science and technology. Experimental results demonstrate that the proposed models outperform existing domain-specific baselines across multiple downstream tasks, thereby validating the critical role of temporal alignment in enhancing the understanding of scientific and technological text.
📝 Abstract
Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025. We also post-train these models in a chronologically-consistent manner using paper and patent citations, creating SciTBERT-CI model family. We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus. To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface. Performance in a variety of classification, regression, and retrieval tasks spanning papers and patents highlights the importance of aligning encoder model representations with the domain distributions of their downstream tasks, and chronologically consistent encoders can match or exceed models trained without temporal constraints.
Problem

Research questions and friction points this paper is trying to address.

pre-trained language models
chronological consistency
lookahead bias
domain bias
science-technology interface
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chronologically consistent language models
SciTBERT
Citation-based post-training
PatRepEval benchmark
Science-technology interface
🔎 Similar Papers