Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决水处理研究知识分散问题,通过开发WaterBERT模型进行领域适应性预训练,提高大规模文献挖掘的语义表示和结构化信息提取能力。
📝 Abstract
Water treatment research is expanding rapidly, but much of the knowledge acquired from this research remains scattered across unstructured literature. The field still lacks a dedicated language model that can efficiently capture water treatment-specific domain semantics for large-scale literature mining. Here, we address this by developing WaterBERT, a domain-adapted encoder model designed for semantic representation and structured information extraction from water treatment texts. WaterBERT was developed by continual pretraining on a large-scale water treatment corpus comprising about 2.97 billion tokens. Three fine-tuned models based on WaterBERT were systematically evaluated on downstream tasks, achieving the best overall performance among general-purpose and domain-specific BERT models, with F1 scores of 90.12% for multiclass treatment process classification, 79.50% for named entity recognition, and 74.04% for relation extraction. Beyond these benchmark tasks, we further demonstrated WaterBERT's advantages for large-scale literature processing. Applied to 5,144 Environmental Science & Technology articles, WaterBERT-BERTopic identified coherent, diverse, and domain-specific research topics without predefined categories. Building on WaterBERT, we processed 693,211 abstracts at substantially lower cost than commercial LLMs while retaining competitive extraction performance to construct a structured water treatment knowledge graph. The knowledge graph was then integrated with lexical and dense retrieval to develop a Water Knowledge-Enhanced Retrieval System (WaterKERS), which achieved a relevance score of 77.7, substantially outperforming text-based retrieval baselines (54.7-64.5). Through WaterBERT, this study provides a compact and scalable semantic foundation for large-scale information processing and evidence mapping in water treatment research.
Problem

Research questions and friction points this paper is trying to address.

Water Treatment
Semantic Representation
Large-Scale Literature Mining
Domain-Adapted Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain-Adaptive Pretraining
Water Treatment Semantic Representation
Large-Scale Literature Mining
WaterBERT
Structured Information Extraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mudi Zhai
UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia
Ruihong Qiu
Ruihong Qiu
ARC DECRA Fellow, Lecturer (Assistant Professor) @The University of Queensland
GraphLarge Language Models
Q
Qingyun Zeng
Microsoft Copilot Studio AI, Redmond, WA 98052, United States; Departments of Mathematics & Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA 19104, United States
T. David Waite
T. David Waite
UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia
B
Bing-Jie Ni
UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia
Haoran Duan
Haoran Duan
Tsinghua/Newcastle/Durham University
Multimodal AIGenerative AI