SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of training data scarcity, privacy constraints, and reliance on external APIs in medical question answering by proposing SAGE, a novel framework guided by lightweight semantic anchors such as MeSH. SAGE introduces a pioneering iterative evolution mechanism that integrates atomic and relational concepts, enabling local small language models to autonomously synthesize high-quality medical training data offline without requiring large-scale corpora or cloud support. Experimental results demonstrate that models fine-tuned on SAGE-synthesized data significantly outperform conventional methods across multiple benchmarks while substantially improving data efficiency. By eliminating dependencies on extensive resources and external services, this work establishes a new paradigm for developing medical AI systems in resource-constrained and privacy-sensitive environments.
📝 Abstract
Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-quality, expert-annotated training data. This challenge is exacerbated by stringent privacy requirements and the impracticality of utilizing large open-source corpora or proprietary cloud APIs within resource-limited clinical settings. To address these obstacles, we introduce SAGE (\textit{Semantic Anchor-Guided Evolution}), a novel data synthesis framework that enables small, locally deployed models to generate high-quality medical training data. SAGE leverages lightweight, publicly available taxonomies such as MeSH as semantic anchors, imposing a structured prior to effectively guide and ground the data generation process. At its core, SAGE iteratively interleaves atomic (individual concept-based) and associative (relation-based) synthesis, bootstrapping training data from minimal seeds. This approach eliminates the need for large collections of medical documents or reliance on external APIs, providing a practical solution for on-premises data creation. Extensive experiments across multiple medical question-answering benchmarks demonstrate that models fine-tuned with SAGE-synthesized data consistently outperform those trained using self-derived or conventional document-based paradigms, highlighting tangible improvements in data efficiency and resource utilization for medical LLM development. Code is available at https://github.com/DIaacKr/SAGE.
Problem

Research questions and friction points this paper is trying to address.

Medical Question Answering
Data Scarcity
Privacy Constraints
Resource-limited Clinical Settings
Training Data Synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Anchor-Guided Evolution
Medical QA Data Synthesis
MeSH Taxonomy
Iterative Bootstrapping
On-premises Deployment
🔎 Similar Papers