Learnware of Language Models: Specialized Small Language Models Can Do Big

📅 2025-05-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) face significant reuse bottlenecks in data-scarce, privacy-sensitive, and computationally expensive scenarios. Method: This paper introduces the learnware paradigm—systematically adapted to language modeling for the first time—and proposes a learnware modeling framework tailored for small language models (SLMs). The framework enables cross-domain reuse of ~100 domain-specific SLMs (e.g., finance, healthcare, mathematics), each with up to 8B parameters, via capability specification-based representation, supporting privacy-preserving, on-demand discovery and zero-shot scheduling without sharing raw training data. Contribution/Results: Experiments demonstrate consistent and substantial improvements over baseline SLMs across all domains: +14% accuracy over Qwen1.5-110B on financial tasks and superior performance to Flan-PaLM-540B on healthcare tasks. These results validate the effectiveness and advancement of the learnware paradigm for lightweight, secure, and efficient model reuse.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
The learnware paradigm offers a novel approach to machine learning by enabling users to reuse a set of well-trained models for tasks beyond the models' original purposes. It eliminates the need to build models from scratch, instead relying on specifications (representations of a model's capabilities) to identify and leverage the most suitable models for new tasks. While learnware has proven effective in many scenarios, its application to language models has remained largely unexplored. At the same time, large language models (LLMs) have demonstrated remarkable universal question-answering abilities, yet they face challenges in specialized scenarios due to data scarcity, privacy concerns, and high computational costs, thus more and more specialized small language models (SLMs) are being trained for specific domains. To address these limitations systematically, the learnware paradigm provides a promising solution by enabling maximum utilization of specialized SLMs, and allowing users to identify and reuse them in a collaborative and privacy-preserving manner. This paper presents a preliminary attempt to apply the learnware paradigm to language models. We simulated a learnware system comprising approximately 100 learnwares of specialized SLMs with 8B parameters, fine-tuned across finance, healthcare, and mathematics domains. Each learnware contains an SLM and a specification, which enables users to identify the most relevant models without exposing their own data. Experimental results demonstrate promising performance: by selecting one suitable learnware for each task-specific inference, the system outperforms the base SLMs on all benchmarks. Compared to LLMs, the system outperforms Qwen1.5-110B, Qwen2.5-72B, and Llama3.1-70B-Instruct by at least 14% in finance domain tasks, and surpasses Flan-PaLM-540B (ranked 7th on the Open Medical LLM Leaderboard) in medical domain tasks.
Problem

Research questions and friction points this paper is trying to address.

Applying learnware paradigm to specialized small language models (SLMs)
Addressing data scarcity and privacy in specialized language tasks
Enhancing performance by reusing SLMs without exposing user data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reuse specialized small language models via learnware paradigm
Identify suitable models using privacy-preserving specifications
Outperform large models in specialized domains efficiently
💼 Related Jobs
No related jobs found.
Zhi-Hao Tan
Zhi-Hao Tan
Nanjing University
machine learninglearnware
Z
Zi-Chen Zhao
National Key Laboratory for Novel Software Technology, Nanjing University, China; School of Artificial Intelligence, Nanjing University, China
H
Hao-Yu Shi
National Key Laboratory for Novel Software Technology, Nanjing University, China; School of Artificial Intelligence, Nanjing University, China
X
Xin-Yu Zhang
National Key Laboratory for Novel Software Technology, Nanjing University, China; School of Artificial Intelligence, Nanjing University, China
Peng Tan
Peng Tan
National Key Laboratory for Novel Software Technology, Nanjing University, China; School of Artificial Intelligence, Nanjing University, China
Y
Yang Yu
National Key Laboratory for Novel Software Technology, Nanjing University, China; School of Artificial Intelligence, Nanjing University, China
Zhi-Hua Zhou
Zhi-Hua Zhou
Nanjing University
Artificial IntelligenceMachine LearningData Mining