Score
Designs, builds, and evaluates methods to transfer and adapt pretrained foundation models to new tasks and domains, including domain alignment, representation transfer, prompt-based adaptation, attaching lightweight task-specific heads, and parameter-efficient fine-tuning (e.g., freezing backbones or tuning minimal parameters). Also develops and analyzes corresponding adaptation, prompting, transfer, and evaluation techniques specifically for tabular foundation models.
This study addresses the insufficient exploration of relationships and trade-offs among different paradigms for cross-task adaptation of large language models by constructing the first unified analytical framework. Methodologically, it systematically integrates parameter-efficient fine-tuning, in-context learning, and embedding injection techniques, establishing a comprehensive taxonomy along the dimensions of weights, prompts, and embeddings. Employing a systematic review methodology, the work thoroughly analyzes the strengths, limitations, and intrinsic connections of each paradigm. The findings reveal the evolutionary logic underlying these three paradigms, yielding a holistic taxonomic landscape that clarifies their core advantages and constraints while identifying key open problems. Ultimately, this research provides strategic guidance for future investigations into the efficient adaptation of large models.
General-purpose foundation models struggle to adapt to domain-specific patterns and requirements. Method: This study systematically establishes a theoretical and methodological framework for Domain-Specific Foundation Models (DSFMs), proposing the first cross-industry reusable DSFM customization framework that integrates pretraining-finetuning, instruction alignment, domain adaptation, knowledge injection, and efficient parameter updating, accompanied by a multi-dimensional evaluation system. Results: The framework is validated across ten+ domains—including finance, healthcare, and manufacturing—yielding a comprehensive application landscape; it identifies three universal bottlenecks: data scarcity, computational constraints, and regulatory compliance barriers. This work fills a critical gap in DSFM surveys and delivers the first authoritative, academically rigorous methodology guide and practical reference for industry–academia–research collaboration in DSFM development.
Parameter-efficient fine-tuning (PEFT) methods for large foundation models suffer from unclear cross-modal adaptation mechanisms and a lack of systematic classification frameworks. Method: This paper introduces the first unified PEFT taxonomy spanning language, vision, and multimodal foundation models, systematically analyzing the adaptation principles and architectural distinctions of mainstream techniques—including LoRA, Adapter, Prompt Tuning, and IA³—within the Transformer architecture. Through comparative analysis, it identifies evolutionary trends and critical research gaps. Contribution/Results: We publicly release an authoritative, curated PEFT literature repository on GitHub, featuring a structured knowledge graph and practical implementation guidelines. This resource significantly lowers the barrier to task-specific customization of large models and establishes foundational support for advancing PEFT theory and enabling robust cross-modal applications.
In few-shot domain adaptation, two key challenges hinder performance: (1) hyperparameter tuning is impractical due to scarce validation data, and (2) models lack robustness under distribution shift. To address these, we propose Soup-Adapter—a hyperparameter-agnostic, distributionally robust adapter ensemble method. Our approach extends the CLIP adapter paradigm to DINOv2 for the first time and introduces a reparameterizable adapter soup: multiple adapters are trained independently via distinct paths, their outputs are averaged at inference, and their parameters are concatenated and reparameterized to stabilize optimization. Experiments demonstrate that Soup-Adapter consistently outperforms single-adapter baselines across a wide range of hyperparameters and significantly improves both accuracy and robustness on cross-domain few-shot tasks. This work establishes a new, efficient, practical, and generalizable paradigm for few-shot domain adaptation.
To address structural distortion and topological inconsistency in high-dimensional parameter spaces (e.g., 4D tensors) induced by low-rank approximation in parameter-efficient fine-tuning, this paper proposes a structure-preserving low-rank core space modeling method. Unlike conventional low-rank adapters (e.g., LoRA), which are restricted to linear weight matrices, our approach explicitly models and preserves the intrinsic topological structure of the original high-dimensional parameter space—achieving compact and accurate reconstruction of N-dimensional parameter updates via high-order tensor decomposition. Evaluated across CV, NLP, and multimodal benchmarks, the method yields an average accuracy improvement of 1.8% under identical parameter budgets, while reducing structural distortion by 37%, significantly outperforming existing baselines.
Standard fine-tuning of foundation models suffers from low downstream adaptation efficiency and fails to recover the optimal adaptable parameter set. Method: We propose the first PEFT co-optimization framework that explicitly integrates meta-learning (MAML-style) into the foundation model’s retraining phase, using a LoRA-inspired low-rank adaptation structure. Contribution/Results: We theoretically prove that standard retraining is inherently suboptimal in adaptability, whereas our method strictly recovers the optimal adaptable parameters and provides a generalization error bound. Experiments on RoBERTa with the ConvAI2 dialogue continuation task demonstrate significant improvements in zero-shot and few-shot rapid adaptation performance, empirically validating the theoretical guidance.
This study systematically evaluates the applicability of the foundation model (FM) paradigm across three scientific domains—genomics, satellite imagery, and time-series analysis—to assess whether FMs can supplant traditional supervised learning. Method: We construct a cross-modal benchmark framework employing lightweight architectures (e.g., Wide ResNet, U-Net), automated hyperparameter optimization, and standardized training protocols to rigorously compare domain-specific FMs against strong supervised baselines. Contribution/Results: Across all tasks, carefully tuned supervised models match or exceed state-of-the-art domain-specific FMs; large-scale pretraining yields no consistent empirical gains. This work provides the first multi-modal scientific validation that the FM paradigm remains immature for these domains. We open-source two automated evaluation workflows and underscore the necessity—and benchmarking value—of strong supervised baselines in scientific AI assessment.
This work addresses the limitations of tabular foundation models in small-scale specialized domains, where performance is hindered by data scarcity, high dimensionality, distributional shifts, and the absence of mechanisms to effectively incorporate structured priors such as knowledge graphs. To overcome these challenges, the authors propose a knowledge-guided fine-tuning approach that, for the first time, integrates structural attention priors derived from knowledge graphs with low-rank adaptation (LoRA) to efficiently inject domain knowledge during fine-tuning. Experiments on lightweight tabular models—including TabPFN and TabICL—demonstrate that the method substantially outperforms standard fine-tuning on specialized tasks, while yielding only marginal gains on general tasks. These results validate the efficacy of structured knowledge injection and reveal that continued fine-tuning may lead to the collapse of pre-trained knowledge.
This study investigates how to achieve full-stack transfer learning in robotics across multiple levels of abstraction, from high-level language instructions to low-level motor skills. By unifying the analysis of common transfer mechanisms in large language models (LLMs), vision-language models (VLMs), and vision-language-action models (VLAs), this work offers the first systematic exploration—from the perspective of robotic transfer learning—of the role played by foundation models and Transformer architectures in bridging high-level semantic understanding with low-level action generation. The findings demonstrate that foundation models constitute a critical pathway toward full-stack transfer in robotics, while also highlighting key challenges such as the need for high-quality data collection and the establishment of standardized transfer benchmarks.
Large language models (LLMs) face significant challenges in task adaptation under resource-constrained and closed-source API settings, where conventional parameter-efficient fine-tuning (PEFT) methods are inapplicable due to their reliance on direct model parameter access and high computational overhead. Method: This paper proposes a lightweight, parameter-free knowledge injection framework that enables task-specific adaptation without accessing the LLM’s internal parameters. Its core innovation is the “Specialized Small Model (SSM) Collaboration Paradigm,” integrating knowledge distillation from the LLM, distribution-aware task modeling, and zero-parameter coupling between the SSM and the LLM. Contribution/Results: Experiments demonstrate that our approach matches PEFT-level performance across diverse downstream tasks while reducing GPU memory consumption by over 90% and inference latency by 85%. Crucially, it operates entirely within black-box API environments—requiring no model weights, gradients, or architectural access—thus enabling seamless integration with proprietary, closed-source LLM APIs.
This study systematically investigates the effectiveness and applicability of fine-tuning in Tabular Foundation Models (TFMs), aiming to answer when and how fine-tuning can improve performance, calibration, and fairness. Through comprehensive evaluation across benchmarks including TALENT, OpenML-CC18, and TabZilla, the work compares zero-shot inference, meta-learning, full supervised fine-tuning (SFT), and parameter-efficient fine-tuning (PEFT). The findings reveal that zero-shot TFMs already exhibit strong performance, while SFT often degrades accuracy or calibration quality. Meta-learning and PEFT yield only marginal gains under specific data conditions. This work is the first to demonstrate that fine-tuning is not universally beneficial and provides practical, data-characteristic-driven guidelines for its application, clearly delineating the boundaries of its advantages.
This study addresses the misalignment between large language models (LLMs) and human experts in child-oriented instructional tasks, where such discrepancies can impair learning outcomes. For the first time, it quantitatively evaluates the alignment of multiple mainstream foundation models in out-of-distribution educational settings, revealing shared biases across models and demonstrating that ensemble methods may exacerbate negative alignment with desired learning outcomes. Through cross-model behavioral correlation analysis, expert-weighted ensembling, benchmark comparisons, and robust alignment metrics, the research finds that 50% of misalignment errors are shared across models, pointing to pretraining data as a critical source of these systematic deviations. This work provides crucial insights into the limitations of current LLMs for educational AI applications, highlighting the need for targeted alignment strategies in pedagogical contexts.