A Foundational Multi-Modal Model for Few-Shot Learning

πŸ“… 2025-08-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address data scarcity, high annotation costs, and poor generalization in few-shot learning (FSL) for scientific domains such as biomedicine and materials science, this paper introduces M3Fβ€”the first multimodal few-shot learning framework tailored for scientific computing. Methodologically, we construct M3FD, the first science-oriented multimodal few-shot benchmark dataset, encompassing images, medical scans, tabular data, and time-series signals; and propose LMMM, a modular large language and multimodal model architecture enabling cross-modal knowledge transfer and task-adaptive fine-tuning. In terms of contributions and results, M3F significantly outperforms conventional meta-learning approaches across diverse scientific tasks, empirically demonstrating the strong cross-domain generalization capability of large models in FSL. Furthermore, we publicly release the M3FD dataset, the M3F codebase, and associated tooling to foster reproducible, community-driven research in scientific few-shot learning.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionData Mining & Knowledge Management: Mining of Visual, Multimedia & Multimodal Data

Application Category

Web Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
πŸ“ Abstract
Few-shot learning (FSL) is a machine learning paradigm that aims to generalize models from a small number of labeled examples, typically fewer than 10 per class. FSL is particularly crucial in biomedical, environmental, materials, and mechanical sciences, where samples are limited and data collection is often prohibitively costly, time-consuming, or ethically constrained. In this study, we present an innovative approach to FSL by demonstrating that a Large Multi-Modal Model (LMMM), trained on a set of independent tasks spanning diverse domains, task types, and input modalities, can substantially improve the generalization of FSL models, outperforming models based on conventional meta-learning on tasks of the same type. To support this, we first constructed a Multi-Modal Model Few-shot Dataset (M3FD, over 10K+ few-shot samples), which includes 2D RGB images, 2D/3D medical scans, tabular and time-course datasets, from which we manually curated FSL tasks such as classification. We further introduced M3F (Multi-Modal Model for Few-shot learning framework), a novel Large Multi-Modal Model framework tailored for data-constrained scientific applications. M3F supports a wide range of scientific data types through a modular pipeline. By fine-tuning the model on M3FD, M3F improves model performance, making LMMM feasible for real-world FSL deployment. The source code is located at https://github.com/ptdang1001/M3F. To democratize access to complex FSL data and promote reproducibility for public usage, M3FD is paired with a flexible and user-friendly tool that enables efficient querying, task-specific sampling, and preprocessing. Together, our dataset and framework offer a unified, scalable solution that significantly lowers the barrier to applying LMMMs in data-scarce scientific domains.
Problem

Research questions and friction points this paper is trying to address.

Improves few-shot learning with multi-modal models
Addresses data scarcity in scientific domains
Enhances generalization across diverse tasks and modalities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Multi-Modal Model for diverse FSL tasks
Multi-Modal Few-shot Dataset with 10K+ samples
Modular pipeline supporting various scientific data
πŸ”Ž Similar Papers
No similar papers found.