Score
Designs, builds, and analyzes datasets and processing pipelines that integrate two or more complementary data modalities (such as text, images, audio, video, or sensor time‑series), including collection, alignment, preprocessing, multimodal representation learning, and evaluation to support downstream models and analyses.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
A critical shortage exists of high-quality, large-scale, semantically aligned acoustic–visual–textual multimodal datasets. Method: This paper proposes an end-to-end framework for constructing video-based multimodal data, integrating three key components: (i) video content filtering, (ii) cross-modal synchronization triplet extraction (audio–frame–subtitle), and (iii) fine-grained description synthesis leveraging image-to-text generation models—ensuring temporal and semantic alignment across all three modalities. Contribution/Results: The resulting publicly released dataset spans diverse real-world scenarios and substantially advances performance on cross-modal retrieval and joint embedding learning tasks, achieving state-of-the-art results across multiple benchmarks. By providing a scalable, high-fidelity resource, this work establishes a new foundation for training and evaluating foundational multimodal models.
Existing models are largely restricted to unimodal or bimodal architectures, limiting their capacity to efficiently integrate the rich diversity of visual modalities present in real-world scenarios. To address this, we propose a low-supervision, fully automated multimodal video data pipeline that enables programmable composition and joint learning across heterogeneous visual modalities—including RGB, depth, optical flow, and edge maps. Our key contributions are: (1) PHG-MAE, a lightweight multimodal self-supervised encoder (<1M parameters), which leverages pretrained expert models and knowledge distillation to achieve performance on par with 300M-parameter large models; and (2) seamless integration of off-the-shelf modules (e.g., DPT) to enable real-time semantic segmentation and near-real-time depth estimation from handheld or webcam video on commodity hardware. Extensive experiments demonstrate the pipeline’s efficiency, scalability, and strong generalization under resource-constrained conditions.
This paper addresses the emerging challenge of multimodal time series forecasting by introducing the first systematic solution: the Multimodal Time Series Benchmark (MMTSB), encompassing textual, visual, and metadata modalities. It supports both conventional historical-dependency forecasting and zero-shot cold-start forecasting. Methodologically, we propose novel techniques for aligning heterogeneous multimodal data, cross-modal temporal annotation, controllable cold-start dataset partitioning, and a standardized evaluation protocol. Empirical results—first to quantitatively demonstrate significant performance gains from external modalities in short-sequence and cold-start forecasting—reveal principled relationships between modality utility and intrinsic data characteristics. All datasets, code, and analytical results are publicly released to advance time series forecasting toward realistic, complex scenarios.
To address inefficiencies in model incremental updates, unfair policy evaluation, and high retraining costs under continual data growth, this paper proposes an end-to-end adaptive machine learning platform. Methodologically: (1) it introduces a declarative domain-specific language (DSL) to uniformly model data selection strategies (e.g., coreset, uncertainty sampling) and trigger policies (e.g., drift-aware scheduling); (2) it establishes the first composite model evaluation framework enabling fair, cross-policy comparison; and (3) it implements a co-optimization mechanism integrating sample-level fine-grained data selection with high-throughput training. Contributions include an open-source, extensible system architecture, a standardized benchmark ecosystem, and abstracted ML pipeline interfaces. Experiments demonstrate significant improvements in training throughput and substantial reductions in retraining overhead—while preserving model accuracy—and enable reproducible analysis across diverse strategy combinations.
Existing time-series analysis models are restricted to numeric modalities and struggle to incorporate domain-specific textual knowledge, resulting in limited modeling capacity and a lack of high-quality multimodal benchmarks. To address this, we introduce Time-MMD—the first large-scale multimodal time-series dataset spanning nine diverse domains—featuring fine-grained semantic alignment between numeric sequences and domain-specific textual descriptions. We further propose a multi-domain multimodal time-series benchmark that systematically tackles modality contamination and the absence of cross-modal alignment. Additionally, we open-source MM-TSFlib, the first modular library for multimodal time-series forecasting, enabling joint modeling and granular evaluation. Experiments demonstrate that our approach reduces average MSE by over 15% on multimodal forecasting tasks, with gains reaching 40% in text-rich scenarios. This work advances time-series analysis from unimodal paradigms toward human-AI collaborative multimodal reasoning. Code and data are publicly available.
This work addresses the long-standing challenge in multiphase transport and thermal systems research—namely, fragmented data and non-reusable raw instrument files that hinder reproducibility and benchmarking. To overcome this, the authors propose the S+TD spatiotemporal dimensional classification framework and establish an open-source ecosystem that integrates multimodal data, including boiling imagery, high-speed video, infrared thermography, and CFD field outputs, into the publicly available NED3 dataset. Complementing this data infrastructure, they develop a suite of tools—SeqReg for sequence regression, BubbleID for bubble identification, CFDTwin for digital twin modeling, and IRISApp for thermal analysis—to support computer vision, acoustic decoding, and surrogate modeling. The platform enables applications such as non-intrusive heat flux estimation and establishes a reproducible, interoperable paradigm for AI-driven thermal-fluid research.
This study addresses the critical scarcity of high-quality, multimodal, and structured open datasets in the energy domain that are essential for advancing large language model applications. To bridge this gap, the authors introduce mAIEnergy, the first comprehensive multimodal energy dataset integrating heterogeneous sources—including policy documents, scientific literature, power system records, meteorological data, and infrastructure information—across four modalities: text, images, time series, and geospatial data. Through standardized preprocessing, unified metadata management, and strict adherence to FAIR principles, mAIEnergy provides a consistent, reproducible corpus comprising 50,000 textual documents, 20,000 images, 25 million time-series entries, and 2 million geospatial records. This resource establishes a foundational infrastructure for AI-driven energy research and decision-making.
本文介绍NeMo数据设计器,一种用于生成多模态合成数据的开源框架,通过声明式配置和插件系统提高数据集多样性和工作流可重复性。
为解决视频预训练数据管道封闭问题,提出VIDAFORGE,一个可执行五阶段工作流的开放研究基础设施,通过对比不同数据配方对模型性能的影响来优化视频预训练。
This work addresses the limitations of existing visual data mining tools, which often operate as isolated applications and cannot be readily embedded into web environments, thereby hindering the sharing and interactive integration of analytical workflows. To overcome this, the paper introduces a “component exposition” paradigm and presents a web-based collaborative visual analytics environment that enables users to construct machine learning pipelines through modular components. The system supports real-time state propagation and dynamic exploration across components by integrating modular visual programming, reactive dataflow, and web-embedding technologies. Crucially, any component within a workflow can be seamlessly embedded into external web pages, abstracting away underlying complexity while enabling customizable views and narrative-driven data experiences. Deployment in data literacy education demonstrates that this approach significantly lowers the barrier for users to understand and apply machine learning techniques.