Score
Designs and builds multimodal datasets and benchmarks that combine and align heterogeneous modalities (e.g., images, text/ocr, audio/transcripts, video, 3D poses/boxes, gaze/screen/smartwatch sensor streams, and robotic demonstrations), including protocols for data collection, synchronization, simulated vs real splits, and synthetic-data generation. Develops annotation schemas and tooling (e.g., OCR/context/evidence labels, segmentation, 3D detection, provenance and relevance scoring), along with quality-control, preprocessing, feature-extraction, filtering, and assembly steps to produce task-specific corpora such as meme annotations, audio–transcript corpora, provenance datasets, and dexterous-manipulation/demonstration collections.
This paper systematically analyzes four core challenges in multimodal alignment and fusion—cross-modal misalignment, semantic modality gap, computational bottlenecks, and data noise/heterogeneity—that hinder model robustness, generalizability, and scalability. To address them, the authors propose a unified taxonomy covering 200+ studies, revealing an integrated “alignment–fusion” co-design paradigm. They further introduce three principled solution pathways: (1) noise-robust learning, (2) heterogeneous representation modeling, and (3) few-shot cross-modal transfer—unifying contrastive learning, cross-modal attention, latent-space alignment, graph neural networks, and differentiable architecture search. Extensive experiments on social media analysis, medical imaging, and sentiment recognition validate the efficacy of the framework. The work distills actionable design principles for scalable, robust, and generalizable multimodal learning, establishing a new benchmark for both theoretical research and industrial deployment.
Existing multimodal evaluation benchmarks inadequately reflect real-world, heterogeneous daily usage scenarios and lack systematic assessment across diverse tasks and output formats. Method: We introduce the first fine-grained, real-scenario-oriented multimodal benchmark—comprising 505 practical scenarios and 8,000+ samples—supporting 16 input/output modalities and 40+ output formats (e.g., numbers, code, JSON, free-form text). We propose a four-dimensional capability reporting framework—“Application–Input–Output–Skill”—replacing monolithic multiple-choice evaluation with task-driven, format-aware, interpretable assessment. The benchmark integrates expert crowdsourced scenario sampling, 40+ customized automated metrics, multi-format parsers, and interactive visualization tools. Contribution/Results: Comprehensive evaluation of state-of-the-art vision-language models reveals, for the first time, their fine-grained capability boundaries and long-tail deficiencies across modality combinations and task types.
Existing surveys predominantly examine isolated components of multimodal pipelines and lack empirically grounded, pedagogically oriented integration frameworks for teaching and learning contexts. Method: This study introduces the first taxonomy and analytical framework covering five core modalities—natural language, video, sensor data, human-centered signals, and environmental logs—and proposes a novel “mid-fusion” paradigm for multimodal data integration. It further innovates by applying citation graph pruning to achieve structured, high-precision literature synthesis. Contribution/Results: Through systematic review, taxonomic modeling, and multimodal fusion design, we demonstrate that multimodal synergy enables detection of fine-grained learning behaviors imperceptible to unimodal analysis. While prediction accuracy remains largely unchanged, interpretability improves significantly, yielding deeper insights into learners’ cognitive-affective states and training outcomes.
Traditional multimodal approaches rely on modality-specific encoders and late fusion, limiting scalability and cross-modal generalization. This work proposes a unified paradigm that reformulates diverse multimodal tasks—including text, image, audio, and video processing—as “next-frame prediction,” with all inputs and outputs represented as serialized video frames, enabling end-to-end, single-model inference. It pioneers the task-redefinition strategy—previously confined to NLP—within multimodal learning, eliminating modality-specific design in favor of a fully modality-agnostic architecture. The framework integrates cross-modal tokenization, sequential frame representation, autoregressive modeling, and a shared Transformer decoder, supporting joint multi-task pretraining. Empirical evaluation demonstrates strong zero-shot and few-shot generalization across text-to-text, image-to-text, video-to-video, video-to-text, and audio-to-text tasks. Crucially, adaptation requires only lightweight task-specific heads, underscoring its flexibility and efficiency.
To address the commercial deployment challenges of multimodal AI in document understanding and code generation—stemming from limited training data and restrictive licensing—this paper introduces BigDocs: the first high-quality, traceable, and license-compliant open multimodal dataset for documents and code (7.5 million samples across 30 task categories) and its associated benchmark, BigDocs-Bench (featuring 10 real-world tasks, e.g., Screenshot2HTML and Image2LaTeX). We propose novel evaluation paradigms, including GUI-aware and image-driven code generation. Our data curation pipeline integrates automated content analysis, license-compliance filtering, structured metadata tracing, and human verification. Models trained on BigDocs achieve an average 25.8% performance gain over GPT-4o across multiple tasks, with human evaluations strongly favoring their outputs.
This work addresses the challenge that certain modalities may introduce interference under specific inputs in multimodal fusion, thereby degrading model performance. To mitigate this issue, the authors propose a plug-and-play pre-fusion calibration module that leverages cross-modal summary contrast to extract supportive and conflicting cues, generating instance-level and dimension-level modulation signals. These signals dynamically enhance beneficial features while suppressing misleading information. The method achieves, for the first time, fine-grained, conflict-aware modulation prior to fusion and is compatible with both sequential and convolutional architectures. It consistently improves performance across five benchmark tasks—including emotion understanding and action recognition—and demonstrates enhanced robustness and consistency under modality missingness and data perturbations.
Current large multimodal language models (LMMs) lack systematic evaluation of their multimodal interaction capabilities. To address this gap, this work proposes MIBench, a structured benchmark that, for the first time, assesses LMMs along two dimensions—modality source bias and multimodal collaborative generation—across three cognitive levels: recognition, comprehension, and reasoning. The benchmark comprises 32 task categories and over 10,000 sample pairs, organized into a unified framework (con_v, con_t, task). Evaluation using MIBench reveals pervasive issues in existing LMMs, including strong text-dominant bias and weak collaborative generation ability. Notably, even native multimodal models exhibit fundamental deficiencies in basic interaction mechanisms. These findings provide critical insights and concrete directions for future research in multimodal interaction modeling.
This study addresses the challenge that multimodal models struggle to reliably bind textual prompts (e.g., “image”) to their corresponding input modalities, resulting in inadequate source modality tracking. The work formally defines and empirically investigates the “source modality monitoring” problem for the first time, introducing an evaluation paradigm grounded in target-modality information retrieval. By integrating syntactic manipulation with semantic perturbation, the authors systematically assess binding mechanisms across eleven prominent vision-language models. Their findings reveal that semantic cues dominate the binding process when modality distributions exhibit significant divergence, consistently outweighing syntactic signals. These insights offer critical evidence and a novel perspective for enhancing the reliability and robustness of multimodal agents.
This work proposes a unified vision model that transcends task-specific architectures traditionally required in computer vision by formulating diverse visual tasks—including detection, segmentation, and geometric prediction—as multimodal generation problems. The model is driven solely by natural language instructions (optionally augmented with visual prompts) to produce text, images, or hybrid outputs from a single architecture, eliminating the need for specialized heads or modular designs. It introduces the SenseNova-Vision Corpus, a large-scale dataset of vision-language instruction-response pairs, and leverages off-the-shelf pretrained multimodal models refined through instruction tuning and joint multimodal generation strategies. This end-to-end framework supports compositional, language-defined tasks and achieves performance on par with or superior to specialized systems across structured understanding, dense geometric prediction, segmentation, and multi-view geometry benchmarks.