Score
Designs and implements methods that combine contextual embeddings with engineered or structured features into unified representations (via concatenation, projection, attention, or other learned fusion layers). Builds and evaluates predictive models and pipelines that consume these fused representations—e.g., regressors or rankers—and analyzes performance changes (RMSE, ranking metrics) resulting from the fusion.
Existing research lacks a systematic analysis of modality–language backbone integration mechanisms in multimodal large language models (MLLMs). Method: We systematically survey 125 MLLMs published between 2021 and 2025, proposing the first large-language-model-centric three-dimensional taxonomy—architectural integration, representation learning, and training paradigms—to unify cross-modal alignment pathways. Leveraging bibliometric analysis, architectural decoupling, and fine-grained modeling of embeddings and loss functions, we characterize evolutionary patterns in fusion granularity, joint representation design, and objective function development. Contribution/Results: This work fills a critical theoretical gap in structured analysis of MLLM fusion mechanisms, identifies shared bottlenecks—including misaligned modality-specific feature hierarchies and suboptimal cross-modal optimization objectives—and delivers a reusable theoretical framework and practical guidelines for developing robust, scalable multimodal foundation models.
To address the challenge of simultaneously achieving high accuracy, efficiency, and generation quality in large language models (LLMs) under multi-task settings, this paper proposes a context-aware dynamic parameter partitioning mechanism. The method dynamically partitions model parameters into context-sensitive subregions in real time based on gradient signals, enabling task-driven adaptive parameter allocation and knowledge fusion—without external fine-tuning or auxiliary modules. Its core innovation lies in introducing the first gradient-driven dynamic partitioning paradigm, empowering the model to autonomously recalibrate and specialize, thereby overcoming the limitations of conventional static parameter optimization. Experiments demonstrate substantial improvements across diverse language understanding and generation benchmarks: +2.1% accuracy, −8.7% perplexity, 32% reduction in memory footprint, 24% decrease in training time, and enhanced logical coherence and contextual consistency in generated text.
This paper addresses two fundamental challenges in foundation model research: the opaque nature of representation mechanisms and diminishing returns from scaling. To resolve these, we propose the “contexture” theory—a unified characterization of representation learning wherein optimal representations maximize mutual information between inputs and contextual variables, with peak generalization achieved at moderate contextual strength. We establish the first unified mathematical framework proving that scaling bottlenecks stem primarily from contextual *quality*, not scale. We introduce two general-purpose context-aware learning objectives—SVME and KISE—and a multi-context fusion strategy. Leveraging information theory and statistical learning theory, we derive a generalization bound for representation learning, unifying theoretical explanations across supervised, self-supervised, and generative pretraining paradigms. Empirical validation confirms that mainstream pretraining objectives implicitly optimize contexture. Our work provides both theoretical foundations and practical guidelines for designing efficient, context-driven pretraining frameworks.
Current foundational models lack a systematic theoretical characterization of their learned representations. Method: We propose Contexture Theory, which formalizes mainstream representation learning as the approximation of the top-d singular functions of an expectation operator—induced by associations between inputs and contextual variables. Our framework unifies supervised, self-supervised, and manifold learning under a common representational mechanism; identifies contextual quality—not parameter count—as the fundamental determinant of diminishing returns in model scaling; integrates singular function analysis, expectation operator modeling, context utility quantification, and generalization-theoretic proofs; and introduces a downstream-task-agnostic metric for context quality assessment. Results: Empirical evaluation across multiple real-world datasets demonstrates that our context quality metric exhibits strong correlation with downstream task performance, establishing a new, interpretable, and quantitatively assessable paradigm for representation learning.
Lack of a unified, reproducible evaluation benchmark hinders rigorous validation of effectiveness and robustness in deep model fusion. To address this, we introduce FusionBench—the first comprehensive benchmark specifically designed for deep model fusion—covering 26 tasks across open-vocabulary image classification, text classification, and generation. It integrates 74 fine-tuned models and 16 fusion methods. We formally define, for the first time, a systematic fusion evaluation paradigm comprising three categories: prediction-level ensembling, parameter-level merging, and component-level hybridization, accompanied by standardized task suites, model configurations, and evaluation protocols. Implemented in PyTorch, FusionBench supports both LoRA and full-parameter fine-tuning and incorporates mainstream algorithms including Task Arithmetic, DARE, SLERP, and Pareto Merging. Empirical analysis reveals significant differences in robustness across fusion methods under distribution shift. The codebase and documentation are publicly available and have been widely adopted as the de facto standard in the field.
This work addresses the challenge of detecting generative AI–produced content. We propose an unsupervised, interpretable embedding-space analysis method: semantic embeddings of text or images are extracted using pre-trained large language or multimodal models; subsequently, dimensionality reduction (e.g., PCA) uncovers an intrinsic, low-dimensional distributional shift between AI-generated and human-created samples—rendering them highly separable without supervision. This phenomenon is systematically validated for the first time and endowed with human-interpretable semantic meaning (e.g., topic coherence, syntactic redundancy). Experiments across diverse generative models—including ChatGPT, Gemini, and Stable Diffusion—demonstrate that high-accuracy separation is achieved solely from raw embeddings and unsupervised projection, without fine-tuning, labeled data, or model-specific detectors. Our approach thus significantly enhances both generalizability and interpretability of AI-content detection.
This work addresses the lack of principled understanding in existing literature regarding the choice between cross-attention and feature concatenation strategies for multimodal fusion, which has largely relied on empirical heuristics. Through controlled experiments and theoretical analysis, we demonstrate for the first time that feature alignment quality is the key determinant of fusion strategy performance: under pre-aligned features, concatenation consistently outperforms cross-attention by 4.1–5.1 percentage points across all data scales, with its advantage becoming more pronounced as alignment degrades. Building on this insight, we develop a theoretical decision framework grounded in sample complexity and validate our findings using features extracted from ResNet-18 and CLIP ViT-B/32 on controlled datasets.
Existing multimodal embedding models struggle to uniformly support text, images, video, and audio, often omitting audio or covering only a subset of modalities. This work proposes a unified four-modality embedding space built upon a frozen vision–language foundation model, augmented with a lightweight audio tower connector and modality-gated deep adapters, requiring no updates to the base model parameters. By aligning only audio with text, the approach enables cross-modal retrieval across all pairs—including audio–image—while fully preserving the original performance on pre-existing modalities. Both generations of the proposed model are trained within hours on a single GPU and achieve strong results in audio–text and audio–image retrieval. The code, model weights, and evaluation tools are publicly released.
This work addresses the limitations of existing meta-learning approaches for predicting machine learning pipeline performance (PPE) and estimating dataset similarity (DPSE), which predominantly rely on dataset meta-features while overlooking rich historical experiments and pipeline metadata, thereby failing to effectively model interactions between datasets and pipelines. To overcome this, the study introduces knowledge graph embedding into meta-learning for the first time, constructing a unified knowledge graph that integrates datasets, pipelines, and large-scale experimental results from 144,177 OpenML experiments. By jointly leveraging meta-features and empirical performance records, the proposed method—KGmetaSP—explicitly captures the complex interactions between datasets and pipelines. Remarkably, KGmetaSP achieves substantial improvements in both PPE prediction accuracy and DPSE retrieval effectiveness using a single, general-purpose meta-model, establishing a novel paradigm for cross-dataset meta-learning.