late fusion

Designs and implements architectures and algorithms that combine outputs, embeddings, or predictions from multiple pretrained or concurrently trained encoders at the decision/output level to produce a single fused representation or final prediction. This includes building gated fusion mechanisms and networks, ensemble/decision-level fusion strategies, and post-hoc integration methods that learn per-encoder weights or pooling rules and that are evaluated for robustness and ranking/performance effects.

latefusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Exploring Synergistic Ensemble Learning: Uniting CNNs, MLP-Mixers, and Vision Transformers to Enhance Image Classification

Apr 12, 2025
MB
Mk Bashar
🏛️ Michigan State University | University of Illinois Urbana-Champaign | Islamic University of Technology | Motional

Single-architectural vision models face inherent performance bottlenecks in image classification. Method: This paper proposes a cross-paradigm ensemble framework that preserves architectural integrity by integrating three heterogeneous architectures—CNN (ResNet), MLP-Mixer, and Vision Transformer—via weighted averaging and majority voting, without architectural modification or feature-space alignment. Crucially, the framework implicitly isolates their respective feature spaces while enabling synergistic gains. Contribution/Results: It introduces the first complementary analysis paradigm grounded in architectural orthogonality. Evaluated end-to-end on ImageNet, the ensemble surpasses prior single-model SOTA in top-1 accuracy while reducing overall inference latency—establishing a new benchmark for efficient, high-accuracy image classification.

Enhancing image classification by combining CNNs, MLP-Mixers, and Vision TransformersExploring complementarity between architectures via ensemble learning techniquesImproving accuracy and latency in ImageNet classification with ensemble methods

Fine-tuning Aligned Classifiers for Merging Outputs: Towards a Superior Evaluation Protocol in Model Merging

Dec 18, 2024
FK
Fanshuang Kong
🏛️ Beihang University | Tongji University

In model merging, representational misalignment—modeled as an orthogonal transformation—exists between fused outputs and fine-tuned classifiers in the feature space, leading to evaluation distortion and suboptimal performance. To address this, we propose a novel few-shot unsupervised classifier alignment paradigm: using only a small number of unlabeled samples, it calibrates classifier weights via orthogonal transformation to achieve feature-space alignment. Based on this, we establish a more reliable evaluation protocol for merged models. Experiments across multiple classification tasks demonstrate that our method significantly improves the accuracy of merged models and yields evaluations that more faithfully reflect the intrinsic capabilities of merging methods. This work introduces a new benchmark for model merging that jointly ensures effectiveness and evaluability.

Address misalignment in model merging outputsEnhance classification via orthogonal transformationPropose FT-Classifier for superior evaluation

Dynamic Post-Hoc Neural Ensemblers

Oct 06, 2024
SP
Sebastian Pineda Arango
🏛️ University of Freiburg | ELLIS Institute Tübingen | University of Technology Nürnberg

Existing ensemble methods (e.g., greedy or random ensembling) employ static, sample-agnostic base-model weights, limiting expressive capacity and generalization. This paper proposes the Dynamic Posterior Neural Ensemble (DPNE), the first framework to introduce Random Prediction Dropout (RPD)—a theoretically grounded diversity regularization mechanism that provably enhances base-model diversity under a derived lower bound. DPNE employs a neural architecture to learn sample-wise adaptive weights end-to-end, eliminating restrictive pre-specified weight structures. Extensive experiments across CV, NLP, and tabular tasks demonstrate that DPNE significantly outperforms state-of-the-art ensemble baselines, effectively mitigating overfitting while improving both in-distribution accuracy and out-of-distribution robustness. The core contributions are: (1) dynamic posterior weight modeling conditioned on input samples, and (2) theoretically guaranteed diversity regularization via RPD.

Addressing limitations of constant-weight ensembling approachesEnhancing accuracy and robustness of ensemble methodsImproving generalization via regularized dynamic ensembling

This work addresses the lack of a unified Python framework for ensemble learning methods grounded in Composite Fusion Analysis (CFA), particularly in integrating Rank-Score Characteristic (RSC) functions with Cognitive Diversity (CD). To bridge this gap, we propose InFusionLayer—a general-purpose machine learning architecture inspired by CFA that, for the first time, unifies RSC and CD mechanisms within a single framework compatible with PyTorch, TensorFlow, and Scikit-learn. Requiring only a small set of base models, our approach achieves substantial performance gains in both unsupervised and supervised multi-class classification tasks. Extensive experiments across multiple computer vision benchmarks validate its efficacy, and the open-sourced implementation facilitates the practical adoption and broader dissemination of CFA within mainstream deep learning ecosystems.

Classifier generationCognitive diversityCombinatorial Fusion Analysis

Harnessing The Collective Wisdom: Fusion Learning Using Decision Sequences From Diverse Sources

Aug 21, 2023
TB
Trambak Banerjee
🏛️ University of Kansas | Fudan University | Yale University

Integrating hypothesis testing results across heterogeneous multi-source studies—some reporting only binary significance decisions, others only FDR control levels—poses a fundamental challenge for rigorous, unified FDR control. Method: We propose the Integrated Ranking and Thresholding (IRT) framework, which operates solely on binary rejection decisions, a prespecified global FDR level, and the set of hypotheses—requiring neither raw data, p-values, nor effect sizes. IRT employs nonparametric evidence aggregation and a ranking-driven thresholding mechanism, circumventing traditional meta-analysis assumptions of statistical homogeneity and reliance on shared summary statistics. Contribution/Results: IRT is the first method to achieve theoretically guaranteed strong FDR control under non-shared statistical summaries. We prove its FDR control property rigorously; simulations demonstrate superior performance over state-of-the-art integration methods; and real-world application to multi-center genome-wide association studies confirms its practical utility and robustness.

Combining findings across diverse data sourcesEnsuring overall false discovery rate controlFusing evidence from multiple testing procedures

Latest Papers

What's happening recently
View more

This work addresses the trade-off between computational cost and accuracy in neural network ensembles, where existing ensemble methods are computationally expensive while conventional weight aggregation techniques often sacrifice performance. To bridge this gap, the authors propose a “partial fusion” framework that formulates weight aggregation as a generalized pruning process. By leveraging partial optimal transport, the method matches and fuses the most similar neurons across models, allowing for neuron deletion, isolation, or linear combination. A similarity metric at the neuron level enables a controllable balance between computational overhead and model accuracy. Experiments demonstrate that partial fusion significantly reduces computational costs while preserving accuracy close to that of full ensembles; furthermore, its single-model variant outperforms traditional pruning approaches.

computational costmodel accuracyneural network ensembles

This work addresses the trade-off between domain expertise and robust reasoning in large language models: general-purpose models lack specialized knowledge, while domain-specific models often exhibit weak reasoning capabilities and poor robustness. To reconcile these limitations without additional training, the authors propose a novel inference-time capability fusion framework that reinterprets speculative decoding’s “draft–verify” mechanism as an adaptive routing strategy based on distributional divergence. Specifically, the method dynamically allocates control between a general and a domain-specific model by monitoring their output distribution discrepancies in real time using Jensen–Shannon divergence, enabling seamless collaboration. Evaluated on scientific benchmarks such as GPQA and ChemBench, the approach significantly outperforms individual models, demonstrating consistent performance gains on complex scientific reasoning tasks.

domain specializationinference-time collaborationlarge language models

This work proposes a biologically plausible multilayer neuronal network model that moves beyond the simplified weighted-sum neuron paradigm of conventional artificial neural networks, which struggles to capture the true learning mechanisms of biological systems. The proposed architecture employs cascaded adaptive combiners to enable efficient online, streaming learning without relying on backpropagation. By integrating neuron-level computations that more closely mirror biological reality with practical machine learning principles, the model establishes a concise and scalable learning framework. Empirical evaluation on image classification tasks demonstrates competitive performance, confirming both the efficacy and practicality of the approach.

alternative to backpropagationbiological neuron modelscascaded adaptive combiners

This work addresses the lack of systematic understanding in efficiently fusing large language models (LLMs) fine-tuned with lightweight adapters in multi-task learning, particularly regarding the trade-offs among ensembling, merging, and routing strategies. The study systematically evaluates three parameter-efficient fusion approaches—output ensembling, parameter averaging, and input-dependent routing—and demonstrates that non-uniform fusion consistently outperforms uniform methods, with routing yielding significant performance gains despite its higher computational cost. To reconcile this efficiency–performance trade-off, the authors propose a low-overhead expert selection mechanism that combines clustering with greedy subset selection, achieving near-optimal performance while substantially reducing computational overhead, thereby striking an effective balance between model efficacy and efficiency.

ensemblingmergingmulti-task learning

This work addresses the diminished understanding of neural network fundamentals caused by the widespread use of high-level deep learning libraries. To bridge this gap, the authors construct a complete neural network framework from scratch, eschewing automatic differentiation and prebuilt modules. The implementation explicitly details forward and backward propagation, incorporates multiple activation functions, L2 regularization, and advanced optimizers such as Adam. Designed to balance pedagogical clarity with engineering scalability, the framework demonstrates numerical stability, correctness, and generalization capability on multiclass classification tasks. It thus provides a reproducible and extensible tool for both research and instruction, fostering deeper insight into the core principles of deep learning.

deep learning librarieseducational gapfundamental understanding

Hot Scholars

MN

Mubashir Noman

MBZUAI
Image ProcessingObject Tracking / ClassificationRemote Sensing Change DetectionComputer
YL

Yunhui Liu

Nanjing University
Graph Machine Learning
GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
RS

Rusab Sarmun

University of Dhaka
Artificial IntelligenceDeep LearningComputer VisionReinforcement Learning
JG

Jihao Gu

University College London
Computer Vision