combinatorial fusion analysis

Designs and implements search and evaluation procedures that enumerate or heuristically explore combinations of models and fusion rules to identify optimal ensemble subsets and fusion configurations (including rank-score fusion and other combinatorial selection methods). Builds validation workflows that evaluate candidate fusions on held-out data and quantify performance gains and uncertainty (for example with held-out metrics and bootstrap confidence intervals) to select the best fusion configuration.

combinatorialfusionanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.

binary classificationcomputational costearly stopping

An Integrated Fusion Framework for Ensemble Learning Leveraging Gradient-Boosting and Fuzzy Rule-Based Models

Nov 01, 2024
JL
Jinbo Li
🏛️ China Unicom | Henan University | University of Macau | University of Alberta | Nantong University

Fuzzy rule models offer strong interpretability but suffer from poor scalability and susceptibility to overfitting in complex tasks and large-scale data scenarios. To address these limitations, this paper proposes a novel ensemble framework integrating gradient boosting with fuzzy rule-based base learners. We introduce a dynamic control factor that adaptively adjusts the weights of fuzzy base models in each boosting iteration, simultaneously serving as a regularizer and performance optimizer. Additionally, we design a validation-set-driven, sample-level correction mechanism to enhance generalization and ensemble diversity. Experimental results demonstrate that our approach significantly mitigates overfitting, reduces rule complexity (e.g., fewer rules and shorter antecedents), and preserves high model interpretability and maintainability. The method thus provides a practical pathway for deploying interpretable AI in complex industrial applications.

Addresses overfitting and complexity in fuzzy rule-based systems using dynamic controlCombines gradient boosting with fuzzy models to enhance performance and interpretabilityImproves scalability and maintenance of ensemble models through adaptive tuning mechanisms

This work addresses the lack of a unified Python framework for ensemble learning methods grounded in Composite Fusion Analysis (CFA), particularly in integrating Rank-Score Characteristic (RSC) functions with Cognitive Diversity (CD). To bridge this gap, we propose InFusionLayer—a general-purpose machine learning architecture inspired by CFA that, for the first time, unifies RSC and CD mechanisms within a single framework compatible with PyTorch, TensorFlow, and Scikit-learn. Requiring only a small set of base models, our approach achieves substantial performance gains in both unsupervised and supervised multi-class classification tasks. Extensive experiments across multiple computer vision benchmarks validate its efficacy, and the open-sourced implementation facilitates the practical adoption and broader dissemination of CFA within mainstream deep learning ecosystems.

Classifier generationCognitive diversityCombinatorial Fusion Analysis

Green AI in Action: Strategic Model Selection for Ensembles in Production

May 21, 2024
NN
Nienke Nijkamp
🏛️ Delft University of Technology | Deloitte Risk Advisory BV | University of Amsterdam

To address the energy–accuracy trade-off arising from multi-model parallel inference in real-time AI integration systems, this paper proposes the first energy-efficiency–oriented ensemble model selection framework. Methodologically, it integrates static and dynamic model selection strategies, incorporating energy-aware inference scheduling and a multi-objective accuracy–energy optimization mechanism. The static strategy reduces energy consumption to 62% of the baseline while improving F1 score; the dynamic strategy achieves 76% energy consumption with further F1 gain. With the accuracy–energy trade-off mechanism, energy drops dramatically to 14% and 57%, respectively, with negligible accuracy loss. Our key contribution lies in the first deep embedding of energy-efficiency modeling into the ensemble inference pipeline, enabling joint optimization of computational resources and predictive performance. The framework delivers substantial energy savings and robust accuracy guarantees in production-grade AI systems.

Balancing AI model accuracy with energy consumption in ensemblesOptimizing model selection strategies for energy-efficient AI systemsReducing energy usage in ensemble learning without sacrificing accuracy

Harnessing The Collective Wisdom: Fusion Learning Using Decision Sequences From Diverse Sources

Aug 21, 2023
TB
Trambak Banerjee
🏛️ University of Kansas | Fudan University | Yale University

Integrating hypothesis testing results across heterogeneous multi-source studies—some reporting only binary significance decisions, others only FDR control levels—poses a fundamental challenge for rigorous, unified FDR control. Method: We propose the Integrated Ranking and Thresholding (IRT) framework, which operates solely on binary rejection decisions, a prespecified global FDR level, and the set of hypotheses—requiring neither raw data, p-values, nor effect sizes. IRT employs nonparametric evidence aggregation and a ranking-driven thresholding mechanism, circumventing traditional meta-analysis assumptions of statistical homogeneity and reliance on shared summary statistics. Contribution/Results: IRT is the first method to achieve theoretically guaranteed strong FDR control under non-shared statistical summaries. We prove its FDR control property rigorously; simulations demonstrate superior performance over state-of-the-art integration methods; and real-world application to multi-center genome-wide association studies confirms its practical utility and robustness.

Combining findings across diverse data sourcesEnsuring overall false discovery rate controlFusing evidence from multiple testing procedures

Latest Papers

What's happening recently
View more

This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.

benchmarkingdataset selectionmodel ranking

This study addresses the underexplored engineering challenges in existing machine learning evaluation frameworks, where operational issues and their root causes have lacked systematic investigation. To bridge this gap, the work formally establishes evaluation engineering as a distinct research direction within software engineering. Through an empirical analysis of 57 frameworks and a comprehensive categorization of 16,560 reported issues across a newly proposed five-stage workflow model, the study reveals that 41.4% of problems originate in the specification phase, while 61.7% of classified issues stem from missing functionality, inadequate documentation, and insufficient input validation. The findings yield a structured taxonomy of evaluation-related problems and provide empirical evidence to inform the design and improvement of robust evaluation systems.

empirical studyevaluation harnessesmachine learning

This study addresses the performance limitations in credit card fraud detection caused by the scarcity, high cost, and imbalanced distribution of fraudulent samples. To overcome these challenges, the authors propose a Combinatorial Fusion Analysis (CFA) framework during validation that searches for an optimal subset among seven base classifiers—including Random Forest, XGBoost, and LightGBM—and integrates them via a Diversity-Enhanced Fusion Weighted Scoring (DEF WtScore) strategy. By treating CFA as a model selection and weighting mechanism rather than full ensemble fusion, the approach effectively mitigates information leakage. Evaluation employs an unbiased 60/20/20 data split and Bootstrap confidence intervals. On the IEEE-CIS dataset, the method achieves an AUC-ROC of 0.9405, AUPRC of 0.6699, and F1-score of 0.6373, significantly outperforming the best individual model, soft voting, and stacking ensembles.

Combinatorial Fusion Analysiscredit-card fraud detectionimbalanced data

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Hot Scholars

EY

Eric Yu

Professor, University of Toronto
information systems engineeringconceptual modelingsoftware engineeringrequirements engineering
HL

Hui Li

Jiangnan University
Informtion fusionmulti-modal processing
HZ

Hector Zenil

Associate Professor @ King’s College London & Researcher @ The Francis Crick Institute
algorithmic information dynamicscausalityalgorithmic probabilitymachine intelligence
YX

Yang Xu

Fudan University, School of Computer Science
Software Defined NetworksData Center NetworksDistributed Machine Learning
RD

Ross D. King

Professor of Machine Intelligence, Chalmers University
Automation of ScienceDrug DesignArtificial IntelligenceMachine Learning