Score
Designs and implements search and evaluation procedures that enumerate or heuristically explore combinations of models and fusion rules to identify optimal ensemble subsets and fusion configurations (including rank-score fusion and other combinatorial selection methods). Builds validation workflows that evaluate candidate fusions on held-out data and quantify performance gains and uncertainty (for example with held-out metrics and bootstrap confidence intervals) to select the best fusion configuration.
This work proposes a sequential testing–based early-stopping strategy for binary ensemble classifiers to reduce inference overhead while strictly bounding the divergence rate from predictions of the full ensemble. The approach terminates evaluation as soon as a decisive majority emerges during the sequential assessment of base models. Under three optimality criteria, the strategy can be formulated as a linear programming problem, enabling efficient computation of the optimal stopping rule. Experimental results on UCI and Grinsztajn benchmark datasets demonstrate that the method achieves an average speedup exceeding 4× while consistently maintaining prediction divergence below 0.1%.
Fuzzy rule models offer strong interpretability but suffer from poor scalability and susceptibility to overfitting in complex tasks and large-scale data scenarios. To address these limitations, this paper proposes a novel ensemble framework integrating gradient boosting with fuzzy rule-based base learners. We introduce a dynamic control factor that adaptively adjusts the weights of fuzzy base models in each boosting iteration, simultaneously serving as a regularizer and performance optimizer. Additionally, we design a validation-set-driven, sample-level correction mechanism to enhance generalization and ensemble diversity. Experimental results demonstrate that our approach significantly mitigates overfitting, reduces rule complexity (e.g., fewer rules and shorter antecedents), and preserves high model interpretability and maintainability. The method thus provides a practical pathway for deploying interpretable AI in complex industrial applications.
This work addresses the lack of a unified Python framework for ensemble learning methods grounded in Composite Fusion Analysis (CFA), particularly in integrating Rank-Score Characteristic (RSC) functions with Cognitive Diversity (CD). To bridge this gap, we propose InFusionLayer—a general-purpose machine learning architecture inspired by CFA that, for the first time, unifies RSC and CD mechanisms within a single framework compatible with PyTorch, TensorFlow, and Scikit-learn. Requiring only a small set of base models, our approach achieves substantial performance gains in both unsupervised and supervised multi-class classification tasks. Extensive experiments across multiple computer vision benchmarks validate its efficacy, and the open-sourced implementation facilitates the practical adoption and broader dissemination of CFA within mainstream deep learning ecosystems.
To address the energy–accuracy trade-off arising from multi-model parallel inference in real-time AI integration systems, this paper proposes the first energy-efficiency–oriented ensemble model selection framework. Methodologically, it integrates static and dynamic model selection strategies, incorporating energy-aware inference scheduling and a multi-objective accuracy–energy optimization mechanism. The static strategy reduces energy consumption to 62% of the baseline while improving F1 score; the dynamic strategy achieves 76% energy consumption with further F1 gain. With the accuracy–energy trade-off mechanism, energy drops dramatically to 14% and 57%, respectively, with negligible accuracy loss. Our key contribution lies in the first deep embedding of energy-efficiency modeling into the ensemble inference pipeline, enabling joint optimization of computational resources and predictive performance. The framework delivers substantial energy savings and robust accuracy guarantees in production-grade AI systems.
Integrating hypothesis testing results across heterogeneous multi-source studies—some reporting only binary significance decisions, others only FDR control levels—poses a fundamental challenge for rigorous, unified FDR control. Method: We propose the Integrated Ranking and Thresholding (IRT) framework, which operates solely on binary rejection decisions, a prespecified global FDR level, and the set of hypotheses—requiring neither raw data, p-values, nor effect sizes. IRT employs nonparametric evidence aggregation and a ranking-driven thresholding mechanism, circumventing traditional meta-analysis assumptions of statistical homogeneity and reliance on shared summary statistics. Contribution/Results: IRT is the first method to achieve theoretically guaranteed strong FDR control under non-shared statistical summaries. We prove its FDR control property rigorously; simulations demonstrate superior performance over state-of-the-art integration methods; and real-world application to multi-center genome-wide association studies confirms its practical utility and robustness.
本文提出一种基于置信区域的筛选框架,用于解决模拟系统可接受性问题,保证高概率筛选出所有或每个可接受系统,并支持并行化。
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.
This study addresses the underexplored engineering challenges in existing machine learning evaluation frameworks, where operational issues and their root causes have lacked systematic investigation. To bridge this gap, the work formally establishes evaluation engineering as a distinct research direction within software engineering. Through an empirical analysis of 57 frameworks and a comprehensive categorization of 16,560 reported issues across a newly proposed five-stage workflow model, the study reveals that 41.4% of problems originate in the specification phase, while 61.7% of classified issues stem from missing functionality, inadequate documentation, and insufficient input validation. The findings yield a structured taxonomy of evaluation-related problems and provide empirical evidence to inform the design and improvement of robust evaluation systems.
This study addresses the performance limitations in credit card fraud detection caused by the scarcity, high cost, and imbalanced distribution of fraudulent samples. To overcome these challenges, the authors propose a Combinatorial Fusion Analysis (CFA) framework during validation that searches for an optimal subset among seven base classifiers—including Random Forest, XGBoost, and LightGBM—and integrates them via a Diversity-Enhanced Fusion Weighted Scoring (DEF WtScore) strategy. By treating CFA as a model selection and weighting mechanism rather than full ensemble fusion, the approach effectively mitigates information leakage. Evaluation employs an unbiased 60/20/20 data split and Bootstrap confidence intervals. On the IEEE-CIS dataset, the method achieves an AUC-ROC of 0.9405, AUPRC of 0.6699, and F1-score of 0.6373, significantly outperforming the best individual model, soft voting, and stacking ensembles.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.