Score
Designs, builds, and evaluates predictive models and systems that assign discrete labels or categories to inputs based on features or learned representations. Implements and analyzes classifiers, decision boundaries, training pipelines, loss functions, and performance measures (accuracy, precision, recall, confusion matrices) for supervised labeling tasks.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
This paper addresses the fundamental trade-off between zero false negatives and low false positives in dynamic classification. To resolve this, we propose a lightweight multi-model collaborative framework. Methodologically, we introduce a novel self-supervised classification learning mechanism; dynamically partition input data into $N$ mutually exclusive subsets; train independent submodels for parallel prediction; and incorporate a confidence-threshold-based filtering and prediction rejection mechanism—eliminating unreliable predictions without requiring auxiliary verification models. Supervised feedback is further leveraged to iteratively refine model performance. Experiments demonstrate strict zero false negatives and a 37.2% reduction in false positive rate over state-of-the-art ensemble methods under low partitioning error; under high partitioning error, the framework maintains robustness comparable to current best models. Our core contribution is a reliability- and efficiency-aware lightweight paradigm for dynamic classification.
This study addresses the comparative performance evaluation of supervised machine learning models for multiclass text classification. We propose and implement an independent, reproducible text classification benchmarking platform that systematically evaluates artificial neural networks (ANNs), backpropagation networks (BPNs), and classical classifiers—including support vector machines (SVM), random forests, and naive Bayes—on a unified, manually annotated dataset. Experiments follow standardized preprocessing, k-fold cross-validation, and accuracy as the primary evaluation metric. Results reveal substantial performance variation across algorithms on real-world text data, with optimal model selection strongly dependent on dataset characteristics. The principal contribution is a scalable, open benchmarking framework for text classification; empirical findings demonstrate that ANNs consistently achieve superior classification accuracy across most test scenarios, thereby providing data-driven guidance for model selection in practical text classification tasks.
Conventional categorical encoding methods (e.g., one-hot) in industrial process modeling lack semantic expressiveness, failing to capture meaningful relationships among categories such as reactor types or operation sequences. Method: This paper introduces, for the first time, an NLP-inspired semantic-aware categorical embedding framework: category-specific semantic vectors are generated via pre-trained language models and subsequently projected into an interpretable low-dimensional space using PCA or UMAP. Contribution/Results: Unlike conventional encodings, the proposed method explicitly models semantic distances between categories and enables quantitative feature importance analysis. Evaluated on an industrial case study involving cutting tool coatings, it achieves significant predictive performance gains. Moreover, it natively supports heterogeneous inputs—integrating both categorical and numerical features—thereby overcoming a fundamental limitation of existing encoding paradigms that cannot represent categorical similarity.
Existing machine learning frameworks suffer from insufficient formalization of objective functions and lack a unified, cross-domain behavioral design paradigm. Method: We propose an equation-constrained compositional function modeling approach for learners, constructing task graphs and compositional semantic graphs to enable model-agnostic behavioral specification and optimization. We introduce a novel task-oriented pattern language framework and the “manipulator” task paradigm, supporting end-to-end, architecture-agnostic, and adversarial-training-free minimal editing of data attributes. Contribution/Results: Theoretically, our work integrates formal methods and theoretical computer science principles. Empirically, we demonstrate precise, controllable, and interpretable behavioral editing on small-scale models under stable training—without stochastic sampling or data intervention—yielding significant improvements in deployment efficiency and formal verifiability.
This work addresses the challenge of error propagation in intermediate steps of multi-step reasoning, which undermines the reliability of large language models. To mitigate this issue, the authors propose a fine-grained process supervision method grounded in information theory. Their approach introduces a novel technique that leverages information gain to automatically generate step-level labels, employing Monte Carlo sampling to efficiently estimate net information gain. By integrating these labels into a process reward model, the method enhances chain-of-thought reasoning selection. Notably, the label generation complexity is reduced from O(N log N) to O(N), substantially improving scalability. Empirical evaluations across diverse domains—including mathematics, Python programming, SQL, and scientific question answering—demonstrate consistent gains in both accuracy and efficiency of best-of-K reasoning.
To address the insufficient calibration of machine learning models in high-stakes decision-making (e.g., clinical prediction) and the common trade-off between calibration and discriminative performance, this paper proposes a representation-clustering-based ensemble calibration framework. First, it learns sample representations to capture latent-space structure, then performs adaptive clustering to partition heterogeneous subpopulations. Dedicated calibration functions are constructed per cluster, optimized jointly to minimize calibration error while preserving discriminative power. The framework is model-agnostic—compatible with diverse base models and evaluation metrics—and requires no modification to original model architectures. Experiments show it improves calibration accuracy of state-of-the-art calibration methods from 82.28% to 100%, significantly enhancing predictive reliability, while maintaining or even improving discriminative metrics such as AUC. Its core innovation lies in aligning calibration granularity with intrinsic representation structure, and synergistically boosting both calibration and discrimination via multi-calibrator ensembling and joint optimization.
Existing LLM-driven feature engineering methods are not designed for multi-label learning, thus failing to model label dependencies and lacking task-specificity. To address this, we propose FEAML—a novel framework that pioneers the integration of LLM-based code generation into multi-label settings. FEAML automatically constructs highly discriminative features by jointly leveraging metadata and label co-occurrence matrices. It introduces label-dependency-aware prompt engineering and a Pearson correlation-based redundancy detection mechanism, coupled with closed-loop optimization guided by classification accuracy. This yields an interpretable, low-redundancy, and self-optimizing feature generation paradigm. Extensive experiments on multiple standard multi-label benchmark datasets demonstrate that FEAML significantly outperforms conventional feature engineering approaches, achieving substantial average improvements in classification accuracy—thereby validating its effectiveness and generalizability.
To address the degradation of prediction probability calibration in deployed image classification models due to concept drift, this paper proposes an online calibration monitoring method that requires no access to model internals—only predicted probabilities and ground-truth labels. Our approach introduces, for the first time, a Cumulative Sum (CUSUM) control chart with dynamic control limits into calibration monitoring. It computes cumulative deviations of calibration error over time and adaptively adjusts detection thresholds to enable early warning of calibration loss. Compared to static-threshold methods, our framework significantly enhances sensitivity to temporal distribution shifts and accelerates response to emerging miscalibration. We validate its effectiveness and robustness across multiple image classification benchmarks under diverse concept drift scenarios. The proposed method establishes a scalable, black-box-compatible paradigm for trustworthy model deployment, enabling continuous, lightweight calibration assessment without architectural or training modifications.
This study addresses the limitations of existing approaches for automatically extracting machine learning (ML) pipeline structures, which often rely on manual annotations or suffer from insufficient generalization to keep pace with the rapid evolution of the ML ecosystem. The work presents the first systematic evaluation of small language models (SLMs) for reverse-engineering ML pipelines and proposes an SLM-based method for their automatic identification and reconstruction. Through comprehensive comparative experiments across multiple SLMs and rigorous statistical validation using Cochran’s Q, McNemar, and Pearson’s chi-squared tests, the authors demonstrate that the best-performing SLM significantly outperforms current methods and exhibits robustness across diverse classification schemes. This approach uncovers finer-grained patterns in data science practices and overcomes longstanding bottlenecks in scalability and domain adaptability inherent in traditional techniques.