Score
Designs and builds predictive models and end-to-end predictive analytics pipelines that use supervised learning methods (classification, regression) and time-series/autoregressive techniques to forecast outcomes. Implements, evaluates, and operationalizes machine learning models (including Excel or Python implementations) for tasks such as ETA prediction, predictive maintenance, market analytics, and cost modeling, covering model selection, validation, and deployment.
This study addresses the challenge of degraded data quality—specifically outliers and missing values—in long-horizon time series forecasting, which critically undermines model robustness. We establish a unified evaluation framework to systematically benchmark mainstream models—including LSTM, Prophet, XGBoost, and Random Forest—under three realistic data conditions: complete, noisy (outlier-contaminated), and incomplete (missing-value) sequences, with ARIMA as the baseline. Methodologically, we employ sliding-window modeling, multi-step rolling prediction, and adaptive imputation for preprocessing. Our key contributions include: (i) a novel, interpretable algorithm selection guideline grounded in data characteristics and forecasting requirements; and (ii) empirical findings demonstrating that XGBoost reduces average MAE by 23% under noise, while Prophet exhibits superior stability for long-term trend forecasting. The results provide reproducible, principled guidance for industrial-scale time series modeling.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.
In scientific computing applications—such as trajectory prediction, optimal control, and minimum energy path computation—downstream algorithms critically depend on accurate model evaluations. Conventional mean-squared-error-based supervised learning often induces task-specific performance degradation due to misalignment between the loss function and the ultimate algorithmic objective. Method: We propose a task-oriented predictive modeling paradigm that replaces standard regression losses with a surrogate objective: the maximum prediction error over a downstream task support set. Our framework integrates sampling measure modeling, empirical risk discretization, and iterative optimization to directly optimize downstream algorithmic performance. Contribution/Results: This is the first approach to explicitly embed downstream robustness requirements into the training objective. Evaluated across multiple scientific computing benchmarks, it consistently improves both predictive accuracy and algorithmic stability, demonstrating superior generalization under task-relevant perturbations.
Business process optimization remains challenging due to fragmented methodologies across process mining, predictive process monitoring, and process-aware recommendation—each operating in isolation without a unified theoretical foundation or integration framework. Method: This paper proposes a closed-loop optimization framework that systematically integrates Alpha algorithm/Inductive Miner for process discovery, LSTM/Transformer for runtime prediction, collaborative filtering/graph neural networks for action recommendation, and explainable AI (XAI) for interpretability—enabling automated bottleneck identification, anomaly forecasting, and prescriptive optimization from event logs. Contribution/Results: We establish the first unified conceptual boundary, evolutionary taxonomy, and synergy paradigm across the three domains; construct a comprehensive classification schema covering 120+ studies; clarify application scopes and standardized evaluation benchmarks; and deliver an industrially actionable methodology selection guide with validated deployment pathways.
To address the challenge of detecting Advanced Persistent Threats (APTs) that evade traditional rule-based engines, this paper proposes a lightweight, interpretable predictive analytics framework integrating logistic regression and K-means clustering. Designed for low-resource settings with small-scale security event data (Kaggle dataset, *n* = 2,000), it enables real-time threat detection and response. Methodologically, it is the first to synergistically combine these two models in resource-constrained environments and employs SPSS-based statistical tests to validate feature significance. Compared to baseline rule engines, the framework achieves significantly improved threat alert sensitivity (+23.6%) and reduces average response time by 41%, while preserving high model interpretability. It thus delivers actionable, proactive defense decision support for Security Operations Centers (SOCs).
This study addresses key challenges in business and financial forecasting—namely poor reproducibility, low model transparency, and weak cross-environment consistency—by systematically evaluating Meta’s open-source Prophet framework. Under a unified experimental design, Prophet is benchmarked against multiple ARIMA variants and Random Forest models, leveraging its additive structure, standardized workflow, and open implementation. The findings demonstrate that Prophet achieves competitive predictive performance while substantially enhancing reproducibility, auditability, and engineering integration efficiency. Rather than introducing a novel algorithm, this work advocates for Prophet as a transparent, reliable, and collaboration-friendly forecasting methodology, particularly suited for high-stakes decision-making contexts where interpretability and robustness are paramount.
This work addresses the cumbersome and error-prone process of manually extracting samples and labels from relational databases for traditional machine learning modeling. To streamline this workflow, the authors propose PQL, a declarative domain-specific language inspired by SQL that enables users to define diverse predictive tasks—including regression, classification, time-series forecasting, and recommendation—through a single query directly over relational databases, with training labels automatically generated. PQL offers two implementations: one optimized for low-latency, small-scale scenarios and another designed for large-scale data processing. The approach has been validated in real-world applications such as financial fraud detection, product recommendation, and load forecasting, demonstrating its versatility, efficiency, and significant improvements in modeling productivity and scalability.
This study addresses the challenges of modeling and analyzing complex time series arising in astrophysics, meteorology, finance, and other domains by systematically integrating classical statistical methods—such as ARIMA, exponential smoothing, and state-space models—with modern machine learning techniques, including tree-based ensembles, hidden Markov models, Gaussian processes, and deep learning architectures like RNNs, CNNs, and Transformers. By distilling cross-disciplinary modeling principles, the work establishes a unified framework that combines theoretical rigor with practical guidance, offering researchers a comprehensive and extensible toolkit for time series analysis. This approach significantly enhances the capacity to handle temporal data across diverse scientific and applied contexts.
This work proposes the first end-to-end automated artificial intelligence research framework capable of fully automating the development pipeline from algorithmic idea generation to executable machine learning classifiers. The approach integrates structured meta-prompt engineering with large language model–based code generation, augmented by an automated evaluation and iterative optimization mechanism. Experimental results on twenty standard datasets from the Infinity-Bench benchmark demonstrate that multiple novel classifiers autonomously generated by the framework significantly outperform baseline methods implemented in scikit-learn. This study thus achieves, for the first time, complete automation of the entire workflow—from initial algorithmic conception to deployable, runnable code—marking a significant step toward self-driving AI research systems.
Inaccurate aviation fuel demand forecasting hampers supply chain optimization, as existing approaches rely heavily on expert judgment or deterministic models and lack data-driven, long-horizon predictive capabilities. To address this, we propose a hybrid data-driven framework for Copenhagen Airport that integrates multiple exogenous variables and—novelty—synergistically combines an LSTM-based sequence-to-sequence architecture with Facebook Prophet for 30-day rolling forecasts. Evaluated on real-world market data, the hybrid model achieves a 18.7% reduction in MAPE over standalone models and conventional methods, demonstrating superior robustness during periods of high demand volatility. This work empirically validates the efficacy of integrating deep learning with classical time-series modeling for long-term aviation fuel forecasting. Moreover, it delivers a production-ready decision-support tool for fuel distributors, advancing the intelligent,精细化 (fine-grained) transformation of aviation energy supply chains.