build predictive models

Designs and builds predictive models and end-to-end predictive analytics pipelines that use supervised learning methods (classification, regression) and time-series/autoregressive techniques to forecast outcomes. Implements, evaluates, and operationalizes machine learning models (including Excel or Python implementations) for tasks such as ETA prediction, predictive maintenance, market analytics, and cost modeling, covering model selection, validation, and deployment.

buildpredictivemodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$204K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Addressing Challenges in Time Series Forecasting: A Comprehensive Comparison of Machine Learning Techniques

Mar 26, 2025
SA
Seyedeh Azadeh Fallah Mortezanejad
🏛️ Jiangsu University

This study addresses the challenge of degraded data quality—specifically outliers and missing values—in long-horizon time series forecasting, which critically undermines model robustness. We establish a unified evaluation framework to systematically benchmark mainstream models—including LSTM, Prophet, XGBoost, and Random Forest—under three realistic data conditions: complete, noisy (outlier-contaminated), and incomplete (missing-value) sequences, with ARIMA as the baseline. Methodologically, we employ sliding-window modeling, multi-step rolling prediction, and adaptive imputation for preprocessing. Our key contributions include: (i) a novel, interpretable algorithm selection guideline grounded in data characteristics and forecasting requirements; and (ii) empirical findings demonstrating that XGBoost reduces average MAE by 23% under noise, while Prophet exhibits superior stability for long-term trend forecasting. The results provide reproducible, principled guidance for industrial-scale time series modeling.

Comparing ML techniques for time series forecasting accuracyEvaluating algorithms on complete, outlier, and missing value datasetsSelecting optimal TS forecasting method based on data characteristics

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

Learning task-specific predictive models for scientific computing

Jun 04, 2025
JY
Jianyuan Yin
🏛️ National University of Singapore

In scientific computing applications—such as trajectory prediction, optimal control, and minimum energy path computation—downstream algorithms critically depend on accurate model evaluations. Conventional mean-squared-error-based supervised learning often induces task-specific performance degradation due to misalignment between the loss function and the ultimate algorithmic objective. Method: We propose a task-oriented predictive modeling paradigm that replaces standard regression losses with a surrogate objective: the maximum prediction error over a downstream task support set. Our framework integrates sampling measure modeling, empirical risk discretization, and iterative optimization to directly optimize downstream algorithmic performance. Contribution/Results: This is the first approach to explicitly embed downstream robustness requirements into the training objective. Evaluated across multiple scientific computing benchmarks, it consistently improves both predictive accuracy and algorithmic stability, demonstrating superior generalization under task-relevant perturbations.

Addressing limitations of mean square error in task-specific learningDeveloping iterative algorithms for task-specific supervised learningLearning predictive models for non-prediction scientific tasks

Business process optimization remains challenging due to fragmented methodologies across process mining, predictive process monitoring, and process-aware recommendation—each operating in isolation without a unified theoretical foundation or integration framework. Method: This paper proposes a closed-loop optimization framework that systematically integrates Alpha algorithm/Inductive Miner for process discovery, LSTM/Transformer for runtime prediction, collaborative filtering/graph neural networks for action recommendation, and explainable AI (XAI) for interpretability—enabling automated bottleneck identification, anomaly forecasting, and prescriptive optimization from event logs. Contribution/Results: We establish the first unified conceptual boundary, evolutionary taxonomy, and synergy paradigm across the three domains; construct a comprehensive classification schema covering 120+ studies; clarify application scopes and standardized evaluation benchmarks; and deliver an industrially actionable methodology selection guide with validated deployment pathways.

Optimize business process performancePredict future process behaviorSupport data-driven decision-making

To address the challenge of detecting Advanced Persistent Threats (APTs) that evade traditional rule-based engines, this paper proposes a lightweight, interpretable predictive analytics framework integrating logistic regression and K-means clustering. Designed for low-resource settings with small-scale security event data (Kaggle dataset, *n* = 2,000), it enables real-time threat detection and response. Methodologically, it is the first to synergistically combine these two models in resource-constrained environments and employs SPSS-based statistical tests to validate feature significance. Compared to baseline rule engines, the framework achieves significantly improved threat alert sensitivity (+23.6%) and reduces average response time by 41%, while preserving high model interpretability. It thus delivers actionable, proactive defense decision support for Security Operations Centers (SOCs).

Evaluating key network features for accurate threat classificationIncorporating contextual features to improve early threat detectionReal-time cyber-attack detection using predictive analytics methods

Latest Papers

What's happening recently
View more

This study addresses key challenges in business and financial forecasting—namely poor reproducibility, low model transparency, and weak cross-environment consistency—by systematically evaluating Meta’s open-source Prophet framework. Under a unified experimental design, Prophet is benchmarked against multiple ARIMA variants and Random Forest models, leveraging its additive structure, standardized workflow, and open implementation. The findings demonstrate that Prophet achieves competitive predictive performance while substantially enhancing reproducibility, auditability, and engineering integration efficiency. Rather than introducing a novel algorithm, this work advocates for Prophet as a transparent, reliable, and collaboration-friendly forecasting methodology, particularly suited for high-stakes decision-making contexts where interpretability and robustness are paramount.

business analyticsfinancial analyticsforecasting

This work addresses the cumbersome and error-prone process of manually extracting samples and labels from relational databases for traditional machine learning modeling. To streamline this workflow, the authors propose PQL, a declarative domain-specific language inspired by SQL that enables users to define diverse predictive tasks—including regression, classification, time-series forecasting, and recommendation—through a single query directly over relational databases, with training labels automatically generated. PQL offers two implementations: one optimized for low-latency, small-scale scenarios and another designed for large-scale data processing. The approach has been validated in real-world applications such as financial fraud detection, product recommendation, and load forecasting, demonstrating its versatility, efficiency, and significant improvements in modeling productivity and scalability.

declarative languagemachine learningpredictive modeling

This study addresses the challenges of modeling and analyzing complex time series arising in astrophysics, meteorology, finance, and other domains by systematically integrating classical statistical methods—such as ARIMA, exponential smoothing, and state-space models—with modern machine learning techniques, including tree-based ensembles, hidden Markov models, Gaussian processes, and deep learning architectures like RNNs, CNNs, and Transformers. By distilling cross-disciplinary modeling principles, the work establishes a unified framework that combines theoretical rigor with practical guidance, offering researchers a comprehensive and extensible toolkit for time series analysis. This approach significantly enhances the capacity to handle temporal data across diverse scientific and applied contexts.

ForecastingMachine LearningStatistical Modeling

This work proposes the first end-to-end automated artificial intelligence research framework capable of fully automating the development pipeline from algorithmic idea generation to executable machine learning classifiers. The approach integrates structured meta-prompt engineering with large language model–based code generation, augmented by an automated evaluation and iterative optimization mechanism. Experimental results on twenty standard datasets from the Infinity-Bench benchmark demonstrate that multiple novel classifiers autonomously generated by the framework significantly outperform baseline methods implemented in scikit-learn. This study thus achieves, for the first time, complete automation of the entire workflow—from initial algorithmic conception to deployable, runnable code—marking a significant step toward self-driving AI research systems.

AI automationautomate AI researchend-to-end framework

Data-driven jet fuel demand forecasting: A case study of Copenhagen Airport

Nov 04, 2025
AC
Alessandro Contini
🏛️ Technical University of Denmark

Inaccurate aviation fuel demand forecasting hampers supply chain optimization, as existing approaches rely heavily on expert judgment or deterministic models and lack data-driven, long-horizon predictive capabilities. To address this, we propose a hybrid data-driven framework for Copenhagen Airport that integrates multiple exogenous variables and—novelty—synergistically combines an LSTM-based sequence-to-sequence architecture with Facebook Prophet for 30-day rolling forecasts. Evaluated on real-world market data, the hybrid model achieves a 18.7% reduction in MAPE over standalone models and conventional methods, demonstrating superior robustness during periods of high demand volatility. This work empirically validates the efficacy of integrating deep learning with classical time-series modeling for long-term aviation fuel forecasting. Moreover, it delivers a production-ready decision-support tool for fuel distributors, advancing the intelligent,精细化 (fine-grained) transformation of aviation energy supply chains.

Addressing the lack of machine learning studies for fuel predictionEvaluating data-driven models for 30-day fuel demand horizonForecasting jet fuel demand to optimize aviation supply chains

Hot Scholars

EF

Emilio Ferrara

Professor of Computer Science at the University of Southern California
Human-Centered AISocial ComputingNetwork ScienceAI Safety
SB

Stella Biderman

EleutherAI
Natural Language ProcessingArtificial IntelligenceLanguage ModelingDeep Learning
PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering
MS

Maarten Sap

Carnegie Mellon University
Natural Language ProcessingArtificial IntelligenceCommonsense ReasoningEthics in AI
CB

Conrad Borchers

Carnegie Mellon University
Educational Data MiningLearning AnalyticsIntelligent Tutoring SystemsSelf-Regulated Learning