data mining

Designs and implements algorithms, pipelines, and software to extract patterns, associations, clusters, anomalies, summaries, and predictive models from large or complex datasets. Builds and evaluates preprocessing, feature engineering, statistical and machine‑learning methods, pattern‑recognition and visualization components to validate, interpret, and operationalize the discovered knowledge.

datamining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$187K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Beyond algorithm hyperparameters: on preprocessing hyperparameters and associated pitfalls in machine learning applications

Dec 04, 2024
CS
Christina Sauer
🏛️ LMU Munich | Munich Center for Machine Learning | Medical University of Vienna

This paper identifies a systemic issue in machine learning: preprocessing hyperparameters—such as missing-value imputation strategies—are frequently overlooked yet substantially bias model evaluation. Current practice often involves informal, post-hoc tuning of preprocessing steps, leading to optimistic performance estimates and irreproducible results. To address this, the authors formally distinguish and empirically analyze the coupling effects between algorithmic and preprocessing hyperparameters. Using a modular supervised learning workflow model, controlled variable experiments, replication of canonical case studies, and bias diagnostics, they quantify the resulting optimistic bias. Key contributions include: (1) establishing preprocessing hyperparameters as equally critical as algorithmic ones; (2) proposing formal modeling principles to eliminate informal preprocessing tuning; and (3) delivering actionable reporting guidelines for ML practitioners, thereby significantly enhancing model credibility and reproducibility.

Addresses overlooked preprocessing hyperparameters in ML model tuningAims to improve predictive modeling quality and reportingHighlights pitfalls in informal preprocessing optimization practices

MLPrE -- A tool for preprocessing and exploratory data analysis prior to machine learning model construction

Oct 29, 2025
DS
David S Maxwell
🏛️ The University of Texas MD Anderson Cancer Center

To address poor scalability, integration complexity, and inflexible configuration in preprocessing multi-source heterogeneous data for machine learning modeling and graph database construction, this paper proposes a lightweight, modular, JSON-driven automated data preprocessing framework. Built upon Spark DataFrames for efficient distributed processing, the framework defines 69 composable and parallelizable processing stages spanning input parsing, filtering, statistical analysis, feature engineering, and graph-structure transformation. Its declarative JSON-based configuration enables dynamic adaptation to varying data types and scales, significantly enhancing interoperability with workflow orchestration systems such as Apache Airflow. Experimental evaluation across six heterogeneous datasets demonstrates the framework’s generality and scalability: it successfully supports wine quality clustering analysis and end-to-end conversion of phosphosylation site–kinase interaction data into a graph database.

Addressing scalability limitations in existing data processing workflowsPreprocessing diverse data formats for machine learning modelsProviding exploratory analysis capabilities for large-scale datasets

This study addresses the lack of rigorous statistical assessment for the reliability of output structures in complex clustering pipelines that involve multiple data-dependent stages such as anomaly detection, feature selection, and clustering. To bridge this gap, the work systematically applies selective inference to the entire clustering analysis workflow, establishing a statistical framework that enables valid significance testing of final cluster assignments. The proposed method rigorously controls the type I error rate at any pre-specified nominal level and demonstrates strong empirical performance on both synthetic and real-world datasets. By doing so, it provides a principled and reliable foundation for statistical inference in multi-stage, data-driven clustering procedures.

clustering pipelinesdata analysis pipelineselective inference

ExeKGLib: A Platform for Machine Learning Analytics based on Knowledge Graphs

Aug 01, 2025
AK
Antonis Klironomos
🏛️ Bosch Center for AI | Oslo Metropolitan University | RWTH Aachen | University of Mannheim | University of Oslo

To address the challenge that domain experts—lacking machine learning (ML) expertise—struggle to construct high-quality analytical pipelines, this paper proposes a knowledge graph–based low-code ML platform. Methodologically, it encodes ML best practices, algorithmic constraints, and domain semantics into a structured, inferable knowledge graph, enabling visual pipeline orchestration and automated execution via a graphical user interface. Technically, the platform integrates the Python ecosystem, modern GUI frameworks, and a robust pipeline automation engine. Its key contribution lies in being the first to deeply embed a reasoning-capable knowledge graph into ML workflow design, thereby significantly enhancing pipeline executability, transparency, and cross-domain reusability. Empirical validation across multiple real-world scientific and engineering case studies demonstrates the platform’s effectiveness: non-ML practitioners can independently build, interpret, and reuse high-quality analytical workflows.

Enables non-ML experts to build ML pipelines easilyImproves transparency and reusability of ML workflowsUses knowledge graphs to simplify ML pipeline creation

Machine learning model selection lacks formalized methodologies, making it difficult to systematically characterize contextual factors—such as data characteristics and prediction tasks—and their interactions, resulting in opaque, non-adaptive decisions. This paper introduces, for the first time, software product line (SPL) principles into ML model selection, proposing a variability-aware algorithm selection framework. It constructs a configurable feature model that explicitly captures commonalities and variabilities among contextual factors—including dataset size, feature dimensionality, and task type—as well as their logical dependencies. By integrating scikit-learn’s heuristic rules with an instantiation framework, the approach enables interpretable, adaptive, and transparent model recommendations. An empirical case study demonstrates that the method significantly outperforms existing strategies in accuracy, interpretability, and contextual adaptability.

Machine LearningModel SelectionRule Formalization

Latest Papers

What's happening recently
View more

High-Dimensional Data Processing: Benchmarking Machine Learning and Deep Learning Architectures in Local and Distributed Environments

Dec 11, 2025
JJ
José Julián Rodríguez Gutiérrez
🏛️ División de Ingenierías Campus Irapuato-Salamanca

To address the lack of unified benchmarks for model performance evaluation on high-dimensional big data in both local and distributed environments, this work designs an end-to-end evaluation framework covering three representative tasks—Epsilon (numerical regression), RestMex (text classification), and IMDb (movie feature analysis). Leveraging Apache Spark (Scala), we establish a reproducible heterogeneous computing experimental infrastructure to systematically compare traditional machine learning and deep learning models across accuracy, training efficiency, and resource consumption. This study presents the first pedagogically implemented standardized benchmark supporting multiple models, multimodal data, and diverse deployment scenarios, empirically uncovering performance bottlenecks and architectural trade-offs inherent in distributed scaling. The outcomes include an open-source evaluation pipeline, a standardized reporting template, and a reusable teaching paradigm—providing empirical foundations for AI system selection and optimization in big data contexts.

Benchmark machine learning architectures for high-dimensional data processingCompare local and distributed computing environments for big dataImplement workflows for text analysis and classification tasks

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This study addresses a critical gap in clustering interpretability: existing post-hoc explanation methods primarily focus on feature importance or instance-level explanations and struggle to reliably uncover structured patterns within clusters. To systematically evaluate this limitation, the authors conduct the first controlled assessment of multiple explanation techniques—including random forest permutation importance, LIME, and principal component analysis—in synthetic datasets where ground-truth structured patterns are explicitly embedded. Results demonstrate that while these methods partially recover relevant features, none consistently identifies all types of predefined patterns. This reveals a fundamental shortcoming of current interpretability tools in capturing pattern-level cluster structure and underscores the urgent need for dedicated methods designed specifically for detecting and explaining such intra-cluster patterns.

cluster interpretationexplainabilityfeature importance

This study addresses the limited reliability and diagnostic capability of existing AutoClustering systems, which stem from a lack of interpretability regarding how meta-features influence the selection of clustering algorithms and hyperparameters. For the first time, the work systematically reviews the meta-features employed across 22 AutoClustering methods and organizes them into a coherent taxonomy. By integrating global interpretability through decision predicate graphs and local interpretability via SHAP values, the authors conduct a thorough analysis of feature contributions within meta-models. Their investigation uncovers structural biases and consistent patterns in current meta-learning strategies, revealing fundamental limitations of prevailing approaches. These insights not only expose critical shortcomings but also offer actionable interpretability guidelines for designing transparent and trustworthy unsupervised AutoML systems.

AutoClusteringAutoMLexplainability

This work addresses the challenge of effectively analyzing massive, heterogeneous high-performance computing (HPC) logs, which hinders fault diagnosis and performance optimization. The authors propose a scalable log analysis workflow that uniquely integrates frequent pattern mining based on finite-state automata with job-level log correlation. By leveraging the Aho–Corasick automaton for efficient pattern storage and matching, and incorporating system hierarchy and message priority information, the approach enables automated detection and clustering of errors and anomalous events. Experiments on an exascale-class supercomputing system demonstrate that the method accurately identifies characteristic error sequences, reveals distinct failure patterns across different applications, and supports real-time, interpretable monitoring to enhance system resilience.

anomaly detectionHPC logslog analysis

Hot Scholars

SM

Saif M. Mohammad

Principal Research Scientist (NLP,Computational Affective Science), National Research Council Canada
Natural Language ProcessingLexical SemanticsAffective ScienceEmotion Recognition
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion
JZ

Jiayuan Zhou

Principal Researcher, Waterloo Research Centre, Huawei Canada
OSS VulnerabilitiesCrowdsourced Software EngineeringMining Software RepositoriesEmpirical
KS

Kun Su

Google Research
Multimodal LearningAudio/Music GenerationRecommendation system
SM

Saikat Mondal

Doctoral Researcher, University of Saskatchewan, Canada
Empirical Software EngineeringAI4SEAI EngineeringData Analytics