Multiple data-driven missing imputation

📅 2025-07-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the adaptive imputation of short-to-moderate-length (1–5+ points) missing segments in univariate time series—occurring at arbitrary positions (beginning, middle, or end)—where local temporal characteristics (e.g., trend, volatility) vary significantly. To this end, we propose KZImputer, a novel method that dynamically selects the optimal imputation strategy based on both missing-segment location and local signal structure: extrapolative linear fitting for leading gaps, a hybrid of local mean and linear interpolation for interior gaps, and backward trend estimation for trailing gaps. KZImputer demonstrates robust performance even under high missingness rates (>50%), substantially outperforming conventional methods (e.g., linear, spline, and last-observation-carried-forward imputation). Extensive experiments show that it achieves state-of-the-art accuracy in MAE and RMSE, while preserving signal morphology more faithfully—as quantified by dynamic time warping (DTW) distance and spectral similarity. Notably, gains are most pronounced in sparse-data regimes, enhancing reliability for downstream analytical tasks.

Technology Category

Machine Learning: Kernel MethodsData Mining & Knowledge Management: Mining of Spatial, Temporal or Spatio-Temporal DataCognitive Modeling & Cognitive Systems: Neural Spike Coding

Application Category

Graph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalization
📝 Abstract
This paper introduces KZImputer, a novel adaptive imputation method for univariate time series designed for short to medium-sized missed points (gaps) (1-5 points and beyond) with tailored strategies for segments at the start, middle, or end of the series. KZImputer employs a hybrid strategy to handle various missing data scenarios. Its core mechanism differentiates between gaps at the beginning, middle, or end of the series, applying tailored techniques at each position to optimize imputation accuracy. The method leverages linear interpolation and localized statistical measures, adapting to the characteristics of the surrounding data and the gap size. The performance of KZImputer has been systematically evaluated against established imputation techniques, demonstrating its potential to enhance data quality for subsequent time series analysis. This paper describes the KZImputer methodology in detail and discusses its effectiveness in improving the integrity of time series data. Empirical analysis demonstrates that KZImputer achieves particularly strong performance for datasets with high missingness rates (around 50% or more), maintaining stable and competitive results across statistical and signal-reconstruction metrics. The method proves especially effective in high-sparsity regimes, where traditional approaches typically experience accuracy degradation.
Problem

Research questions and friction points this paper is trying to address.

Imputing missing points in univariate time series data
Handling gaps at start, middle, or end of series
Improving accuracy for high missingness rates (50%+)
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hybrid strategy for missing data scenarios
Tailored techniques for gap positions
Linear interpolation and localized statistics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Interregional Academy of Personnel Management | Luxena Ltd.
S
Sergii Kavun
Interregional Academy of Personnel Management, Kyiv, Ukraine; Luxena Ltd., Lead of Data Science Team, Kyiv, Ukraine