Score
Designs and implements inference and optimization methods that process long temporal sequences by running models over overlapping short windows, including the mechanisms to initialize, carry or fuse state between windows and to reconcile outputs at window boundaries. Builds and analyzes algorithms and systems that trade off computation and memory for temporal consistency and low boundary discontinuities when performing repeated optimization or inference on streaming or long-duration time-series data.
Existing data stream learning research often relies on unrealistic assumptions—such as single-pass processing and strict online constraints—leading to ill-defined problem formulations, biased evaluation protocols, and misalignment with industrial requirements. Method: This paper systematically critiques and deconstructs these restrictive assumptions, proposing a “de-paradigmized” framework that centers modeling on concept drift and temporal dependence while abandoning rigid formal stream constraints; algorithmic design integrates time-series analysis, concept drift detection, robust statistical learning, and privacy-preserving techniques—rejecting isolated development of bespoke streaming algorithms. Contribution/Results: The work yields a methodology guide grounded in industrial practice, fostering renewed consensus between academia and industry. It significantly enhances model robustness, interpretability, and privacy compliance in real-world dynamic environments.
This work addresses the high latency, resource contention, and operational overhead caused by frequent state updates in streaming machine learning. The authors propose a probabilistic sparsification strategy that decouples inference from state persistence: while all events contribute to inference scoring, only those deemed highly informative trigger persistence. This approach enables precise control over the persistence path without requiring high-frequency in-memory control planes or cross-node coordination, while preserving unbiasedness of time-aggregated statistics. By integrating approximate statistics from disk-based key-value stores with variance-aware temporal aggregation modeling, the method reduces persistence events by up to 90%, substantially lowering I/O and serialization costs while maintaining or even improving downstream task performance.
This work addresses the lack of multi-step compositional reasoning capability in time series analysis by formally defining and systematically tackling the novel task of *multi-step time series reasoning*—requiring models to jointly support logical decomposition, precise numerical computation, and verifiable, structured reasoning. To this end, we propose the *Program-Augmented Reasoning Agent* (PARA), which integrates in-context learning, self-correction mechanisms, and deterministic program execution: a large language model (LLM) handles high-level semantic understanding, while symbolic program execution ensures computational accuracy and full traceability. Evaluated on a newly constructed benchmark for multi-step time series reasoning, PARA significantly outperforms general-purpose LLMs across accuracy, interpretability, and generalization to complex reasoning patterns. Our approach establishes a new paradigm for time series intelligence—one that is decomposable, executable, and verifiable.
Existing stream processing systems exhibit limited window definition capabilities: they either rely on imperative languages or support only simple time- or count-based windows, making it difficult to precisely specify complex, condition-driven window start/end semantics—leading to semantic ambiguity, poor customizability, and combinatorial overlap explosion. This paper proposes a formal window framework based on Monadic Second-Order logic (MSO), the first to apply MSO to stream window modeling, establishing an equivalent formal triad comprising MSO formulas, regular expressions, and finite automata. It further models window overlap as a static analysis problem and characterizes its decidability boundary. Based on this foundation, we design a semantically precise, user-friendly window definition language and an automatic compilation-and-execution engine. Evaluation in real-world scenarios—including ICU patient monitoring—demonstrates high expressiveness, low ambiguity, and bounded computational overhead.
This work addresses the fragmented and application-level implementation of preprocessing, accelerator invocation, and postprocessing in neural inference on microcontrollers, which lacks system-wide coordination. To overcome this, the authors propose abstracting the inference pipeline as an operating system primitive and introduce SynapticOS—a runtime built atop Zephyr—that enables deterministic execution with zero heap usage and a constant memory footprint (peak: 2,784 bytes) through static memory pools, frame-level reset mechanisms, and phase-order validation. Integrated with priority-based job scheduling (real-time, normal, and best-effort) and PowerQuad DSP optimizations—including self-calibrating FFT and Q15 matrix multiplication—the system achieves 215.8 FPS (4.63 ms per frame) for face detection on the NXP FRDM-MCXN947 platform, yielding a 6.7× speedup over QEMU with software floating-point while incurring only a 20.7 KB Flash overhead and passing all 99 test cases.
研究提出一种决策层方法,通过统计意义的重用、生成或延迟策略,解决流系统中专家模型池的维护问题。
This work addresses the limitations of existing foundation time series models, which suffer from high computational overhead, poor adaptability to dynamic data streams, and an inability to learn continuously—hindering their deployment in resource-constrained environments. To overcome these challenges, we propose TimeBlocks, a novel foundation modeling paradigm that uniquely integrates multi-task generalization, lightweight architecture, and continual calibration capabilities. TimeBlocks dynamically assembles compact models at inference time through a modular pool of model blocks and a routing strategy tailored to incoming data streams. Furthermore, it incorporates StreamCore, a streaming summarization algorithm that enables efficient online calibration. Extensive experiments demonstrate that TimeBlocks achieves significantly higher prediction accuracy than current methods across multiple datasets while maintaining low computational costs, enabling effective real-time forecasting under stringent resource constraints.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
本文针对流数据中超参数优化问题,提出四种边界约束处理策略,通过实验验证其优于现有方法。
Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state-space models, retrieval-augmented predictors, foundation models, and tool-using agents. These developments are typically studied in isolation, organized by architecture or modeling era. We argue that they can instead be viewed through a common question of \emph{how does a time-series model retain and access information beyond its immediate input?} This question is motivated by a fundamental limitation of conventional time-series modeling: information relevant to a prediction may lie far beyond a feasible input window, while compressing history into a fixed-size state can discard information that may become useful later. We formulate this challenge as a \emph{memory} problem and organize existing time-series methods along a spectrum from internal memory, encoded in parameters and fixed-size states, to external memory that is addressable, retrievable, and increasingly maintained by agents. We then develop a unified taxonomy of memory mechanisms and review three classes of external memory, including explicit modules, retrieval augmentation, and agentic stores, under a common framework for what is retained, how it is written and accessed, and how it persists. A cross-cutting analysis maps these mechanisms to time-series tasks and identifies gaps in both methods and evaluation. We conclude by outlining open problems in building memory systems that can selectively retain, retrieve, revise, and forget information as temporal environments evolve. The result is a framework for studying memory as a first-class dimension of time series modeling, independent of the underlying backbone.
研究解决了机器学习系统在线修正时监控器误报问题,通过Huber-style方法减少误报并提高有效性。