READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of fault-type semantic evaluation benchmarks in existing fault diagnosis retrieval and the ambiguity caused by reliance on indirect downstream prediction metrics. We construct a time-series diagnostic retrieval benchmark spanning 12 datasets and propose a semantic relevance criterion based on shared fault types to systematically evaluate classical distance measures, foundation model embeddings, and Gaussian process reranking. Our experiments demonstrate that Gaussian process reranking with limited annotations outperforms complex pretraining strategies and large language model reasoning. This approach significantly improves NDCG@10 across all datasets, achieving gains of 0.11 in reranking performance and 0.16 in overall performance over the strongest baseline, thereby establishing an efficient paradigm for fault retrieval.
📝 Abstract
Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so visually different traces of the same fault count as relevant while similar-looking traces of different faults do not. Treating retrieval as a base retriever followed by a reranker, we evaluate classical distances, symbolic retrievers, self-supervised and foundation-model embedders, and their fusion, plus label-aware and language-model rerankers, under one protocol that varies supervision, pollution, and corpus scale with significance testing. Under a common channel-independent interface, pretrained representations offer no statistically detectable advantage over strong classical and symbolic baselines for search alone. The decisive factor is a small amount of resolved-case supervision at reranking, namely a Gaussian-process reranker that propagates a few neighbor labels in embedding space, which helps far more than more sophisticated representations or language-model reasoning and holds under pollution and at full corpus scale. Guided by these findings, we fuse a normal-residual-scored embedder with a dynamic time warping leg via reciprocal-rank fusion, then rerank with the Gaussian-process reranker, improving NDCG@10 over its own search stage on all 12 datasets, by +0.11 from reranking and +0.16 over the strongest single base retriever.
Problem

Research questions and friction points this paper is trying to address.

Time-series diagnosis
Historical instance retrieval
Benchmark
Multivariate time series
Relevance definition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time-Series Retrieval
Benchmark
Gaussian Process Reranker
Reciprocal Rank Fusion
Dynamic Time Warping
🔎 Similar Papers
No similar papers found.