READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis
This study addresses the lack of fault-type semantic evaluation benchmarks in existing fault diagnosis retrieval and the ambiguity caused by reliance on indirect downstream prediction metrics. We construct a time-series diagnostic retrieval benchmark spanning 12 datasets and propose a semantic relevance criterion based on shared fault types to systematically evaluate classical distance measures, foundation model embeddings, and Gaussian process reranking. Our experiments demonstrate that Gaussian process reranking with limited annotations outperforms complex pretraining strategies and large language model reasoning. This approach significantly improves NDCG@10 across all datasets, achieving gains of 0.11 in reranking performance and 0.16 in overall performance over the strongest baseline, thereby establishing an efficient paradigm for fault retrieval.