🤖 AI Summary
This work investigates the zero-shot and few-shot time-series anomaly detection capabilities of large language models (LLMs). Methodologically, it conducts multi-model comparative experiments (including Llama, Qwen, and GPT series), employs time-series image-based encoding, designs a hypothesis-driven controllable evaluation framework, and applies systematic prompt engineering. The study yields four counterintuitive findings: (1) LLMs perform significantly better when processing time-series as images rather than text; (2) explicit chain-of-thought or reasoning prompts yield no statistically significant improvement; (3) repetition bias and arithmetic reasoning are not primary mechanisms underlying anomaly recognition; and (4) architectural differences lead to substantial performance variation. These results challenge prevailing assumptions in the field and empirically demonstrate that LLMs possess foundational—but non-trivial—time-series anomaly detection capabilities. To foster reproducibility and further research, the authors open-source both the implementation code and a dedicated benchmark dataset, AnomLLM.
📝 Abstract
Large Language Models (LLMs) have gained popularity in time series forecasting, but their potential for anomaly detection remains largely unexplored. Our study investigates whether LLMs can understand and detect anomalies in time series data, focusing on zero-shot and few-shot scenarios. Inspired by conjectures about LLMs' behavior from time series forecasting research, we formulate key hypotheses about LLMs' capabilities in time series anomaly detection. We design and conduct principled experiments to test each of these hypotheses. Our investigation reveals several surprising findings about LLMs for time series: 1. LLMs understand time series better as images rather than as text 2. LLMs did not demonstrate enhanced performance when prompted to engage in explicit reasoning about time series analysis 3. Contrary to common beliefs, LLM's understanding of time series do not stem from their repetition biases or arithmetic abilities 4. LLMs' behaviors and performance in time series analysis vary significantly across different model architectures This study provides the first comprehensive analysis of contemporary LLM capabilities in time series anomaly detection. Our results suggest that while LLMs can understand time series anomalies, many common conjectures based on their reasoning capabilities do not hold. Our code and data are available at `https://github.com/Rose-STL-Lab/AnomLLM/`.