🤖 AI Summary
This study addresses the challenge of fault isolation from raw logs due to the lack of prior knowledge by proposing LOGOS, an unsupervised system that introduces a zero-prior diagnostic paradigm. Requiring neither seed queries nor predefined boundaries, LOGOS mines entity-event co-occurrences and temporal precedence relations to compress massive log volumes into a compact precedence forest structure. Experimental results demonstrate that LOGOS achieves a median processing time of only 4.5 minutes while eliminating 99.8% of noise and attaining a recall of 0.76. Furthermore, it provides warnings up to 16 hours in advance and improves the accuracy of LLM-based root cause analysis to 80%–100%. These findings indicate that LOGOS significantly reduces the reliance on domain expertise for troubleshooting proprietary software systems.
📝 Abstract
Commercial observability platforms rely on domain artifacts like distributed traces, topology maps, and baseline metrics. However, when troubleshooting proprietary software, enterprise operators are left with only raw, unannotated text logs. We explore the extreme boundary of log-only diagnosis: To what extent can we isolate failure propagation using strictly raw text logs? We present LOGOS, an unsupervised system that exploits entity-event co-occurrence and temporal precedence to collapse millions of raw log lines into a compact precedence forest. Evaluated across 25 production enterprise outages and 12 open-source issues, LOGOS operates with zero prior knowledge---requiring no seed queries, observed symptoms, or pre-defined incident boundaries. In a median wall-time of 4.5 minutes, LOGOS eliminates a median 99.8% of background noise, achieves 0.76 mean recall, and detects failure cascades with a 16-hour median diagnosis-verified lead time---consolidating alert floods 124x to enable 80% enterprise (100% open-source) zero-shot LLM root-cause accuracy.