🤖 AI Summary
This study addresses the challenge of field redundancy in data lakes, which significantly constrains the efficiency of AI-driven analytics. To mitigate this issue, we establish a theoretical framework grounded in the Zipf-Mandelbrot distribution and propose a low-cost heuristic filtering method that operates without expert intervention. By integrating histogram-based filtering criteria, power-law distribution modeling, and constrained exponential parameter techniques, the proposed approach rapidly identifies entity fields with high informational value from massive telemetry datasets. This work contributes an efficient, cost-effective, and fully automated mechanism for valuable entity selection, thereby substantially optimizing the data foundation for AI analysis and enhancing large-scale data processing capabilities.
📝 Abstract
Data lakes store large amounts of telemetry, with logs from network sensors, hosts, and applications containing possibly hundreds of fields for every event. Large enterprises are then left with data lakes that cannot be analyzed efficiently with AI. Aggregate analysis looks at persistent shifts in behavior over time. Many of the fields and columns in data lakes are not useful as they do not contain information that is sufficiently diverse or concentrated to support AI analysis. SAIVE is a simple method for examining a few rows in a large table and applies a histogram of histograms filtering criterion to select the fields that for AI analysis is more likely to yield useful results. This paper provides a principled foundation for the SAIVE heuristics by assuming of a Zipf-Mandelbrot power-law distribution of the underlying data. Constraining the Zipf-Mandelbrot exponent alpha to a reasonable range provides a a practical, cheap, expert-free filter for selecting AI valuable entities in large data sets.