๐ค AI Summary
This study addresses the performance degradation in lakehouse tables caused by small-file accumulation, a problem exacerbated by the absence of principled merge-triggering mechanisms. The authors develop an open simulation framework that generates diverse table layouts based on Apache Iceberg and extracts 17 metadata features to predict post-merge file reduction ratios with high accuracy using XGBoost (Rยฒ=0.998). Their work is the first to systematically uncover strong correlations between metadata characteristics and merge efficacy. They further propose a model-free merging strategy relying solely on a single partition-level threshold, which demonstrates remarkable generalization across heterogeneous workloads (Rยฒ=0.976). Experimental results show that this approach substantially improves metadata-intensive query performance while also revealing its subtle impact on the parallelism of full-table scan compaction.
๐ Abstract
Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhile remains threshold-driven, but which metadata features actually determine compaction utility is not well understood. We present an open simulation framework that generates 2,376 Apache Iceberg tables spanning three orders of magnitude in file size, extracts 17 metadata features from manifest files without reading data, and trains XGBoost to predict the continuous file-reduction ratio (R2 = 0.998, RMSE= 0.013). The binary compaction decision turns out to be trivially separable by a single partition-level threshold max_files_per_partition> 4, requiring no learned model. Cross-schema validation on 96 TPC-H tables confirms generalisation without retraining (R2 = 0.976). A query benchmark reveals that compaction benefits metadata-heavy queries but can slow full-scan aggregations by reducing task parallelism. All code and data are publicly available.