🤖 AI Summary
This study investigates the factors influencing the asymptotic time complexity of insertion and lookup operations in double-hashing linear probing tables. By comparing single-choice with various non-greedy insertion strategies across different load factors, and incorporating element eviction mechanisms, it provides a rigorous asymptotic complexity analysis. The work reveals the significant advantages of the "power of two choices" phenomenon, demonstrating that non-greedy strategies can surpass classical theoretical lower bounds. Specifically, without eviction, the lookup time is bounded by O(log ε⁻¹). When eviction is introduced, even under full table capacity, O(1) lookups and O(ε⁻¹/²) insertions are achievable. These findings offer novel theoretical foundations for hash table design.
📝 Abstract
This paper considers the following basic question: If an (ordered) linear-probing hash table is allowed \emph{two} hash functions, instead of one, how does this change the expected insertion and query time, as a function of the load factor $1 - ε$? We prove that the \emph{greedy two-choice insertion strategy} achieves polynomially better bounds than the single choice algorithm, but that one can even do \emph{much better} by using more sophisticated non-greedy strategies. Specifically, we show that there is an insertion strategy that does not evict elements (once an element is inserted, its hash choice is fixed) and that achieves expected query time $O(\log ε^{-1})$ with expected insertion time $O(ε^{-1})$. We then further show that, if one is allowed to evict elements (i.e., to change over time which hash function a given element uses), then it is possible to achieve expected query time $O(1)$ with expected insertion time $O(ε^{-1/2})$. This final result achieves an expected query time of $O(1)$ even when the hash table is filled to $100\%$ full. Combined, the results reveal that there is a surprisingly strong ``power of two choices'' phenomenon for linear-probing hash tables, allowing for a two-choice hash table to achieve significantly better bounds than what might at first seem to be possible.