🤖 AI Summary
This study addresses whether black-box algorithms can exploit structural knowledge to accelerate hash collision detection. Through query complexity analysis and randomized algorithm design, this work proposes an instance-optimal black-box collision detection algorithm that requires no prior knowledge of specific structural information. The primary contribution is a proof that a single algorithm can achieve near-instance-optimality within the birthday bound, constraining query complexity to $O(q \log n)$ while matching established lower bounds. This result demonstrates the impossibility of embedding purely structural backdoors into hash functions, thereby partially resolving a related open problem and establishing definitive security boundaries under black-box access.
📝 Abstract
Can structural knowledge about a hash function help accelerate the (black box) detection of collisions in it? This question is fundamental to cryptography theory given the importance of collision-resistant hash functions, and in this paper we tackle it from the angle of instance optimality, an ultimate notion of beyond worst case algorithm analysis that has gained significant traction in recent years. Instance optimality asks for a single algorithm that, on every input, performs nearly as well as the best correct algorithm that ``knows the structure'' of that specific input. Here we measure algorithms by the number of queries they make to the hash function $f\colon [n]\to [n]$, and we say that an algorithm ``knows the structure'' of the input if, in addition to query access to $f$, it has free access to an unlabeled copy $π^{-1}\circ f\circπ$ of $f$, for an unknown permutation $π$ on $[n]$.
We prove the existence of an (almost) instance-optimal algorithm for collision detection in the regime most interesting from a cryptographic perspective: among functions where finding a collision takes significantly less than $\sqrt{n}$ queries. Specifically, we prove the existence of a single algorithm $A$ that, for any input $f$ in which a structure-aware algorithm can find a collision using $q\leq O(\sqrt{n/\log n})$ queries in expectation, $A$ can find a collision in at most $O(q\log n)$ queries. The $O(\log n)$ multiplicative overhead is tight, matching a lower bound of Ben-Eliezer, Grossman, and Naor [ICALP'25], and partially resolving their main open question. Our result implies, in particular, that it is impossible for a cryptographic designer to plant purely structural backdoors for collision finding (for this unlabeled notion of structure): whatever collisions the designer's secret knowledge finds, the public can find with a multiplicative overhead of $O(\log n)$.