🤖 AI Summary
This study addresses the unresolved optimal work and parallel depth of spooky pebbling for binary tree dependencies in quantum computing under fixed space budgets. To this end, it proposes a block-cleaning strategy tailored for complete binary trees, integrating asymptotic complexity analysis with parallel scheduling techniques to optimize the pebbling process. The key contributions include determining the optimal work and parallel depth under arbitrary space constraints, reducing the time complexity at minimum budget from O(n log n) to Θ(n log n / log log n), thereby resolving a longstanding time-optimality challenge. Furthermore, when applied to RNS point addition trees, the proposed approach achieves asymptotically optimal work while reducing gate counts by approximately 20%–23%.
📝 Abstract
Pebble games model computations under a fixed space budget. Spooky pebbling allows quantum memory to be released by measurement, with the resulting phases corrected later. We study two-input computations with binary-tree dependencies and determine the optimal work and parallel depth for complete binary trees. Our key idea is to clean up the tree in blocks, reducing repeated recomputation of intermediate values. For the complete tree $B_h$ with $n=2^h-1$ vertices and every space budget $h+1\le s\le n$, we give an algorithm with asymptotically optimal work $Θ(nh/\log(s+1))$. At the minimum budget $s=h+1$, this improves the $O(n\log n)$ bound of Kornerup, Sadun, and Soloveichik to $Θ(n\log n/\log\log n)$, resolving their time-optimality question.
We also construct a parallel schedule with optimal depth \[
Θ\!\left(h+\frac ns
\max\!\left\{\frac{h}{\log(s+1)},\,1+\log^*h\right\}\right). \] Our lower bounds hold for every full binary tree. We also show that achieving optimal parallel depth can require asymptotically more work than minimizing work alone. Applying our schedule to the RNS point-addition trees in the public implementation of Chevignard, Fouque, and Schrottenloher reduces their Toffoli/AND gate count by $20.86\%$ for a P-224 instance and $22.66\%$ for a P-256 instance, using the same arithmetic circuits and peak workspace.