BetweenCut: Private Heavy-Node Classification with Doubly Logarithmic Error in Tree Height

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of identifying record-level differentially private heavy nodes in tree-structured data, where existing methods suffer from path accumulation errors that grow significantly with tree height. To overcome this limitation, this work proposes BetweenCut, an algorithm that optimizes threshold comparison strategies to suppress multi-level privacy budget consumption, thereby enabling private heavy node classification in deep tree structures. The primary contribution lies in reducing the additive error bound from logarithmic or square-root levels to a double-logarithmic level of O(log log h), independent of database size. Furthermore, all nodes simultaneously satisfy (ε, δ)-differential privacy constraints, substantially improving identification accuracy in deep-tree scenarios.
📝 Abstract
Finding heavy nodes in a tree---those whose counts exceed a given threshold---is a building block for analysis and learning over structured data. Achieving record-level differential privacy (DP) without sacrificing accuracy is challenging because each record contributes to counts along an entire root-to-leaf path, allowing privacy costs to accumulate across levels. Existing methods account for the multiple threshold comparisons for each record incur additive error margins of $Ω_{\varepsilon,δ}(\log h)$ or $Ω_{\varepsilon,δ}(\sqrt{\log h})$ for tree height $h$. We introduce \textsc{BetweenCut}, an $(\varepsilon,δ)$-DP algorithm with an additive error margin of $O_{\varepsilon,δ}(\log\log h)$, improving the existing bounds for deep trees. This error holds simultaneously for all nodes and is independent of the input database size.
Problem

Research questions and friction points this paper is trying to address.

heavy-node classification
differential privacy
tree structure
additive error
privacy cost accumulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differential Privacy
Heavy-Node Classification
Tree Structure
BetweenCut
Doubly Logarithmic Error