🤖 AI Summary
This study addresses the limitation of existing persistence diagram two-sample tests, which fail to localize regions of topological discrepancy. To overcome this, we propose a simultaneous inference framework for local mean contrasts under a fixed budget. Local contrasts are estimated via additive landmark responses, and confidence intervals are calibrated using the Gaussian multiplier bootstrap. Furthermore, geometric sufficient conditions under which feature displacement induces non-zero contrasts are established. The primary contribution is the construction of difference maps with rigorous family-wise error rate control. Experiments demonstrate that the proposed method achieves simultaneous coverage rates of 94%–98%, significantly outperforming conventional approaches. Additionally, it successfully localizes regions of discrepancy in fused ring systems on the MUTAG benchmark, validating its practical utility for interpretable topological data analysis.
📝 Abstract
Many two-sample tests for populations of persistence diagrams assess global differences without identifying the regions of the birth-death plane that contribute to them. We study simultaneous inference for local mean contrasts when the number of available diagrams is fixed. They are differences in expected weighted feature mass within $\ell_\infty$ neighborhoods at several centers and radii. We estimate these contrasts using additive landmark responses. A Gaussian multiplier bootstrap calibrates simultaneous confidence intervals while allowing unequal group covariances. The neighborhoods whose intervals exclude zero form a map with approximate family-wise error control, and selecting a subset of original intervals for display preserves their joint coverage guarantee. On the simultaneous coverage event, every reported neighborhood lies within twice its radius of the support of the mean-measure difference. A geometric result gives sufficient radius conditions for a displaced feature to produce a nonzero contrast. A comparison of sufficient detection thresholds quantifies the tradeoff between reducing the number of tested coordinates and reserving observations for an independent pilot. In simulations with 40 to 120 diagrams per class, the bands achieved 94%-98% simultaneous coverage under both the strict null and equal means with unequal covariances. In the latter setting, a permutation maximum and the pooled-t implementation of the two-stage persistence-image test of Moon and Lazar rejected in up to 32% and 26% of runs, respectively. In the fixed-budget simulations, spending a third of the observations on a pilot to choose landmarks or radii located changes less often than a prespecified grid at a single radius. On the MUTAG benchmark, the localized region concentrates on rings of fused-ring systems, an exploratory reading.