🤖 AI Summary
This study addresses the ambiguity of predictive error metrics following the removal of structurally similar samples in materials machine learning. To this end, it proposes a "deletion lower bound" benchmark and conditional neighbor bounds, leveraging ridge regression identities to decouple residual fitting from predictive shifts. This approach enables the quantification of retraining equivalence and balances target suppression with model utility. Experiments on Materials Project data demonstrate that reducing redundancy decreases retraining loss by approximately eightfold, confirming that predictive variations are more pronounced within the low-bound region. These findings establish a multidimensional evaluation criterion for data pruning in materials informatics.
📝 Abstract
In materials machine learning, closely related retained structures can sustain accurate property predictions even after removing a specific record, rendering post-deletion prediction error an ambiguous metric for machine unlearning. To resolve this ambiguity, we define the deletion floor as the expected target loss under a specified retraining procedure at the deleted request. Standard indistinguishability constraints yield a sharp interval bounding an update's target loss around this baseline reference. Theoretically, a conditional neighbor bound links a low deletion floor directly to retained fit, prediction regularity, and local label agreement, while an exact ridge identity isolates residual fit from the prediction change induced by record deletion. Empirically, controlled redundancy sweeps show an $\approx 8\times$ drop in median normalized retraining loss when one retained relative remains after deletion. Across two distinct fitting regimes in a paired Materials Project study, the lower-floor regime also exhibits a larger prediction change on more than 50% of the shared requests. Systematic comparisons against approximate updates and the original model decouple deliberate target suppression from preserved overall model utility. Consequently, request-level unlearning evaluations should report reference loss, prediction change, and retained utility together, interpreting post-deletion accuracy against what retraining itself leaves behind.