🤖 AI Summary
This paper addresses fundamental challenges in optimization and data science: streaming convex polyhedral approximation, sparse robust least-squares regression, adversarial optimization, modeling backdoor data poisoning attacks, and robustness analysis of graph clustering under “benign misspecification.” Methodologically, it unifies these problems through a high-dimensional geometric lens, integrating random projections, streaming algorithm design, statistical learning theory, and graph signal processing. Key contributions are: (1) the first provably secure theoretical model for backdoor attacks, yielding explicit safety thresholds; (2) novel algorithms with rigorous approximation and statistical guarantees, significantly improving streaming approximation accuracy and noise resilience; and (3) a characterization of strong consistency—despite model misspecification—for several classical graph clustering methods, precisely delineating their robustness boundaries.
📝 Abstract
We give new results for problems in computational and statistical machine learning using tools from high-dimensional geometry and probability. We break up our treatment into two parts. In Part I, we focus on computational considerations in optimization. Specifically, we give new algorithms for approximating convex polytopes in a stream, sparsification and robust least squares regression, and dueling optimization. In Part II, we give new statistical guarantees for data science problems. In particular, we formulate a new model in which we analyze statistical properties of backdoor data poisoning attacks, and we study the robustness of graph clustering algorithms to ``helpful'' misspecification.