projected dimensionality reduction

Designs and analyzes algorithms and systems that compress high-dimensional vectors by projecting them into lower-dimensional subspaces (using sparse or random projections and related compressed representations) to enable efficient aggregation. Builds projection-aware aggregation routines and weighting schemes that compute client reliability or preserve distance-based robustness while reducing server computation and communication complexity.

projecteddimensionalityreduction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the trade-off between communication overhead and aggregation accuracy in distributed learning when transmitting sparse local models. It proposes the first compression scheme that integrates covering codes with sketching techniques and establishes a dual information-theoretic lower bound based on f-divergence. This bound is tight for binary alphabets and strictly stronger than conventional Fano-type bounds. The proposed method achieves a communication–accuracy trade-off approaching the theoretical limit in frequency estimation tasks. Although it does not attain the bound for general alphabets, it opens a new avenue for future research in this direction.

communication-accuracy tradeoffdistributed learningf-divergence

Dual-Lagrange Encoding for Storage and Download in Elastic Computing for Resilience

Jan 28, 2025
XZ
Xi Zhong
🏛️ University of Florida | Rowland Hall St. Marks High School | New Jersey Institute of Technology

Existing approaches to large-scale matrix multiplication in elastic computing suffer from poor straggler tolerance, high data upload overhead, and excessive storage redundancy—requiring full local storage of input matrices. To address these challenges, this paper proposes a novel two-stage Lagrange coding scheme, the first to jointly optimize both storage and download efficiency in elastic environments. The method supports VM preemption, dynamic scaling, and heterogeneous node scheduling, reducing storage overhead to 1/L of Zhong et al.’s scheme and overcoming limitations of conventional full-storage or fault-intolerant designs. By integrating distributed matrix blocking with heterogeneous and cyclic task assignment strategies, our approach achieves load balancing, minimizes total computation time and computational redundancy, and significantly improves fault tolerance and resource utilization—as empirically validated on AWS EC2.

Large Matrix MultiplicationSlow Machine ToleranceStorage Efficiency

This work addresses the issue of uncontrolled reconstruction errors in SVD-based compression of large collections of matrices when heuristic grouping is employed prior to concatenation. To overcome this limitation, the authors propose a theory-driven compressive clustering framework grounded in spectral analysis of horizontally concatenated matrices. They establish, for the first time, a globally provable upper bound on SVD reconstruction error and derive two novel spectral bounds based on a lower bound for singular value growth. Building upon these theoretical guarantees, they design three clustering algorithms with explicit error control, integrated with incremental approximate SVD to efficiently estimate compression error without explicitly forming the full concatenated matrix. The resulting approach achieves a favorable balance among speed, accuracy, and scalability, significantly enhancing the reliability and practicality of SVD compression in applications such as multi-view learning, signal processing, and neural network compression.

error-constrained clusteringmatrix concatenationreconstruction error

Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation

Mar 11, 2024
JL
Jan Laukemann
🏛️ Friedrich-Alexander-Universität Erlangen-Nüernberg | Intel Labs | University of Oregon | Laboratory for Physical Sciences

This work addresses efficient decomposition of high-dimensional sparse tensors—common in healthcare and cybersecurity—on modern parallel processors, overcoming restrictive assumptions about mode structure or sparsity distribution inherent in conventional compressed formats. We propose ALTO, an adaptive linearization tensor representation that is agnostic to both mode structure and sparsity distribution. Built upon ALTO, we design a parallel decomposition algorithm featuring low synchronization overhead and high data reuse, augmented by dynamic performance modeling and scheduling heuristics for automatic hardware adaptation. Leveraging cache- and memory-aware optimizations on Intel Xeon Scalable platforms, experiments demonstrate that ALTO achieves over 10× speedup versus the best structure-agnostic format and a 5.1× geometric mean speedup versus the best structure-aware format, while incurring only 25% of the latter’s storage overhead.

Efficient decomposition of high-dimensional sparse tensorsOvercoming irregular shapes and data distributions in sparse tensorsReducing memory footprint and synchronization overhead in tensor computations

Latest Papers

What's happening recently
View more

This work addresses the high computational overhead of existing robust aggregation methods in federated learning under Byzantine attacks, which stems from processing high-dimensional gradients and hinders scalability to large models. The authors propose a Projection-based Dimensionality Reduction (PDR) framework that compresses client gradients into a low-dimensional subspace via sparse random projections, substantially reducing server-side aggregation complexity while preserving robustness. PDR is the first method to enable generic acceleration of distance-based robust aggregators at the vector level, approaching the theoretical lower bound on computation. It guarantees optimal convergence rates of $O(1/\sqrt{T})$ in non-convex settings and $O(1/T)$ under strong convexity. Experiments demonstrate that PDR achieves speedups of several orders of magnitude in aggregation time, introducing only controllable approximation error while maintaining both efficiency and convergence performance.

Byzantine attacksComputational OverheadDimensionality Reduction

Clustering high-dimensional discrete data is often hindered by high computational cost, sensitivity to sparsity, and limited methodological applicability. This work proposes a deterministic dimensionality reduction framework that compresses binary, categorical, or count-based high-dimensional discrete data into low-dimensional continuous representations via weighted positional encoding. The resulting mapping is injective, preserving the discriminative structure of the original data; under mild conditions, the compressed variables approximately follow a Gaussian distribution while maintaining inter-cluster distances, thereby ensuring identifiable clustering structures. Empirical evaluations on real-world datasets—including infant names and microbiome profiles—demonstrate that the method achieves high clustering accuracy and substantially outperforms mainstream dimensionality reduction techniques in computational efficiency, offering both strong theoretical guarantees and practical utility.

clusteringcomputational efficiencydata compression

This work addresses the high latency introduced by traditional decompression in scientific data analysis, which undermines the storage and transmission benefits of compression. To overcome this limitation, the authors propose a multi-stage, error-bounded decompression and homomorphic analysis framework. By abstracting a generic compression pipeline, the framework enables hierarchical partial decompression and introduces homomorphic operation algorithms tailored to three representative scientific analysis tasks, allowing computations to be performed directly on intermediate compressed representations without full decompression. Implemented atop four mainstream compressors and evaluated across five real-world datasets, the approach consistently reduces data access latency and significantly improves analytical efficiency across diverse workloads.

access latencyanalytical operationsdata decompression

This work addresses the activation communication bottleneck that limits pipeline-parallel training of large language models under low-bandwidth network conditions. The authors propose MAPL, a novel method that enables each pipeline stage to independently learn and dynamically optimize an orthogonal compression subspace. Task-adaptive low-rank compression is achieved through orthogonal projection learning constrained on the Stiefel manifold, while reconstruction fidelity is enhanced via factorized anchor embeddings combined with residual vector quantization. A streaming codebook synchronization protocol is further introduced to reduce communication overhead. Evaluated on LLaMA models ranging from 150M to 1B parameters, MAPL achieves negligible performance degradation even at high compression ratios, significantly outperforming existing approaches such as Subspace Networks.

activation compressioncommunication bottlenecklow-bandwidth networks

Hot Scholars

KZ

Kaixiong Zhou

Assistant Professor, North Carolina State University
Machine LearningAI4ScienceGraph Data Mining
SZ

Shuang Zhou

University of Minnesota, Hong Kong Polytechnic University
Biomedical InformaticsLarge Language ModelsAI for HealthcareElectronic Health Record
KD

Klaus Dietmayer

Professor für Mess- und Regelungstechnik
TrackingInformation FusionSituation UnderstandingAutomatic Driving
XF

Xiaolan Fu

Institute of Psychology, Chinese Academy of Sciences
Cognitive NeuroscienceEmotionMicroexpression
FL

Fang Liu

Computer Science and Engineering, Nanjing University of Science and Technology
Deep learningImage ProcessingRemote SensingSAR