Linear Dimensionality Reduction for Word Embeddings in Tabular Data Classification

📅 2025-09-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address poor generalization in high-dimensional tabular data classification—caused by excessively high embedding dimensions and scarce labeled samples—this paper proposes Partitioned-LDA, a linear discriminant dimensionality reduction method integrating blockwise covariance estimation with shrinkage regularization. It partitions high-dimensional word embeddings into non-overlapping blocks, estimates local covariance matrices per block, and applies shrinkage correction to mitigate estimation bias under small-sample conditions, thereby enhancing the stability and discriminative power of Linear Discriminant Analysis (LDA). Experiments demonstrate that even a 2-dimensional Partitioned-LDA projection surpasses the classification accuracy of the original high-dimensional embeddings on public benchmark leaderboards, consistently ranking among the top ten. Compared to standard PCA and shrinkage-regularized LDA, Partitioned-LDA exhibits superior robustness and higher dimensionality-reduction efficiency under limited training data. This work establishes a novel, interpretable, lightweight, and high-performance embedding compression paradigm for low-resource tabular classification.

Technology Category

Machine Learning: Dimensionality Reduction/Feature SelectionData Mining & Knowledge Management: Data CompressionNatural Language Processing: Learning & Optimization for NLP

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsUser Modeling, Personalization and Recommendation: User privacy protection in personalized systems
📝 Abstract
The Engineers' Salary Prediction Challenge requires classifying salary categories into three classes based on tabular data. The job description is represented as a 300-dimensional word embedding incorporated into the tabular features, drastically increasing dimensionality. Additionally, the limited number of training samples makes classification challenging. Linear dimensionality reduction of word embeddings for tabular data classification remains underexplored. This paper studies Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA). We show that PCA, with an appropriate subspace dimension, can outperform raw embeddings. LDA without regularization performs poorly due to covariance estimation errors, but applying shrinkage improves performance significantly, even with only two dimensions. We propose Partitioned-LDA, which splits embeddings into equal-sized blocks and performs LDA separately on each, thereby reducing the size of the covariance matrices. Partitioned-LDA outperforms regular LDA and, combined with shrinkage, achieves top-10 accuracy on the competition public leaderboard. This method effectively enhances word embedding performance in tabular data classification with limited training samples.
Problem

Research questions and friction points this paper is trying to address.

Reducing high-dimensional word embeddings in tabular classification
Addressing limited training samples for salary category prediction
Improving linear dimensionality reduction methods for embedding performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Linear dimensionality reduction for embeddings
PCA outperforms raw embeddings appropriately
Partitioned-LDA with shrinkage enhances performance
💼 Related Jobs
No related jobs found.
L
Liam Ressel
ETIT-KIT, Germany
H
Hamza A. A. Gardi
IIIT at ETIT-KIT, Germany