Score
Design and implement compact, discriminative descriptor vectors or embedding functions that represent raw signals and encode semantic correlation structure while being robust to noise. Also build and analyze descriptor-matching methods and pipelines for comparing or matching those descriptors across different inputs.
Single-stage detectors struggle with small-object detection due to semantically impoverished shallow features and low-resolution deep features. To address this, we propose a “cross-layer feature borrowing” paradigm: discriminative deep semantics from large, same-category objects are identified via feature matching, then dynamically integrated into shallow feature maps through weighted aggregation and context-aware fusion—enhancing small-object representations without compromising inference speed or breaking the resolution–semantics trade-off. Our method is built upon the SSD framework and supports end-to-end differentiable training. On the COCO benchmark, it achieves a +4.2% improvement in AP for small objects while maintaining real-time inference speed, demonstrating robustness and generalization across complex scenes.
This work addresses the high-dimensional redundancy inherent in word and sentence embeddings. We propose a semantic compression method based on Discrete Wavelet Transform (DWT), which captures hierarchical semantic structures via multi-scale wavelet decomposition—contrasting with conventional linear dimensionality reduction. To our knowledge, this is the first systematic application of DWT to embedding compression, offering both theoretical novelty and practical utility. Experiments demonstrate that embedding dimensions can be reduced by 50%–93% while preserving semantic similarity performance (degradation <0.5%). Moreover, downstream task accuracy improves by 1.2–2.8% on average across text classification, natural language inference (NLI), and semantic textual similarity (STS). The method is compatible with diverse pre-trained language models—including BERT, RoBERTa, and XLM-R—and exhibits strong cross-lingual generalization.
Embedding models trained independently on similar data capture stable semantic meanings but yield inconsistent representation spaces, hindering interoperability across models. This work addresses compatibility challenges in multimodal search and model upgrades via orthogonal transformation-based embedding alignment. Theoretically, we derive the first tight Procrustes alignment error bound, proving the existence of a near-isometric orthogonal transformation that approximately preserves pairwise inner products—establishing rigorous theoretical foundations for alignment. Methodologically, we employ efficient Procrustes analysis as a post-hoc alignment procedure, preserving the intrinsic geometric structure of each embedding space while enabling cross-model alignment. Experiments demonstrate substantial improvements in model retraining compatibility, text retrieval fusion accuracy, and cross-modal search performance; our method achieves state-of-the-art results in hybrid multimodal search.
To address data representation in supervised multi-class classification, this paper proposes a unified linear feature extraction method, RSLDA-ICS_DLSR, integrating Sparse Linear Discriminant Analysis (SLDA) with Discriminative Least Squares Regression (DLSR). The core contribution lies in a joint optimization framework that enforces row-wise sparsity consistency within classes while enabling flexible integration and parameter tuning of diverse discriminative criteria. The method jointly learns a linear transformation matrix and an orthogonal matrix via sparse regularization, iterative alternating minimization, and multi-strategy initialization. Extensive experiments on benchmark datasets—including AR and Extended YaleB (face), Caltech-101 (object), and MNIST (handwritten digits)—demonstrate that RSLDA-ICS_DLSR consistently outperforms state-of-the-art linear discriminant methods in classification accuracy, validating its superior representation capability and generalization performance.
Prompt-based text embeddings exhibit significant dimensional redundancy, leading to excessive storage and computational overhead. To address this, we systematically investigate the geometric structure of prompt embeddings and its relationship with downstream task performance. We propose a post-hoc compression framework leveraging principal component analysis (PCA) and other dimensionality reduction techniques, complemented by quantitative analysis using maximum likelihood estimation (MLE) of intrinsic dimensionality and anisotropy metrics. Experiments demonstrate that retaining only 0.5% of the original dimensions incurs less than 1% performance degradation on classification and clustering tasks; at 25% dimensionality, average performance loss across classification, clustering, retrieval, and semantic similarity tasks remains below 0.3%. This work is the first to reveal the extremely low intrinsic dimensionality of prompt embeddings and to establish a quantitative correlation between embedding anisotropy and task sensitivity—providing both theoretical foundations and practical solutions for efficient embedding deployment.
Existing binary code similarity detection methods struggle to balance interpretability, generalizability, and scalability: handcrafted features are human-readable but exhibit poor generalization, while embedding-based approaches are robust yet opaque, difficult to verify, and suffer from low retrieval efficiency. This paper introduces the first language-model-based agent framework for structured reasoning over assembly code, enabling semantic parsing that automatically generates interpretable and verifiable structured features—including input/output types, side effects, and critical constants. By integrating inverted indexing with relational indexing, the method achieves efficient and scalable retrieval. Crucially, it requires no training and attains recall@1 of 42% (cross-architecture) and 62% (cross-optimization), matching state-of-the-art supervised embedding methods. When fused with embeddings, it significantly surpasses prior art.
This work addresses the high computational cost of similarity search in large-scale multimodal wildlife datasets, where text-based retrieval is hindered by expensive high-dimensional operations. The authors propose a compact hypercube embedding framework that extends lightweight hashing to cross-modal alignment of text with images and audio for the first time. By applying parameter-efficient fine-tuning to pretrained foundation models such as BioCLIP and BioLingual, the method learns binary embeddings residing in a shared Hamming space. This approach drastically reduces both storage requirements and retrieval latency while achieving retrieval performance on par with or superior to continuous embeddings on the iNaturalist2024 and iNatSounds2024 benchmarks. Moreover, it enhances zero-shot generalization capabilities, demonstrating the effectiveness of discrete representations in complex multimodal ecological tasks.
This work addresses the challenge of recovering cross-model object correspondences when given only two partially overlapping embedding datasets produced by distinct black-box encoders. It identifies and leverages, for the first time, the cross-model isometric consistency in local geometric structures induced by contrastive learning encoders, and proposes a training-free, iterative geometric embedding hashing method. Guided by a small set of seed anchor points, the approach progressively expands correspondences across views by integrating distance-based hashing with Beta-Bernoulli Bayesian posterior aggregation. Experiments demonstrate that the method achieves high-precision and robust vector alignment across diverse encoder pairs, overlap ratios, and anchor conditions, successfully enabling applications such as vector database fusion and cross-model clustering.
To address poor generalization in high-dimensional tabular data classification—caused by excessively high embedding dimensions and scarce labeled samples—this paper proposes Partitioned-LDA, a linear discriminant dimensionality reduction method integrating blockwise covariance estimation with shrinkage regularization. It partitions high-dimensional word embeddings into non-overlapping blocks, estimates local covariance matrices per block, and applies shrinkage correction to mitigate estimation bias under small-sample conditions, thereby enhancing the stability and discriminative power of Linear Discriminant Analysis (LDA). Experiments demonstrate that even a 2-dimensional Partitioned-LDA projection surpasses the classification accuracy of the original high-dimensional embeddings on public benchmark leaderboards, consistently ranking among the top ten. Compared to standard PCA and shrinkage-regularized LDA, Partitioned-LDA exhibits superior robustness and higher dimensionality-reduction efficiency under limited training data. This work establishes a novel, interpretable, lightweight, and high-performance embedding compression paradigm for low-resource tabular classification.
研究解决了小样本判别分析中因平衡k-shot采样导致的退化问题,通过修正KLPCDA变体并评估其在文本分类任务中的表现。