Score
Designs and implements systems that extract hidden activations from neural network layers and represent them as vectors, then compute k-nearest neighbors in that activation space to perform tasks such as classification, retrieval, or prompt labeling without fine-tuning. Work includes selecting layer(s), preprocessing/normalizing activation vectors, choosing distance metrics and nearest-neighbor indexing/search algorithms, and optionally fusing activation-based similarity scores with embedding or other model scores.
High-dimensional vector approximate nearest neighbor search (ANNS) suffers from efficiency bottlenecks due to linear growth of distance computation cost with dimensionality—especially acute for LLM-derived semantic vectors. This work systematically evaluates six dimensionality reduction (DR) techniques—PCA, product quantization, autoencoders, contrastive learning-based DR, LSH variants, and random projection—quantifying their acceleration effects on mainstream ANNS engines (e.g., FAISS, Annoy) under a unified experimental framework. We propose two analyzable DR–search co-design architectures and theoretically derive critical pruning gain thresholds, characterizing the fundamental trade-off between dimensionality compression and retrieval accuracy degradation. Experiments across six public benchmarks show that deep DR methods achieve 1.8–3.5× speedup while maintaining >90% recall. Furthermore, we provide a data-aware guideline for optimal DR technique selection based on intrinsic data properties.
This work addresses the lack of systematic design principles for neural retrieval systems that balance efficiency and effectiveness. It proposes the first vertically layered four-tier framework—spanning representation, granularity, orchestration, and robustness—to structurally characterize key design decisions at each layer and their interdependencies. By integrating Bi- and Cross-encoder architectures, atomic and hierarchical chunking strategies, multi-stage re-ranking, agent-based decomposition, and domain generalization techniques, the study elucidates the mechanistic impact of each design choice on system performance. This approach effectively mitigates critical challenges such as information bottlenecks, semantic blind spots, and temporal drift, thereby offering a practical and actionable optimization pathway for building efficient and robust embedded retrieval systems.
To address slow training and excessive embedding storage overhead in ultra-fine-grained classification with millions of classes, this paper proposes a fast neural training framework based on a preconfigured latent space. The method replaces the conventional learnable classifier head with semantically orthogonal and geometrically uniform target vectors—preconstructed in a low-dimensional space using structured vector systems (e.g., Aₙ root lattices). Coupled with an encoder–ViT architecture, it enables end-to-end training without an explicit classification layer. Evaluated on ImageNet-1K and large-scale datasets containing 500K–600K classes, the approach accelerates convergence by up to 2.3× while compressing the embedding vector repository to less than 10% of that required by standard methods. It thus achieves a favorable trade-off among training efficiency, generalization performance, and deployment practicality.
This work addresses the computational bottleneck in large-scale neural network classification, where label prediction complexity scales linearly (O(n)) with the number of classes—often reaching millions. To overcome this limitation, the authors propose a geometric modeling approach in latent space based on a predefined vector system, which reformulates classification as an O(1) nearest cluster-center search by simply identifying extremal indices in the embedding vector. This method significantly reduces inference complexity without compromising training accuracy and inherently supports recognition of novel classes. Experimental results across multiple large-scale datasets demonstrate up to 11.6× overall inference speedup, substantially enhancing the efficiency of ultra-large-scale classification tasks.
Neural networks face critical limitations in safety-critical domains (e.g., healthcare, industrial control) due to hallucination, high computational cost, catastrophic forgetting, and poor interpretability. To address these challenges, we propose a novel AI paradigm grounded in nearest-neighbor search and hierarchical clustering. Our core innovation is a tree-structured index integrating Kohonen self-organizing maps, enabling efficient, retraining-free model extension and fine-tuning. By replacing parametric generation with explicit memory retrieval, the approach substantially mitigates hallucination while enhancing cognitive alignment and transparency. Evaluated on handwritten digit recognition and image caption translation, the method achieves <0.5% accuracy degradation while accelerating nearest-neighbor search over brute-force enumeration by over 200×. It demonstrates superior efficiency, robustness against distributional shifts, and inherent interpretability—offering a viable alternative to conventional deep learning in high-stakes applications.
Whether layer-wise activations across large language models (LLMs) with heterogeneous architectures share alignable representational geometry remains unclear. Method: We propose a systematic framework based on nearest-neighbor graphs and high-dimensional geometric similarity measures to enable cross-model layer alignment and depth-normalized comparison across 24 open-source LLMs. Contribution/Results: We discover— for the first time—that activation spaces at matched normalized depths exhibit highly consistent local neighborhood structures, forming robust, layerwise-evolving geometric patterns. Crucially, normalized depth—not absolute layer index—predicts cross-model activation similarity: nearest-neighbor matching accuracy at equivalent depths significantly exceeds both random baselines and cross-depth controls. This reveals an implicit, shared computational pathway across diverse LLMs, establishing a geometric foundation for model alignment, knowledge transfer, and interpretability research.
kNN-MT incurs substantial computational redundancy and high decoding latency due to performing k-nearest-neighbor (kNN) retrieval for every generated token. To address this, we propose Dynamic Skipping Mechanism (DSM), the first learnable, token-level retrieval gating module—implemented as a lightweight MLP—that predicts whether kNN lookup is necessary for each token, thereby executing retrieval only where beneficial. DSM is fully plug-and-play, requiring no architectural modifications to existing kNN-MT systems. Evaluated on standard benchmarks including WMT, DSM reduces kNN retrieval cost by 53% with negligible BLEU degradation (<0.5), while identifying and skipping 67–84% of redundant tokens. This yields significant inference speedup without compromising translation quality. Crucially, DSM is system-agnostic and compatible with all mainstream kNN-MT variants.
This work addresses the challenge of efficiently supporting both vector similarity search and arbitrary attribute filtering in high-dimensional approximate nearest neighbor retrieval. The authors propose a lightweight graph-based indexing algorithm that seamlessly integrates attribute filtering into the graph traversal process, overcoming the efficiency bottlenecks of existing methods when handling unseen query vectors combined with complex attribute constraints. Experimental results on multiple real-world datasets demonstrate that the proposed approach significantly outperforms state-of-the-art techniques, achieving substantially faster query latency while maintaining high recall. The method thus offers a compelling balance among flexibility, efficiency, and scalability for hybrid vector-and-attribute search scenarios.
This work addresses the limitations of Recall@k as the dominant evaluation metric in approximate nearest neighbor (ANN) search, which often overestimates retrieval quality and incurs redundant computation. The authors propose 1/Ratio@k—an inverse approximation ratio that is hyperparameter-free and directly computable from benchmark data—as a more principled alternative. Through extensive evaluation of state-of-the-art ANN algorithms across diverse high-dimensional datasets, coupled with efficiency analysis and validation on downstream tasks such as classification and retrieval-augmented generation, the study demonstrates that optimizing 1/Ratio@k significantly reduces computational overhead while preserving practical utility. Moreover, 1/Ratio@k exhibits substantially stronger correlation with real-world effectiveness—measured by label accuracy and semantic similarity—than Recall@k.
Approximate k-nearest neighbor (k-ANN) search in vector databases often yields redundant, insufficiently diverse, and inflexible results. Method: This paper proposes a progressive diversity optimization framework that incurs no additional indexing overhead. It explicitly incorporates diversity constraints into state-of-the-art k-ANN pipelines via a three-stage mechanism—iterative search, dynamic deduplication, and similarity verification—enabling joint optimization of result size and diversity in a single retrieval pass. Users can flexibly specify both the desired result count and diversity strength. Results: Experiments on million-scale benchmarks (LAION-art, Deep1M, Txt2img) demonstrate that, under medium-to-high diversity settings, the method improves recall and mean similarity significantly while incurring less than 5% latency overhead, closely approaching theoretical optimality. It is the first approach to unify accuracy, efficiency, and controllability in k-ANN search.
This work addresses the challenge of efficient approximate nearest neighbor search and maximum inner product search (MIPS) in high-dimensional embeddings from large language models. The authors propose a novel approach that transforms asymmetric MIPS into Euclidean nearest neighbor search via dimension augmentation, combined with Equi-Voronoi Polytopes (EVP) quantization and a Fast Linear Assignment Sorting (FLAS) one-dimensional pre-sorting mechanism. This integration substantially accelerates k-nearest neighbor graph (kNNG) construction and query processing while enhancing memory access locality and cache efficiency. Evaluated in the SISAP 2026 challenge, the method achieves low-latency, high-recall MIPS performance, significantly outperforming existing baselines.