Efficient Latent Semantic Clustering for Scaling Test-Time Computation of LLMs

📅 2025-05-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address semantic redundancy in large language model (LLM) inference-time computation scaling, this paper proposes an external-model-free implicit semantic clustering method. The approach constructs a lightweight, context-aware semantic similarity metric directly from intermediate hidden states of the generator LLM, and integrates dynamic thresholding with hierarchical clustering for end-to-end semantic consistency modeling. Its core contribution lies in the first use of LLM internal hidden representations—without any auxiliary embedding models or post-hoc modules—as the sole basis for semantic clustering. Evaluated across multiple LLMs and diverse benchmark tasks, the method achieves an average 2.3× reduction in inference-time computational cost while matching or surpassing state-of-the-art methods in both clustering accuracy and uncertainty calibration.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Search and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Scaling test-time computation--generating and analyzing multiple or sequential outputs for a single input--has become a promising strategy for improving the reliability and quality of large language models (LLMs), as evidenced by advances in uncertainty quantification and multi-step reasoning. A key shared component is semantic clustering, which groups outputs that differ in form but convey the same meaning. Semantic clustering enables estimation of the distribution over the semantics of outputs and helps avoid redundant exploration of reasoning paths. However, existing approaches typically rely on external models, which introduce substantial computational overhead and often fail to capture context-aware semantics. We propose Latent Semantic Clustering (LSC), a lightweight and context-sensitive method that leverages the generator LLM's internal hidden states for clustering, eliminating the need for external models. Our extensive experiment across various LLMs and datasets shows that LSC significantly improves the computational efficiency of test-time scaling while maintaining or exceeding the performance of existing methods.
Problem

Research questions and friction points this paper is trying to address.

Scaling test-time computation for LLMs' reliability and quality
Semantic clustering without external models' computational overhead
Improving efficiency of output distribution estimation in LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses hidden states for semantic clustering
Eliminates need for external models
Improves computational efficiency significantly
🔎 Similar Papers
2024-09-30arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
S
Sungjae Lee
Department of Computer Science and Engineering, POSTECH, South Korea
H
Hoyoung Kim
Graduate School of Artificial Intelligence, POSTECH, South Korea
Jeongyeon Hwang
Jeongyeon Hwang
Ph.D student at POSTECH
Machine Learning
Eunhyeok Park
Eunhyeok Park
POSTECH
neural network optimizationenergy efficient hardware design
Jungseul Ok
Jungseul Ok
Associate Professor, CSE/AI, POSTECH
Reinforcement LearningMachine Learning