Score
Design and implement representation-learning systems that produce vector embeddings for discrete labels conditioned on their surrounding context, and build classifiers that use those embeddings to estimate labels from observed contexts. This work includes devising embedding algorithms, training objectives, and inference procedures that capture higher-order label connectivity and contextual co-occurrence for accurate label prediction.
This paper addresses two fundamental challenges in foundation model research: the opaque nature of representation mechanisms and diminishing returns from scaling. To resolve these, we propose the “contexture” theory—a unified characterization of representation learning wherein optimal representations maximize mutual information between inputs and contextual variables, with peak generalization achieved at moderate contextual strength. We establish the first unified mathematical framework proving that scaling bottlenecks stem primarily from contextual *quality*, not scale. We introduce two general-purpose context-aware learning objectives—SVME and KISE—and a multi-context fusion strategy. Leveraging information theory and statistical learning theory, we derive a generalization bound for representation learning, unifying theoretical explanations across supervised, self-supervised, and generative pretraining paradigms. Empirical validation confirms that mainstream pretraining objectives implicitly optimize contexture. Our work provides both theoretical foundations and practical guidelines for designing efficient, context-driven pretraining frameworks.
This work investigates the behavioral limits of in-context learning (ICL) with thousands of demonstrations in ultra-long-context language models. Through systematic experiments across multiple models (e.g., Llama, Qwen) and datasets, and employing controlled analytical techniques—including random shuffling, label-based grouping, and demonstration subsampling—we find that: (1) ICL robustness to input ordering significantly increases with context length; (2) clustering examples by label degrades performance; and (3) gains do not arise from joint encoding of multiple demonstrations. Key contributions include: the first empirical demonstration that ICL performance scales continuously with demonstration count up to several thousand in large-label-space tasks; superior effectiveness over fine-tuning under low-to-moderate data regimes; and non-negligible gains achievable without fully utilizing available context capacity—challenging prevailing assumptions about ICL mechanisms.
The evolution of embedding techniques from word vectors to multimodal representations remains fragmented, lacking a unified framework that integrates advances across linguistic, cross-lingual, personalized, and multimodal domains—particularly for embodied multimodal learning in large language models. Method: We systematically survey static and contextual language representations, cross-lingual and personalized modeling, sentence/document embeddings, and multimodal fusion in vision, robotics, and cognitive science. We synthesize recent progress in interpretability, model compression, numerical encoding, and bias mitigation, and propose a novel paradigm emphasizing strong alignment across non-textual modalities and scalable training. Contributions: We construct a comprehensive knowledge graph of end-to-end embedding technologies—from Word2Vec and BERT to GPT, generative topic models, and multimodal alignment/distillation methods—identifying key technical bottlenecks and ethical challenges. This work delivers the first systematic roadmap for multimodal, embodied learning in foundation models.
This work investigates the learnability of high-dimensional embedding vectors from discrete data, focusing on how sample size, token frequency, and embedding–correlation strength jointly govern estimation accuracy. We propose a low-rank approximate Approximate Message Passing (AMP) algorithm grounded in a correlation–similarity coupled probabilistic model. This marks the first systematic integration of the AMP framework into the theoretical analysis of embedding estimation, enabling rigorous characterization of the phase transition boundary for estimation performance. Leveraging tools from high-dimensional statistical inference and random matrix theory, we derive precise quantitative relationships between embedding estimation error and key problem parameters. Extensive experiments on synthetic data and real-world text tasks validate our theoretical predictions, demonstrating substantial improvements in statistical efficiency and robustness—particularly in high-dimensional, sparse regimes.
This work investigates whether large language models (LLMs) can generalize in-context learning (ICL) capabilities—traditionally confined to discrete text—to continuous vector inputs. To this end, we propose Vector-ICL, a novel paradigm that aligns continuous vectors from multimodal black-box encoders into the LLM’s embedding space via lightweight, learnable projectors, enabling tuning-free, cross-modal vector-level ICL. Our key contribution is demonstrating that standard, language-modeling-pretrained LLMs—combined with modality-agnostic projectors—can achieve zero-shot or few-shot generalization to unseen continuous vectors, bypassing tokenization constraints entirely. Evaluated across eight diverse tasks—including text reconstruction, numerical regression, molecular property prediction, and fMRI decoding—Vector-ICL consistently outperforms conventional few-shot ICL baselines and domain-specific models, substantiating the feasibility and promise of LLMs as universal vector processors.
This study addresses the challenge of capturing high-order label dependencies in multi-label classification by proposing the HyperLabel framework. The method constructs a label hypergraph with samples as hyperedges, providing structural priors that transcend pairwise interactions. Building upon hypergraph neural networks, an encoder-decoder architecture is designed to achieve unified cross-modal fusion of feature and structural information through cross-attention and bidirectional message passing mechanisms. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance across seven benchmark datasets. Notably, it yields substantial improvements of 10.3% and 8.2% in Macro-F1 on the Delicious and Bibtex datasets, respectively. These findings validate the effectiveness of explicitly modeling complex label co-occurrence patterns for advancing multi-label classification performance.
Existing graph neural networks struggle to capture high-order class-label connectivity in heterophilic directed graphs, limiting their node classification performance. To address this, this work proposes the Label Context Classifier (LCC), which explicitly models high-order inter-class dependencies in heterophilic settings by generating label context embeddings through four types of directed walks. LCC incorporates an adaptive weighting mechanism that seamlessly integrates with any base GNN architecture. This approach is the first to explicitly encode high-order label structures under heterophily, achieving significant performance gains over state-of-the-art methods across multiple benchmark datasets and substantially improving node classification accuracy.
本文通过统计框架探讨了token预测如何学习有用的表示,使用softmax预测头组织token嵌入,并通过自一致性原则改进上下文表示,从而提高下游任务性能。
Contextual learning (ICL) in few-shot settings critically depends on labeled in-context examples, limiting its applicability when only minimal annotations (e.g., 1–5) are available alongside abundant unlabeled context. Method: We propose In-Context Semi-Supervised Learning (IC-SSL), a novel paradigm enabling Transformers to implicitly leverage unlabeled context for robust, context-aware representation learning. We formally define and theoretically analyze IC-SSL, and introduce a unified framework integrating context embedding modeling, contrastive representation alignment, and manifold-constrained pseudo-label distillation. Contribution/Results: Our method achieves an average 12.3% improvement over supervised ICL baselines under ultra-low labeling rates, significantly enhancing generalization. It provides both theoretically grounded interpretability for unsupervised context utilization and a practical technical pathway for label-efficient in-context learning.
This work addresses the mismatch between the geometric structure of pretrained embeddings and downstream tasks in practical systems such as digital governance, where label scarcity, domain shift, and the infeasibility of retraining large models degrade nearest-neighbor retrieval performance. To bridge the gap between unsupervised post-processing and fully supervised projection, the authors propose a label-efficient method that leverages limited supervision to softly align embeddings with class prototypes while preserving embedding dimensionality. This approach optimizes local neighborhood structure and significantly enhances similarity-based retrieval and lightweight classification under extremely low-label regimes. Experimental results demonstrate a 25.7% improvement in local neighborhood quality over raw embeddings and outperform strong unsupervised post-processing baselines by more than 21.1%.