Score
Designs, implements, or evaluates an encoder that jointly represents example–query pairs by masking portions of the pair and training with masked-language-style objectives so the model can bidirectionally encode both example and query contexts. This includes architectures and training procedures for masked pair encoding that enable in-context prediction from masked pretraining and allow empirical comparison to causal-transformer in-context learning performance.
This work investigates the in-context learning (ICL) capabilities of masked language models and establishes, for the first time, a unified theoretical framework that encompasses both masked and causal language models. By introducing a statistical learning approach based on empirical measures and incorporating Wasserstein-type regularization conditions with excess risk analysis, the study provides a unified characterization of the generalization behavior of these two pretraining objectives under task distribution shift. Theoretically, it proves that both architectures achieve comparable performance upper bounds in k-shot ICL settings. Empirically, Masked Pretrained Encoders (MPE) demonstrate function-learning performance on par with GPT-2–style causal Transformers. Furthermore, the paper proposes an optimal task allocation strategy under data budget constraints, offering principled guidance for efficient ICL.
This paper addresses two fundamental challenges in foundation model research: the opaque nature of representation mechanisms and diminishing returns from scaling. To resolve these, we propose the “contexture” theory—a unified characterization of representation learning wherein optimal representations maximize mutual information between inputs and contextual variables, with peak generalization achieved at moderate contextual strength. We establish the first unified mathematical framework proving that scaling bottlenecks stem primarily from contextual *quality*, not scale. We introduce two general-purpose context-aware learning objectives—SVME and KISE—and a multi-context fusion strategy. Leveraging information theory and statistical learning theory, we derive a generalization bound for representation learning, unifying theoretical explanations across supervised, self-supervised, and generative pretraining paradigms. Empirical validation confirms that mainstream pretraining objectives implicitly optimize contexture. Our work provides both theoretical foundations and practical guidelines for designing efficient, context-driven pretraining frameworks.
This paper identifies a critical limitation of Transformer self-attention: its selective focus inherently discards substantial input information—particularly detrimental to retrieval tasks, which require high-fidelity, near-bijective representations. To address this, we propose Masked Mixers, which replace self-attention with masked causal convolutions. We provide the first information-theoretic analysis revealing attention’s fundamental information loss and introduce *global invertibility* as a key criterion for representation fidelity and training efficiency. Experiments demonstrate that Masked Mixers train faster and yield more accurate representations than Transformers on small-context (<512 tokens) generative tasks. Moreover, lightweight Masked Mixers significantly outperform large state-of-the-art Transformer models on cross-scale retrieval benchmarks—validating their dual superiority in language modeling and embedding learning.
Existing knowledge distillation methods for in-context learning (ICL) with small language models focus solely on aligning student outputs with teacher predictions, neglecting the teacher’s input example preferences—thereby limiting generalization. To address this, we propose BiAlign, the first framework that jointly models and optimizes student imitation of teacher behavior along *both* input example preference and output distribution dimensions. Methodologically, BiAlign integrates token-level probabilistic distillation with a novel ranking loss to explicitly capture differential sensitivity to demonstration examples. Evaluated across language understanding, reasoning, and code generation tasks, BiAlign consistently outperforms state-of-the-art distillation approaches—achieving superior robustness and ICL transfer efficiency with fewer parameters and lower computational overhead. This work establishes a new paradigm for enhancing the ICL capability of compact models through bidirectional behavioral alignment.
This work investigates the behavioral limits of in-context learning (ICL) with thousands of demonstrations in ultra-long-context language models. Through systematic experiments across multiple models (e.g., Llama, Qwen) and datasets, and employing controlled analytical techniques—including random shuffling, label-based grouping, and demonstration subsampling—we find that: (1) ICL robustness to input ordering significantly increases with context length; (2) clustering examples by label degrades performance; and (3) gains do not arise from joint encoding of multiple demonstrations. Key contributions include: the first empirical demonstration that ICL performance scales continuously with demonstration count up to several thousand in large-label-space tasks; superior effectiveness over fine-tuning under low-to-moderate data regimes; and non-negligible gains achievable without fully utilizing available context capacity—challenging prevailing assumptions about ICL mechanisms.
This study investigates whether encoder-decoder models possess an intrinsic capacity for automatic speech recognition (ASR) context adaptation. Through controlled experiments across six architectures, the authors compare organized and interleaved demonstration paradigms while decoupling the contributions of lexical and speaker information. The findings confirm that context adaptation constitutes a general emergent capability of such models, functioning effectively out-of-the-box without fine-tuning. Moreover, organized demonstrations are shown to be more stable and effective than their interleaved counterparts. The proposed approach achieves consistent context adaptation performance, yielding relative improvements of up to 30% and 23% under oracle and initial hypothesis settings, respectively.
This study addresses the limited understanding of how few-shot prompting drives in-context learning through function vectors (FVs). The authors propose a causal decomposition framework that, for the first time, expresses FVs as linear combinations of example-specific subvectors, revealing a unified mechanism of additive superposition and context-adaptive reweighting. Leveraging techniques such as causal intervention, separation of Query-Key and Value pathways, and attention tracing, they empirically validate the effectiveness of this additive approximation across multiple tasks and models. Their analysis further identifies Query-Key alignment as a critical factor for enhancing FV quality, particularly in ambiguous scenarios where it substantially improves model performance.
This work addresses the inherent conflict between in-context learning (ICL) and in-weight learning (IWL) in Transformer models by proposing CoQE, a dual-representation architecture that decouples contextual information and input samples at the representation level. CoQE explicitly encodes context into a task representation space and samples into a distinct sample representation space, leveraging dual-space linear modeling grounded in duality theory and an enhanced Transformer structure. Experiments on few-shot classification and pseudo-arithmetic tasks demonstrate that CoQE significantly improves ICL performance while effectively preserving IWL capabilities, thereby validating the efficacy and generality of its co-optimization strategy for both learning paradigms.
To address the insufficient robustness of oversampled baseband signal demodulation under impulsive noise channels, this work challenges the conventional paradigm that treats intersymbol contribution (ISC) as interference, instead modeling ISC as deterministic contextual information embedded within the waveform. We propose a Transformer-based masked symbol modeling framework: complex-valued baseband sequences undergo random masking and reconstruction pretraining, leveraging bidirectional self-attention to capture long-range waveform dependencies and learn implicit “waveform syntax.” This enables semantic-level contextual reasoning over physical-layer signals, allowing the receiver to infer severely distorted symbol segments from neighboring samples under impulsive noise corruption. Experimental results demonstrate that the proposed context-aware demodulator significantly improves bit error rate performance, validating the feasibility and effectiveness of transforming ISC into structured prior knowledge.
Existing theories of in-context learning typically assume that demonstration examples in prompts are independent and identically distributed, thereby overlooking the pervasive temporal correlations present in real-world sequences. This work develops a theoretically tractable model based on linear attention, integrating a linear regression theory sandbox with an actual Transformer architecture to systematically investigate how temporal dependencies within prompts affect in-context learning. We reveal for the first time that such temporal correlations induce an “effective context length,” rendering correlated prompts equivalent to shorter i.i.d. ones. Moreover, we find that when queries are temporally aligned with the context, Softmax attention substantially outperforms linear attention, highlighting the critical importance of aligning attention mechanisms with task-specific structural properties.