masked pair encoder

Designs, implements, or evaluates an encoder that jointly represents example–query pairs by masking portions of the pair and training with masked-language-style objectives so the model can bidirectionally encode both example and query contexts. This includes architectures and training procedures for masked pair encoding that enable in-context prediction from masked pretraining and allow empirical comparison to causal-transformer in-context learning performance.

maskedpairencoder

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work investigates the in-context learning (ICL) capabilities of masked language models and establishes, for the first time, a unified theoretical framework that encompasses both masked and causal language models. By introducing a statistical learning approach based on empirical measures and incorporating Wasserstein-type regularization conditions with excess risk analysis, the study provides a unified characterization of the generalization behavior of these two pretraining objectives under task distribution shift. Theoretically, it proves that both architectures achieve comparable performance upper bounds in k-shot ICL settings. Empirically, Masked Pretrained Encoders (MPE) demonstrate function-learning performance on par with GPT-2–style causal Transformers. Furthermore, the paper proposes an optimal task allocation strategy under data budget constraints, offering principled guidance for efficient ICL.

causal language modelsexcess riskin-context learning

Contextures: The Mechanism of Representation Learning

Apr 28, 2025
RZ
Runtian Zhai
🏛️ Carnegie Mellon University

This paper addresses two fundamental challenges in foundation model research: the opaque nature of representation mechanisms and diminishing returns from scaling. To resolve these, we propose the “contexture” theory—a unified characterization of representation learning wherein optimal representations maximize mutual information between inputs and contextual variables, with peak generalization achieved at moderate contextual strength. We establish the first unified mathematical framework proving that scaling bottlenecks stem primarily from contextual *quality*, not scale. We introduce two general-purpose context-aware learning objectives—SVME and KISE—and a multi-context fusion strategy. Leveraging information theory and statistical learning theory, we derive a generalization bound for representation learning, unifying theoretical explanations across supervised, self-supervised, and generative pretraining paradigms. Empirical validation confirms that mainstream pretraining objectives implicitly optimize contexture. Our work provides both theoretical foundations and practical guidelines for designing efficient, context-driven pretraining frameworks.

Developing unified theory for diverse pretraining methodsOptimizing context variables to improve model performanceUnderstanding representation learning mechanisms in foundation models

Masked Mixers for Language Generation and Retrieval

Sep 02, 2024
BL
Benjamin L. Badger
🏛️ Guidehouse

This paper identifies a critical limitation of Transformer self-attention: its selective focus inherently discards substantial input information—particularly detrimental to retrieval tasks, which require high-fidelity, near-bijective representations. To address this, we propose Masked Mixers, which replace self-attention with masked causal convolutions. We provide the first information-theoretic analysis revealing attention’s fundamental information loss and introduce *global invertibility* as a key criterion for representation fidelity and training efficiency. Experiments demonstrate that Masked Mixers train faster and yield more accurate representations than Transformers on small-context (<512 tokens) generative tasks. Moreover, lightweight Masked Mixers significantly outperform large state-of-the-art Transformer models on cross-scale retrieval benchmarks—validating their dual superiority in language modeling and embedding learning.

Addresses information loss in transformers due to attention mechanisms.Demonstrates masked mixers' superiority in retrieval tasks over transformers.Proposes masked mixers to improve input representation accuracy.

Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context Learning

Dec 28, 2023
CQ
Chengwei Qin
🏛️ Nanyang Technological University | Princeton University | Salesforce Research

Existing knowledge distillation methods for in-context learning (ICL) with small language models focus solely on aligning student outputs with teacher predictions, neglecting the teacher’s input example preferences—thereby limiting generalization. To address this, we propose BiAlign, the first framework that jointly models and optimizes student imitation of teacher behavior along *both* input example preference and output distribution dimensions. Methodologically, BiAlign integrates token-level probabilistic distillation with a novel ranking loss to explicitly capture differential sensitivity to demonstration examples. Evaluated across language understanding, reasoning, and code generation tasks, BiAlign consistently outperforms state-of-the-art distillation approaches—achieving superior robustness and ICL transfer efficiency with fewer parameters and lower computational overhead. This work establishes a new paradigm for enhancing the ICL capability of compact models through bidirectional behavioral alignment.

Aligning input preferences between student and teacher modelsImproving in-context learning abilities of compact modelsReducing computational demands while maintaining performance

In-Context Learning with Long-Context Models: An In-Depth Exploration

Apr 30, 2024
AB
Amanda Bertsch
🏛️ Carnegie Mellon University | Tel Aviv University

This work investigates the behavioral limits of in-context learning (ICL) with thousands of demonstrations in ultra-long-context language models. Through systematic experiments across multiple models (e.g., Llama, Qwen) and datasets, and employing controlled analytical techniques—including random shuffling, label-based grouping, and demonstration subsampling—we find that: (1) ICL robustness to input ordering significantly increases with context length; (2) clustering examples by label degrades performance; and (3) gains do not arise from joint encoding of multiple demonstrations. Key contributions include: the first empirical demonstration that ICL performance scales continuously with demonstration count up to several thousand in large-label-space tasks; superior effectiveness over fine-tuning under low-to-moderate data regimes; and non-negligible gains achievable without fully utilizing available context capacity—challenging prevailing assumptions about ICL mechanisms.

Compares in-context learning with example retrieval and finetuning.Explores in-context learning with long-context models.Investigates properties of in-context learning and long-context models.

Latest Papers

What's happening recently
View more

This study investigates whether encoder-decoder models possess an intrinsic capacity for automatic speech recognition (ASR) context adaptation. Through controlled experiments across six architectures, the authors compare organized and interleaved demonstration paradigms while decoupling the contributions of lexical and speaker information. The findings confirm that context adaptation constitutes a general emergent capability of such models, functioning effectively out-of-the-box without fine-tuning. Moreover, organized demonstrations are shown to be more stable and effective than their interleaved counterparts. The proposed approach achieves consistent context adaptation performance, yielding relative improvements of up to 30% and 23% under oracle and initial hypothesis settings, respectively.

Automatic speech recognitionDemonstrationEncoder-decoder models

This study addresses the limited understanding of how few-shot prompting drives in-context learning through function vectors (FVs). The authors propose a causal decomposition framework that, for the first time, expresses FVs as linear combinations of example-specific subvectors, revealing a unified mechanism of additive superposition and context-adaptive reweighting. Leveraging techniques such as causal intervention, separation of Query-Key and Value pathways, and attention tracing, they empirically validate the effectiveness of this additive approximation across multiple tasks and models. Their analysis further identifies Query-Key alignment as a critical factor for enhancing FV quality, particularly in ambiguous scenarios where it substantially improves model performance.

attention reweightingcausal mechanismfew-shot learning

This work addresses the inherent conflict between in-context learning (ICL) and in-weight learning (IWL) in Transformer models by proposing CoQE, a dual-representation architecture that decouples contextual information and input samples at the representation level. CoQE explicitly encodes context into a task representation space and samples into a distinct sample representation space, leveraging dual-space linear modeling grounded in duality theory and an enhanced Transformer structure. Experiments on few-shot classification and pseudo-arithmetic tasks demonstrate that CoQE significantly improves ICL performance while effectively preserving IWL capabilities, thereby validating the efficacy and generality of its co-optimization strategy for both learning paradigms.

In-Context LearningIn-Weight LearningLearning Conflict

To address the insufficient robustness of oversampled baseband signal demodulation under impulsive noise channels, this work challenges the conventional paradigm that treats intersymbol contribution (ISC) as interference, instead modeling ISC as deterministic contextual information embedded within the waveform. We propose a Transformer-based masked symbol modeling framework: complex-valued baseband sequences undergo random masking and reconstruction pretraining, leveraging bidirectional self-attention to capture long-range waveform dependencies and learn implicit “waveform syntax.” This enables semantic-level contextual reasoning over physical-layer signals, allowing the receiver to infer severely distorted symbol segments from neighboring samples under impulsive noise corruption. Experimental results demonstrate that the proposed context-aware demodulator significantly improves bit error rate performance, validating the feasibility and effectiveness of transforming ISC into structured prior knowledge.

Applying masked symbol modeling to learn waveform syntax for robust demodulationDemodulating oversampled baseband signals in impulsive noise channelsLeveraging inter-symbol overlap as contextual information for signal interpretation

Existing theories of in-context learning typically assume that demonstration examples in prompts are independent and identically distributed, thereby overlooking the pervasive temporal correlations present in real-world sequences. This work develops a theoretically tractable model based on linear attention, integrating a linear regression theory sandbox with an actual Transformer architecture to systematically investigate how temporal dependencies within prompts affect in-context learning. We reveal for the first time that such temporal correlations induce an “effective context length,” rendering correlated prompts equivalent to shorter i.i.d. ones. Moreover, we find that when queries are temporally aligned with the context, Softmax attention substantially outperforms linear attention, highlighting the critical importance of aligning attention mechanisms with task-specific structural properties.

architectural mismatchattention mechanismseffective context length

Hot Scholars

XW

Xia Wang

Research Assistant, Vanderbilt University
autonomous vehiclescyber physical systemscomputer visionlarge language model