response-only loss masking

Designs and implements loss-masking schemes and training/evaluation pipelines that exclude context or prompt tokens so that only model-generated response tokens contribute to the computed loss and resulting gradients. This competence includes creating token-alignment and mask representations, integrating those masks into loss computation and optimization, and analyzing how response-only masking affects gradient flow, evaluation metrics, and prediction accuracy.

response-onlylossmasking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Token Masking Improves Transformer-Based Text Classification

May 16, 2025
XX
Xianglong Xu
🏛️ University of Pittsburgh

To address overfitting in Transformer-based text classification models, this paper proposes a lightweight stochastic token masking regularization technique: during training, each input token is replaced with [MASK] with probability $p = 0.1$, introducing controllable input perturbation. Theoretically, we first model this method as implicit gradient averaging and implicit ensemble, revealing its dual mechanism—simultaneously mitigating overfitting and strengthening inter-token dependencies. Crucially, the approach requires no architectural modifications or additional parameters. Extensive cross-model validation on mBERT, Qwen2.5-0.5B, and TinyLlama-1.1B demonstrates consistent superiority over standard regularizers (e.g., Dropout, Label Smoothing) across language identification and sentiment analysis tasks, yielding average accuracy gains of 1.2–2.8 percentage points. The method exhibits strong generalizability and robustness across diverse model scales and linguistic tasks.

Enhancing transformer text classification via token maskingOptimizing masking rates for diverse NLP tasksReducing overfitting through stochastic input perturbations

While existing masked diffusion language models possess bidirectional denoising capabilities, they struggle to effectively support prompt infilling, limiting their utility in prompt engineering. This work identifies this limitation as stemming from the training paradigm rather than the model architecture itself and proposes a simple yet effective full-sequence joint masking supervised fine-tuning approach—simultaneously masking both prompts and responses—to unlock the model’s prompt infilling capacity. Without any architectural modifications, the method enables the model to automatically generate high-quality prompt templates from only a few examples. It demonstrates strong transferability across diverse downstream tasks and model variants, achieving performance on par with or even surpassing manually crafted prompts, while remaining complementary to existing prompt optimization techniques.

diffusion language modelsmasked language modelingprompt engineering

Morphing Tokens Draw Strong Masked Image Models

Dec 30, 2023
TK
Taekyung Kim
🏛️ NAVER AI Lab

In masked image modeling (MIM) pretraining of Vision Transformers (ViTs), tokenization and local masking induce spatially inconsistent reconstruction supervision, degrading representation discriminability. This work is the first to systematically identify and address this spatial inconsistency issue, proposing Dynamic Token Morphing (DTM): a context-aware, dynamic token aggregation mechanism that generates spatially coherent reconstruction targets. DTM introduces no additional parameters or computational overhead and is plug-and-play across diverse MIM frameworks. On ImageNet-1K and ADE20K, DTM achieves significant gains over state-of-the-art MIM methods—yielding lower training loss and more stable convergence. When transferred to downstream tasks such as iNaturalist, it delivers consistent performance improvements. The core contribution is the first lightweight, parameter-free, and framework-agnostic solution specifically designed to resolve the spatial inconsistency problem in MIM.

Addresses spatial inconsistency in masked image modeling supervisionImproves discriminative representation learning in Vision TransformersReduces training costs while enhancing MIM performance

Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training

Jun 12, 2023
LB
L. Baraldi
🏛️ University of Modena and Reggio Emilia | NVIDIA AI Technology Center | IIT-CNR

To address input noise and pretraining-finetuning mismatch caused by masked modeling in vision transformer pretraining, this paper proposes MaPeT. Methodologically, MaPeT jointly models structural dependencies among image patches via autoregressive masking and random block permutation—eliminating distributional shift induced by conventional random masking. It further introduces auxiliary positional embeddings to mitigate positional information inconsistency between pretraining and finetuning. Additionally, we design a k-CLIP visual tokenizer that maps image patches to discrete CLIP-aligned semantic tokens. Experiments demonstrate that MaPeT achieves state-of-the-art (SOTA) performance on ImageNet among models of comparable parameter count. The code and pretrained models are publicly released.

Image UnderstandingPre-trainingVisual Transformers

Masking Augmentation for Supervised Learning

Jun 20, 2023
BH
Byeongho Heo
🏛️ NAVER AI Lab

In supervised learning, strong random masking augmentation often induces training instability. To address this, we propose Masked Sub-model (MaskSub): a dual-model framework where the main model undergoes standard supervised training, while a dedicated sub-model handles masked inputs; a class-wise self-distillation with relaxed loss is introduced to mitigate optimization oscillations. This is the first method to stably enable strong masking augmentation in purely supervised settings—without requiring contrastive objectives or reconstruction targets. Our core innovations are the synergistic dual-model architecture and the relaxed distillation loss, which jointly balance augmentation strength and training stability. Extensive experiments across diverse architectures—including DeiT-III, MAE, CLIP, ResNet, and Swin—and multiple training paradigms demonstrate consistent performance gains, accelerated loss convergence, and superiority over state-of-the-art baselines. The code is publicly available.

Adopting masking augmentations in supervised learningImproving performance across diverse models and scenariosMitigating unstable training with masking augmentations

Latest Papers

What's happening recently
View more

This study addresses the high sensitivity of Joint Embedding Predictive Architectures (JEPA) to masking strategies and the absence of theoretical explanations for the performance disparity between block-wise and scattered masking. By introducing a wavelet-basis linear measurement perspective, this work reveals that mask geometry effectively prevents representational collapse in the target encoder by preserving irrecoverable coarse-scale information. Validated through 151 pretraining runs, the proposed theory quantifies the mechanism of irrecoverable content, confirming that block-wise masking significantly outperforms random and stripe masking. Furthermore, we demonstrate that freezing the target encoder narrows the performance gap across masking strategies, whereas exposing partial targets substantially enhances representation fidelity.

Joint-embedding predictive architectureMasking strategyMeasurement theory

This study addresses the limitation that existing encoded prompt injection evaluations focus solely on harmful responses, where high refusal rates obscure models' inability to distinguish benign from malicious requests, thereby fostering a false sense of security. We present the first systematic evaluation of benign responses under encoding by introducing a benign-arm control to quantify the degradation of the harm gap. Employing homograph encoding and multi-model benchmarking across SFT, DPO, and RLVR training pipelines, we conduct controlled experiments to disentangle the independent effects of protocol-level versus character-level transformations. Our findings confirm that encoding destroys rather than enhances the harm gap. We identify twelve evaluation deficiencies and reveal that certain vulnerabilities originate at the protocol layer. Ultimately, we demonstrate that corrected metrics more faithfully reflect model robustness against encoded adversarial prompts.

benign armencoded-prompt evaluationharm gap

This study addresses the unpredictable robustness of CLIP under mask pruning by identifying "spurious inversion" as the critical factor underlying unstable masking performance. To this end, we introduce the Spurious Inversion Metric (SIM) and propose a label-free, pre-deployment diagnostic framework that integrates semantic masking, text similarity analysis, and asynchronous GPU batch partitioning for efficient evaluation. Experimental results demonstrate that SIM significantly predicts masking efficacy, enabling optimized models to match or surpass baseline performance. This work thereby offers a reliable solution for the robust deployment of compressed CLIP architectures.

Compressed CLIPMasking-based token pruningPre-deployment diagnostic

本文针对JEPA训练效率低的问题,提出了一种名为M-JEPA的执行架构,通过分离与掩码相关的计算和路由,减少了计算、内存流量和同步开销,提高了训练速度。

GPU efficiencyJEPAmask-specific branches

This study investigates whether the performance gains of mask-replace training in zero-shot text-to-speech (TTS) stem solely from self-correction capabilities during inference. To address this, we propose DeMaR, a model integrating mask-replace training with confidence-based sampling. By disabling inference-time revisions on the LibriTTS dataset, we demonstrate that the observed improvements primarily originate from noise context augmentation and replacement supervision signals during training, rather than relying on token revision mechanisms at inference. Experimental results indicate that, under identical conditions, DeMaR achieves significantly lower word error rates compared to both autoregressive and pure mask-diffusion baselines. Crucially, this advantage persists even when inference-time revisions are disabled. These findings offer new insights into the training dynamics of discrete diffusion models for TTS.

discrete diffusion modelsmask-and-replace trainingself-correction

Hot Scholars

YG

Yang Guan

Software Engineer, Google Inc.
Networks
HZ

Hongyu Zhang

Chongqing University
Software EngineeringMining Software RepositoriesData-driven Software EngineeringSoftware Analytics
MS

Minjoon Seo

Config Intelligence; KAIST
Artificial IntelligenceLanguage Modeling
ZZ

Zibin Zheng

IEEE Fellow, Highly Cited Researcher, Sun Yat-sen University, China
BlockchainSmart ContractServices ComputingSoftware Reliability
XJ

Xiaogang Jin

Professor of the State Key Lab of CAD&CG, Zhejiang University
Computer AnimationComputer GraphicsVirtual RealityDigital Fashion