multi-level supervision

Designs and implements training regimes, loss functions, and annotation mappings that apply supervision signals at multiple representation levels and stages — for example pixel, feature, token, sentence, structural or intermediate-stage outputs — and can include level-wise constraints such as tying parameter slices or embeddings. Builds models and training pipelines that expose and supervise intermediate outputs (multi-stage or progressive pipelines), and analyzes how per-level supervision affects optimization, hierarchical feature learning, disambiguation of ambiguous inputs, and final task performance.

multi-levelsupervision

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Guiding Through Complexity: What Makes Good Supervision for Hard Reasoning Tasks?

Oct 27, 2024
XH
Xuan He
🏛️ Tsinghua University | University of California, Los Angeles

This work investigates how to effectively enhance large language models’ performance on high-difficulty reasoning tasks—such as mathematical proof generation and complex logical deduction—using weak supervision sources (e.g., non-expert annotators or off-the-shelf AI systems). Addressing the trade-off between supervision quality and task difficulty, we establish a key empirical finding: supervision for hard tasks with high step-level error rates yields superior model performance compared to error-free supervision on easy tasks; critically, step-level error rate proves a more informative training signal than final-answer accuracy. Building on this insight, we propose a collaborative supervision paradigm that jointly leverages subtasks and hard tasks, integrated with supervised data mixing, dynamic reweighting, and empirically grounded evaluation. Our approach achieves up to 30% absolute accuracy gains on challenging benchmarks such as MATH. All code and datasets are publicly released.

Effective supervision by weak teacher modelsImpact of step-wise error ratesImproving LLMs on hard reasoning tasks

This work addresses a key limitation of conventional large language model training, which relies on token-level next-token prediction and consequently struggles to distinguish between semantically equivalent expressions that differ in surface form, leading to a bias toward superficial patterns rather than deep semantic understanding. To overcome this, the authors propose elevating the training objective to the conceptual level by introducing a systematic concept-level supervision signal through a concept mapping framework—e.g., unifying surface variants like “mom” and “mother” under a shared concept such as MOTHER. Their approach integrates a concept alignment loss with a multi-surface aggregation strategy, encouraging the model to prioritize semantic correctness. Experiments demonstrate that the resulting concept-aware models achieve lower perplexity, superior performance across multiple NLP benchmarks, and enhanced robustness in domain transfer scenarios.

concept-level traininglarge language modelsnext-token prediction

DiagrammaticLearning: A Graphical Language for Compositional Training Regimes

Jan 02, 2025
ML
Mason Lary
🏛️ University at Buffalo | University of Florida | Harvard University | Air Force Research Laboratory

Modern deep learning training pipelines involve heterogeneous components—such as multi-task heads, distillation objectives, or multimodal encoders—yet lack a unified formalism for modeling inter-component dependencies and enabling holistic optimization. Method: We propose *Learning Diagrams*, the first categorical framework for declaratively specifying training workflows as composable, compiler-ready graphical structures. It enables constraint-driven component composition and automatic synthesis of joint loss functions that enforce predictive consistency across submodels. Contribution/Results: Learning Diagrams unifies diverse paradigms—including few-shot multi-task learning, knowledge distillation, and multimodal learning—under a single abstraction, supporting dynamic in-training and post-hoc reconfiguration. Integrated with PyTorch and Flux.jl, our open-source implementation includes a graph compiler and demonstrates cross-paradigm expressivity and compositional modeling efficacy across canonical benchmarks. Empirical results show substantial improvements in systematicity, modularity, and interpretability of complex model construction.

Deep LearningModel Components InteractionSystematic Representation

Layer by Layer: Uncovering Hidden Representations in Language Models

Feb 04, 2025
OS
Oscar Skean
🏛️ University of Kentucky | Mila | University of Montreal | New York University | University of California, Los Angeles | Meta | Wand.AI

This work challenges the conventional assumption that final-layer representations in large language models (LLMs) are optimal, revealing instead that intermediate-layer hidden states encode richer and more robust semantic information. Method: We propose the first multidimensional representation quality evaluation framework integrating information-theoretic measures (mutual information, compression ratio), manifold geometry, and perturbation invariance—designed for cross-architectural (Transformer/SSM) and cross-modal (text/vision) validation. Contribution/Results: Evaluated on 32 text embedding benchmarks, intermediate-layer embeddings consistently outperform final-layer counterparts by an average of 4.2%, demonstrating both statistical consistency and strong generalization across tasks and architectures. This study provides the first empirical evidence establishing the superiority of intermediate-layer representations, thereby introducing a new paradigm for efficient representation extraction, model compression, and interpretability research.

Analyzing hidden representations in intermediate layers of language modelsDemonstrating mid-layer embeddings outperform final-layer in various tasksProposing metrics to quantify representation quality in model layers

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

Latest Papers

What's happening recently
View more

This study systematically investigates whether representation collapse during multi-stage post-training of large language models leads to degraded adaptability, weakened out-of-distribution generalization, and deteriorated calibration. To this end, the authors construct a comprehensive measurement framework encompassing hidden states, logits, token trajectories, and LoRA updates, thereby establishing— for the first time—the causal relationship between representation collapse and declines in model plasticity, generalization, and calibration. Building on these insights, they propose lightweight intervention strategies, including mixed-domain replay, feature refreshing, representation diversity regularization, and decorrelation of LoRA updates. These methods significantly enhance continual learning capability while preserving behavioral gains, effectively mitigating representation collapse.

adaptation plasticitylarge language modelsrepresentation collapse

Current post-training of language models relies on abstract scalar rewards, which lack transparency regarding the instructional content of preference data and can lead models to learn spurious correlations, resulting in undesirable behaviors such as excessive stylization or sycophancy. This work proposes a data-centric post-training framework that, for the first time, leverages interpretability methods to explicitly model latent conceptual signals within preference data. By analyzing and identifying key features that distinguish preferred from non-preferred responses prior to optimization, the approach integrates interpretability protocols, statistical hypothesis testing, and fine-grained interventions at both feature and data levels. This enables effective diagnosis and suppression of harmful learning signals, significantly reducing off-target behaviors across multiple benchmarks while enhancing model safety and controllable personality.

interpretabilitylearning signalpost-training

This study investigates the effectiveness boundaries of on-policy distillation in training reasoning models, with a focus on critical factors such as teacher selection, self-distillation context, and token-wise optimal policies. To this end, the authors propose a training-free diagnostic framework that, for the first time, quantifies at the token level the alignment between distillation signals and ideal gradient directions. This is achieved through a combination of ideal node gradient derivation and a scalable directional rollout algorithm to efficiently compute gradient alignment scores (measured via cosine similarity). The analysis reveals that distillation signals are more informative when the student errs but may introduce noise along correct reasoning paths. Furthermore, the optimal distillation configuration is highly dependent on both student capability and task characteristics, with no universally optimal setting.

on-policy distillationper-token supervisionreasoning models

This work investigates how neural networks internalize explicit computational procedures into their weights through chain-of-thought (CoT) supervision and examines the implications of this internalization for out-of-distribution generalization. Focusing on the parity learning task, the study provides the first provable analysis of internalization: a single-layer simplified Transformer, guided by CoT, first learns the target function and subsequently internalizes it autoregressively as CoT tokens are gradually removed, enabling direct computation of the parity function. Both theoretical and empirical results demonstrate that the task is nearly impossible to learn from data alone without CoT supervision. While internalization preserves in-distribution performance, it substantially degrades out-of-distribution generalization, revealing an intrinsic trade-off between internalization and robustness to distributional shift.

chain-of-thoughtinternalizationout-of-distribution performance

Hot Scholars

YF

Yanwei Fu

Fudan University
Computer visionmachine learningMultimedia
LL

Lijian Lin

Tencent ARC Lab
Computer VisionVisual Tracking,Video Object Detection
XQ

Xiaojuan Qi

Assistant Professor, The University of Hong Kong
3D VisionDeep learningArtificial IntelligenceMedical Image Analysis
JC

Jong Chul Ye

Professor, Chung Moon Soul Chair, Graduate School of AI, KAIST
machine learningcomputational imagingmedical imagingsignal processing
KJ

Kaixun Jiang

Fudan University
Computer VisionAdversarial Examples