intermediate supervision

Designing intermediate losses, gating mechanisms, and progressive curricula so each step of a multi-step pipeline is supervised to remove a specific degradation factor, enabling structured acquisition of subskills and improved end-to-end performance.

intermediatesupervision

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Towards Scalable Exact Machine Unlearning Using Parameter-Efficient Fine-Tuning

Jun 24, 2024
SB
Somnath Basu Roy Chowdhury
🏛️ UNC Chapel Hill | Google DeepMind | Columbia University | Independent | Google Research

This paper addresses the challenges of “exact removal of specific training samples” in machine unlearning—namely, high computational overhead from retraining, significant latency, and degradation in model performance. To this end, we propose the Sequence-Aware Sharded and Stratified Training (S3T) framework. S3T employs hierarchical sequential training, disjoint partitioning of data subsets, and layer-wise parameter isolation, enabling theoretically rigorous, zero-loss exact unlearning via deactivation of only affected layers. It is the first method to support high-concurrency deletion requests while guaranteeing zero service interruption. Integrated with parameter-efficient fine-tuning (PEFT) and multi-sequence joint optimization, S3T achieves substantial improvements across multiple benchmarks: 92% reduction in deletion latency, <0.3% accuracy loss, 100% service availability, and formal theoretical guarantees on deletion equivalence and performance consistency.

Efficiently remove data influence without full retraining.Enhance deletion capabilities while maintaining model performance.Minimize system downtime during model component retraining.

Step-Opt: Boosting Optimization Modeling in LLMs through Iterative Data Synthesis and Structured Validation

Jun 21, 2025
YW
Yang Wu
🏛️ C2DL | Institute of Automation | Chinese Academy of Sciences | School of Artificial Intelligence | University of Chinese Academy of Sciences | Dalian Minzu University

Large language models (LLMs) struggle with complex problem comprehension and precise mathematical modeling in operations research (OR) optimization tasks. To address this, we propose Step-Opt-Instruct—a novel framework that iteratively generates OR problems of progressively increasing complexity and incorporates a structured, stepwise validation mechanism to effectively prevent error propagation and enhance synthetic data quality. We apply supervised fine-tuning using this framework on LLaMA-3-8B and Mistral-7B. Experimental results demonstrate state-of-the-art performance across three major benchmarks—NL4OPT, MAMO, and IndustryOR—with a 17.01% improvement in micro-averaged accuracy on complex problems. The approach significantly strengthens generalization capability for multi-constraint, multi-objective decision-making tasks, advancing the frontier of natural-language-to-optimization modeling.

Enhancing LLMs for complex optimization modeling tasksGenerating high-quality data via iterative synthesis and validationImproving accuracy in Operations Research problem-solving

Small-object detection performance is hindered by fragmented optimization across stages in conventional pipeline-based detectors. To address this, we propose PLUSNet, an end-to-end co-optimization framework introducing the novel “Purify–Label–Utilize” paradigm. Specifically, we design a hierarchical feature purifier to suppress noise; develop a multi-criterion dynamic label assignment mechanism to improve positive/negative sample quality; and introduce a frequency-domain decoupled detection head for fine-grained feature modeling. All modules are lightweight, modular, and seamlessly integrate with mainstream detectors. Extensive experiments on MS COCO, VisDrone, and other benchmarks demonstrate consistent and significant gains in small-object AP (+3.2–5.8 points), validating the effectiveness of joint upstream-downstream optimization and strong generalizability across diverse scenarios and architectures.

Enhancing downstream task performance effectivelyImproving feature purification and sample labelingOptimizing small object detection pipeline holistically

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

Latest Papers

What's happening recently
View more

Diffusion models suffer from high inference costs due to their large network size and multi-step denoising process, and existing compression methods struggle to simultaneously achieve structural simplicity and performance retention. This work proposes a synergistic compression framework that integrates structured pruning with lightweight teacher-aligned restoration and single-step distillation (SiDA), enabling seamless integration of step distillation without requiring retraining after pruning for the first time. Built upon the EDM2-XS architecture, the pruned model achieves an FID of 3.12 on ImageNet-512 with only 98.8M parameters and a single forward pass at 20% sparsity, and an FID of 4.26 at 30% sparsity—significantly outperforming current baselines while substantially reducing inference overhead without compromising generation quality.

Diffusion PruningInference CostModel Compression

This work addresses the challenge that large language models struggle to make reliable single-step decisions in tensor program optimization, primarily due to the absence of verifiable step-level supervision and interpretability in existing datasets. To bridge this gap, the authors introduce Step-TP, the first step-level optimization dataset enabling closed-loop reasoning. Step-TP leverages atomic and composable optimization strategies, coupled with verifiable intermediate representations, explicit state transitions, and a strategy filtering mechanism, to construct structured chain-of-thought reasoning paths. This approach not only ensures token efficiency and deterministic decompilation to TVM TIR but also substantially enhances the model’s reliability in making single-step decisions within complex optimization spaces.

chain-of-thought reasoningcombinatorial optimizationlarge language models

This work addresses time series characterized by irreversible state evolution—such as equipment degradation, task completion, or neural dynamics—and proposes a novel “latent compass” representation. By leveraging self-supervised contrastive learning, the method constructs a structured latent space in which each time series is mapped onto a manifold trajectory between two orthogonal prototype vectors. State progression and operational mode are disentangled via polar coordinates (θ, r), enabling transparent and interpretable modeling without requiring labeled data. Evaluated on industrial degradation, robotic tasks, and neural activity datasets, the approach achieves performance on par with or superior to black-box deep models in endpoint prediction, multi-step forecasting, and phase separation tasks—even when paired with simple linear regression—while substantially enhancing model interpretability and computational efficiency.

interpretable representationsirreversible state transitionslatent geometry

Existing machine unlearning methods often suffer from either over-unlearning, which degrades model utility, or under-unlearning, which leaves residual privacy risks, primarily due to the absence of precise guidance signals. To address this, this work proposes the GSUO framework, which introduces, for the first time, a task-aware, fine-grained, and differentiable guidance mechanism that dynamically adjusts unlearning intensity based on the memorization strength of individual samples. GSUO supports diverse unlearning scenarios, including random subsets and class-level removal. Extensive experiments demonstrate that GSUO significantly outperforms 14 baseline methods in terms of unlearning efficacy, model generalization, and computational efficiency, offering a highly effective, reliable, and versatile solution for machine unlearning.

guidance signalmachine unlearningmemorization strength

This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.

compression and accelerationconstraint-drivendeployment constraints

Hot Scholars

ZL

Zhiyong Li

Professor of Computer Science, Hunan University
computer vision,object detection
SW

Sibo Wang

The Chinese University of Hong Kong
Databases
JR

Josh Rosen

Northeastern University
computational social sciencecomplex systemscausal inferencehuman mobility