progressive supervision weighting

Design and implement adaptive supervision schemes that dynamically weight and schedule multiple supervisory signals—such as visual perception losses versus task-correctness or reward signals—over training time and across individual samples. Build dependency-guided and sample-adaptive mechanisms that progressively enhance visual supervision, align supervision intensity to each sample’s necessity, and mitigate bias from uneven reward or supervision distributions.

progressivesupervisionweighting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Making Evidence Actionable in Adaptive Learning

Nov 17, 2025
AM
Amirreza Mehrabi
🏛️ Purdue University

Adaptive learning systems often achieve precise diagnosis but suffer from weak pedagogical interventions, leading to delayed or mismatched instructional responses. To bridge the gap between diagnosis and teaching, this paper proposes a teacher-led feedback闭环 system that transforms knowledge-component–level assessment evidence into empirically validated micro-interventions. We introduce three novel safeguards—complete closure of ability gaps, cognitive load control under attention constraints, and anti-redundancy diversity preservation—formulated as a constrained binary integer programming problem. Our hybrid solver balances the richness-latency trade-off by jointly modeling ability estimation, prerequisite dependencies, difficulty windows, and diversity constraints. Evaluated in a large-scale physics course (N ≈ 1,000), the system achieves near-universal mastery across skills, reduces redundant recommendations by 12 percentage points, improves task difficulty distribution uniformity, maintains low computational overhead, and ensures subgroup fairness.

Achieves full skill coverage while reducing redundancy across diverse learnersConverts assessment evidence into vetted micro-interventions for adaptive learningFormalizes intervention assignment with coverage and anti-redundancy constraints

Change detection with adaptive sampling for binary responses

Dec 17, 2025
YY
Yanqing Yi
🏛️ Memorial University of Newfoundland | National Chengchi University

This paper addresses online change detection for binary (conforming/non-conforming) responses in multi-production-line systems. We propose a dynamic monitoring method based on adaptive sampling. Our approach innovatively formulates change detection as a Markov decision process (MDP) with average reward, where the reward function jointly incorporates the likelihood ratio test statistic and sampling information; the optimal sampling policy is derived via Bellman iteration. The method enables real-time identification of potentially out-of-control lines and prioritizes sampling accordingly. Experimental results demonstrate that, when sample size ≥ 20, the proposed method achieves significantly higher statistical power than proportional random sampling. Moreover, both the average sampling proportion allocated to anomalous lines and overall detection performance improve monotonically with increasing sample size or greater disparity in out-of-control/ in-control probabilities across lines.

Adaptive sampling detects changes in multi-line systemsImproves statistical power over equal randomization for binary responsesMethod uses Markov decision process for optimal sampling allocation

Reinforcement learning in reasoning tasks suffers from sparse outcome supervision and difficulty in credit assignment across intermediate steps, while existing process supervision relies on costly human annotations that hinder scalability. This work proposes a novel paradigm termed “supervision internalization,” which leverages a self-reflection mechanism to automatically identify and correct failed reasoning trajectories, thereby generating fine-grained process-level supervision signals endogenously from only outcome feedback—without requiring external annotations. This approach enables precise credit assignment and significantly improves both policy training efficiency and reasoning performance, offering a scalable pathway toward fine-grained reinforcement learning for complex reasoning tasks.

credit assignmentoutcome supervisionprocess supervision

An Adaptive Method for Weak Supervision with Drifting Data

Jun 02, 2023
AM
Alessio Mazzetto
🏛️ Brown University

This work addresses the challenge of label inference degradation in non-stationary environments, where weak supervision sources (e.g., crowd annotations, heuristic rules) exhibit time-varying accuracy due to concept drift. We propose an adaptive sliding window mechanism that requires no prior assumptions about drift patterns. Our method online estimates the real-time accuracy of each weak source and dynamically selects the optimal window size via an explicit bias–variance trade-off, balancing responsiveness to drift against statistical stability. Unlike conventional approaches relying on fixed windows or explicit drift detection, ours is the first to integrate non-stationary statistical inference directly into the weak supervision learning framework. Experiments on synthetic and real-world crowdsourced datasets demonstrate strong robustness across diverse drift patterns—including abrupt, gradual, and periodic drift—and yield significant improvements in final label inference accuracy.

Adaptive weak supervision for non-stationary data driftDynamic window sizing for optimal label inferenceNo a priori assumptions on drift magnitude needed

Latest Papers

What's happening recently
View more

This work investigates whether it is statistically justified to dynamically invoke auxiliary signals—such as LLM inference—in recommendation systems on a per-instance basis, distinguishing between average-effectiveness and instance-level adaptive strategies. The authors introduce a theoretical lower bound based on reward signal-to-noise ratio (reward-SNR), proving that instance-level acquisition policies can only be reliably learned when SNR > ρ*(N) ≈ 2.8/√N; otherwise, apparent structural patterns are merely noise artifacts. Using the Structured Hypothesis Embeddings (SHE) framework—which integrates frozen LLM confidence scores, offline policy evaluation, and multi-granularity routing experiments—the study evaluates three benchmarks: MIND, REES46, and Amazon-Beauty. All datasets fall below the SNR threshold, demonstrating that learned dynamic strategies perform no better than random, thereby revealing a lack of statistical foundation for current practices in adaptive auxiliary signal invocation.

learnable structureLLM acquisitionper-instance policy

This work addresses the challenge of online reinforcement learning fine-tuning for pretrained vision-language-action models under sparse binary rewards, where existing approaches struggle to provide effective per-timestep supervision and conflate feasibility with efficiency objectives, leading to credit assignment errors in mixed intervention trajectories. The authors propose Hierarchical Advantage-Weighted Behavioral Cloning (HABC), which trains separate critic heads for feasibility and efficiency on distinct data subsets and employs a state-adaptive gating mechanism to dynamically fuse their advantages into fine-grained actor loss weights. Additionally, an intervention-aware credit assignment scheme assigns outcome labels only to autonomously executed segments. By decoupling and dynamically integrating these two advantage signals for the first time, HABC substantially improves success rates on three real-world, contact-rich bimanual tasks—from 36%, 44%, and 12% to 92%, 88%, and 38%, respectively.

credit assignmenthierarchical advantageonline RL fine-tuning

This work addresses the limitations of conventional training pipelines, which struggle to dynamically mitigate issues such as overfitting, loss imbalance, and unsafe exploration due to reliance on fixed policies or single-axis schedulers. The authors propose a large language model–based bounded supervisory controller that leverages structured telemetry snapshots to monitor training dynamics in real time and generates verifiable multi-parameter adjustment commands within a constrained action space. This enables closed-loop regulation of learning rate, regularization strength, loss weighting, and exploration strategy. Notably, it introduces pattern-constrained large language models into training supervision for the first time, supporting asynchronous, auditable multi-axis interventions applicable to both supervised and reinforcement learning. Experiments demonstrate a 60% loss reduction on TinyStories with effective overfitting correction, marked alleviation of overly conservative or unsafe exploration in robotic manipulation tasks, and generation of traceable intervention logs.

adaptive trainingexploration collapseloss imbalance

Hot Scholars

XW

Xueqian Wang

Tsinghua University
Information FusionTarget DetectionRadar ImagingImage Processing
PZ

Pengyu Zeng

清华大学
人工智能、深度学习
IM

Ioannis Mademlis

National and Kapodistrian University of Athens
Machine LearningComputer VisionAutonomous RoboticsHuman-Computer Interaction
GP

Georgios Papadopoulos

PhD candindate Imperial College London
Causal InferenceTime SeriesMachine LearningBiostatistics