judge-guided adaptation

Designs and implements adaptation procedures that use a separate "judge" model's feedback to optimize, repair, or otherwise update a target model (e.g., a generator or conditioner), including in-the-loop optimization and held-out-judge evaluation to avoid circularity. Work covers building parameter-efficient update methods, integrating judges during training or inference, and analyzing targeted repairs and robustness under degradation, including cases where the judge may be a multimodal or vision-language model.

judge-guidedadaptation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of reliably optimizing single-image 3D generation quality without access to ground-truth 3D annotations by leveraging vision-language models (VLMs). To this end, the authors propose a debiased VLM-based evaluation protocol that constructs robust pairwise quality guidance through separation of training and evaluation phases, correction of positional bias, and three targeted failure-recovery mechanisms. Built upon dual VLM backbones—Qwen2.5-VL-7B and InternVL3-8B—the approach integrates parameter-efficient fine-tuning, conditional repair, and debiased scoring. Using only publicly available models and data, the method achieves a 0.50 win rate against strong baselines even in severely degraded scenarios, demonstrating the effectiveness of the proposed protocol and revealing inherent performance ceilings in lightweight adaptation strategies.

3D mesh qualityde-biased evaluationparameter-efficient adaptation

This study addresses the challenge of evaluating program modifications during test-time agent evolution due to the absence of ground-truth rewards. We propose JET, a method that evolves executable judges on source task trajectories and transfers them in a frozen manner to target domains, guiding unsupervised program adaptation without requiring external evaluators or model weight updates. The core innovation lies in introducing a cross-task transfer mechanism for executable judges, integrating program evolution with trajectory-based judgment logic construction. Experiments demonstrate that JET achieves a 13% average reward improvement and a 36% relative increase in exact success rate under WebShop cold-start settings, while also validating the existence of search bottlenecks in the PushT task.

agent programsexecutable judgeprogram adaptation

This study investigates the effectiveness of large language models as judges (LLM-as-a-judge) in providing optimization signals for closed-loop table recognition. Leveraging the FinTabNet and OmniDocBench datasets, and combining deterministic TEDS evaluation with structure-preserving instruction constraints, a 2×2 controlled experiment reveals that LLM-provided judgment signals are weak and unreliable—random selection often outperforms score-based iterative strategies. Although structure-preserving constraints substantially reduce severe structural errors, they fail to improve overall performance. The work is the first to demonstrate a critical disconnect between an LLM’s evaluative capability and its utility in optimization, thereby challenging the validity of relying solely on judgment scores for iterative refinement and underscoring the necessity of incorporating stronger verification signals to achieve effective closed-loop optimization.

closed-loop regenerationevaluation abilityLLM-as-a-judge

This work addresses the susceptibility of large language models used as evaluators (LLM-as-judge) to systematic biases, which undermines assessment reliability. The authors propose CyclicJudge, a novel method that introduces the first bias analysis framework based on variance decomposition, disentangling evaluation scores into distinct components attributable to the scenario, the generated response, the judge, and residual noise. By incorporating a round-robin assignment mechanism, CyclicJudge completely eliminates judge-induced bias without increasing the per-evaluation computational cost. Experimental results on MT-Bench demonstrate that the proposed approach effectively removes systematic bias, significantly enhancing both consistency and fairness in model evaluations.

benchmark evaluationjudge biasLLM-based evaluation

Evolutionary Retrofitting

Oct 15, 2024
MV
Mathurin Videau
🏛️ Meta AI | TAU | INRIA | LISN | Univ Gustave Eiffel | CNRS | LIGM | University of Rouen Normandy | LITIS | Khalifa University | Thales - CortAIx-Labs

To address the challenge of optimizing model parameters post-training using non-differentiable real-world feedback (e.g., user ratings, BLEU, word error rate), this paper proposes AfterLearnER: a gradient-free framework that, during inference or post-training, directly optimizes a critical subset of parameters via evolutionary algorithms (e.g., CMA-ES), requiring only数十 to hundreds of scalar feedback evaluations—enabling anytime optimization and human-feedback-driven dynamic adaptation. Theoretically, it introduces a generalization bound analysis to ensure robustness against overfitting. Empirically, AfterLearnER is evaluated across diverse tasks—including depth estimation, speech resynthesis, Doom gameplay, code translation, and latent diffusion—demonstrating significant improvements over conventional fine-tuning on practical metrics such as BLEU, image quality, and game scores. It is the first method to achieve gradient-free, few-shot, and high-generalization post-training refinement.

Applying evolutionary retrofitting to improve performance on threshold-based metricsOptimizing trained models using evolutionary algorithms on non-differentiable error signalsRefining model parameters post-training with limited validation data feedback

Latest Papers

What's happening recently
View more

This study investigates whether large language model (LLM)-based scientific agents rely primarily on initial task context or experimental feedback when making decisions during neural operator adaptation. Specifically, this work examines the capacity of LLMs to select fine-tuning configurations under limited computational budgets. Through controlled intervention experiments, we decouple the mechanisms underlying prior knowledge and feedback signals, benchmarking the approach against random search and Bayesian optimization. Our findings demonstrate that LLM-driven decision-making responds simultaneously to task descriptions and empirical observations, effectively integrating task priors with feedback sensitivity. In most scenarios, the LLM agent outperforms traditional baselines, achieving lower test errors while utilizing experimental feedback more efficiently. These results establish a novel paradigm for transfer learning in partial differential equations.

Experimental FeedbackLarge Language ModelsNeural Operators

This study addresses the version-dependence bias inherent in fixed LLM judges evaluating agent iterations, which risks misjudgment and violates error invariance under varying task conditions. Leveraging the SWE-bench and tau-bench benchmarks alongside statistical multiple-testing corrections and bootstrap confidence intervals, we reveal that stronger model capabilities paradoxically increase false positive rates and cause cross-domain calibration failures. We demonstrate that paired auditing outperforms standalone judging or legacy-version calibration. Empirically, 32 of 60 evaluation units exhibit detectable discrepancies, while legacy-version calibration inflates errors to 19.5 points, underscoring the necessity of human review for reliable agent assessment.

Agent EvaluationEvaluation ReliabilityLLM-as-a-Judge

This study addresses the challenge of auditing biases in multimodal large language models (MLLMs) used for image editing evaluation, where judges are susceptible to irrelevant cues and struggle to verify quality preservation. We propose EditJudgeBias, a benchmark that systematically audits MLLM judge biases across invariance, consistency, and stability dimensions. This is achieved through counterfactual data safeguarded by calibrated validators and a noise-floor comparison mechanism. Our findings reveal that no single metric can comprehensively characterize robustness. Furthermore, we demonstrate that all evaluated judges are vulnerable to spurious extraneous cues: fabricated majority opinions inflate scores, and swapping candidate orders induces preference reversals in 60.9% of cases.

Automated JudgesBias AuditingImage Editing Evaluation

研究探讨了LLM评判模型被评分系统操纵的问题,提出并测试了先提交答案再评判的方法,发现该方法的效果取决于评判模型自身解决问题的能力。

commit-first judgingevaluation frameworksLLM judges

Hot Scholars

ES

Ellis Solaiman

Reader/Professor, SFHEA, FBCS, Newcastle University
TrustResilienceInternet of ThingsAI
YZ

Yousong Zhu

Associate Professor, Chinese Academy of Sciences, Institute of Automation
Multimodal Large Language ModelsSelf-supervised LearningObject Detection
YZ

Yufei Zhan

Institute of Automation, Chinese Academy of Science
Computer VisionLarge Multimodal ModelsGrounding and Detection
NF

Nuno Fachada

U. Lusófona, CTS-Center of Technology and Systems / UNINOVA
artificial intelligencegamesmodeling and simulationresearch software