Score
Designs and implements adaptation procedures that use a separate "judge" model's feedback to optimize, repair, or otherwise update a target model (e.g., a generator or conditioner), including in-the-loop optimization and held-out-judge evaluation to avoid circularity. Work covers building parameter-efficient update methods, integrating judges during training or inference, and analyzing targeted repairs and robustness under degradation, including cases where the judge may be a multimodal or vision-language model.
This work addresses the challenge of reliably optimizing single-image 3D generation quality without access to ground-truth 3D annotations by leveraging vision-language models (VLMs). To this end, the authors propose a debiased VLM-based evaluation protocol that constructs robust pairwise quality guidance through separation of training and evaluation phases, correction of positional bias, and three targeted failure-recovery mechanisms. Built upon dual VLM backbones—Qwen2.5-VL-7B and InternVL3-8B—the approach integrates parameter-efficient fine-tuning, conditional repair, and debiased scoring. Using only publicly available models and data, the method achieves a 0.50 win rate against strong baselines even in severely degraded scenarios, demonstrating the effectiveness of the proposed protocol and revealing inherent performance ceilings in lightweight adaptation strategies.
This study addresses the challenge of evaluating program modifications during test-time agent evolution due to the absence of ground-truth rewards. We propose JET, a method that evolves executable judges on source task trajectories and transfers them in a frozen manner to target domains, guiding unsupervised program adaptation without requiring external evaluators or model weight updates. The core innovation lies in introducing a cross-task transfer mechanism for executable judges, integrating program evolution with trajectory-based judgment logic construction. Experiments demonstrate that JET achieves a 13% average reward improvement and a 36% relative increase in exact success rate under WebShop cold-start settings, while also validating the existence of search bottlenecks in the PushT task.
This study investigates the effectiveness of large language models as judges (LLM-as-a-judge) in providing optimization signals for closed-loop table recognition. Leveraging the FinTabNet and OmniDocBench datasets, and combining deterministic TEDS evaluation with structure-preserving instruction constraints, a 2×2 controlled experiment reveals that LLM-provided judgment signals are weak and unreliable—random selection often outperforms score-based iterative strategies. Although structure-preserving constraints substantially reduce severe structural errors, they fail to improve overall performance. The work is the first to demonstrate a critical disconnect between an LLM’s evaluative capability and its utility in optimization, thereby challenging the validity of relying solely on judgment scores for iterative refinement and underscoring the necessity of incorporating stronger verification signals to achieve effective closed-loop optimization.
This work addresses the susceptibility of large language models used as evaluators (LLM-as-judge) to systematic biases, which undermines assessment reliability. The authors propose CyclicJudge, a novel method that introduces the first bias analysis framework based on variance decomposition, disentangling evaluation scores into distinct components attributable to the scenario, the generated response, the judge, and residual noise. By incorporating a round-robin assignment mechanism, CyclicJudge completely eliminates judge-induced bias without increasing the per-evaluation computational cost. Experimental results on MT-Bench demonstrate that the proposed approach effectively removes systematic bias, significantly enhancing both consistency and fairness in model evaluations.
To address the challenge of optimizing model parameters post-training using non-differentiable real-world feedback (e.g., user ratings, BLEU, word error rate), this paper proposes AfterLearnER: a gradient-free framework that, during inference or post-training, directly optimizes a critical subset of parameters via evolutionary algorithms (e.g., CMA-ES), requiring only数十 to hundreds of scalar feedback evaluations—enabling anytime optimization and human-feedback-driven dynamic adaptation. Theoretically, it introduces a generalization bound analysis to ensure robustness against overfitting. Empirically, AfterLearnER is evaluated across diverse tasks—including depth estimation, speech resynthesis, Doom gameplay, code translation, and latent diffusion—demonstrating significant improvements over conventional fine-tuning on practical metrics such as BLEU, image quality, and game scores. It is the first method to achieve gradient-free, few-shot, and high-generalization post-training refinement.
This study investigates whether large language model (LLM)-based scientific agents rely primarily on initial task context or experimental feedback when making decisions during neural operator adaptation. Specifically, this work examines the capacity of LLMs to select fine-tuning configurations under limited computational budgets. Through controlled intervention experiments, we decouple the mechanisms underlying prior knowledge and feedback signals, benchmarking the approach against random search and Bayesian optimization. Our findings demonstrate that LLM-driven decision-making responds simultaneously to task descriptions and empirical observations, effectively integrating task priors with feedback sensitivity. In most scenarios, the LLM agent outperforms traditional baselines, achieving lower test errors while utilizing experimental feedback more efficiently. These results establish a novel paradigm for transfer learning in partial differential equations.
本文探讨了通过实用方法和无代码工具MUSE整合多源反馈,以支持计算设计,增强设计师的信心与灵活性。
This study addresses the version-dependence bias inherent in fixed LLM judges evaluating agent iterations, which risks misjudgment and violates error invariance under varying task conditions. Leveraging the SWE-bench and tau-bench benchmarks alongside statistical multiple-testing corrections and bootstrap confidence intervals, we reveal that stronger model capabilities paradoxically increase false positive rates and cause cross-domain calibration failures. We demonstrate that paired auditing outperforms standalone judging or legacy-version calibration. Empirically, 32 of 60 evaluation units exhibit detectable discrepancies, while legacy-version calibration inflates errors to 19.5 points, underscoring the necessity of human review for reliable agent assessment.
This study addresses the challenge of auditing biases in multimodal large language models (MLLMs) used for image editing evaluation, where judges are susceptible to irrelevant cues and struggle to verify quality preservation. We propose EditJudgeBias, a benchmark that systematically audits MLLM judge biases across invariance, consistency, and stability dimensions. This is achieved through counterfactual data safeguarded by calibrated validators and a noise-floor comparison mechanism. Our findings reveal that no single metric can comprehensively characterize robustness. Furthermore, we demonstrate that all evaluated judges are vulnerable to spurious extraneous cues: fabricated majority opinions inflate scores, and swapping candidate orders induces preference reversals in 60.9% of cases.
研究探讨了LLM评判模型被评分系统操纵的问题,提出并测试了先提交答案再评判的方法,发现该方法的效果取决于评判模型自身解决问题的能力。