Score
Design and implement evaluation pipelines that iteratively hold out each subject as the test set while training and tuning models on the remaining subjects, compute and report per-subject performance metrics, and aggregate those results to estimate subject-independent generalization. This involves creating subject-based data splits, orchestrating the train/validation/test runs for each held-out subject, and summarizing variability across subjects.
Existing evaluation methodologies struggle to balance theoretical rigor with practical scalability, resulting in prohibitively high experimental costs. This paper proposes a computational evaluation theory for parametric agents, transcending limitations of conventional frameworks. Our key contributions are threefold: (1) the first unified upper-bound theory linking generalized evaluation error to causal effect estimation error; (2) a meta-learner that enables consistent modeling across heterogeneous agent spaces; and (3) simultaneous guarantees of statistical consistency and predictive efficiency. Evaluated across 12 canonical scenarios, our approach reduces evaluation error by 24.1%–99.0% and accelerates computation by 3–7 orders of magnitude relative to empirical experimentation or simulation. These gains substantially enhance both the scalability and reliability of agent evaluation.
研究通过使用独立实现的训练栈作为差异预言机来验证语言模型训练过程,以解决神经网络训练中的oracle问题。
This work addresses the limitations of evaluating language model training efficiency solely through end-point metrics under compute-constrained settings, which often overlooks instability and negative returns during training. The authors propose a multidimensional evaluation paradigm based on training trajectories, conducting repeated-measures experiments within a fixed token budget to track the training dynamics of a 4.26M-parameter Llama-style model on the TinyStories corpus. Combining repeated-measures ANOVA with interval-level telemetry analysis, they observe rapid early convergence—validation loss drops from 8.36 to 2.80 within approximately 4M tokens—followed by non-monotonic degradation, with loss rebounding to 3.90 and no discernible stable training phase. These findings suggest that indiscriminately increasing training tokens may be counterproductive, offering a novel perspective for efficient low-resource training evaluation.
This work addresses the critical challenge of evaluating model generalization in high-stakes scenarios with scarce labels, where existing methods lack reliable, label-free metrics for pre-deployment model selection and post-deployment performance monitoring. To bridge this gap, the study introduces, for the first time, the internal causal circuit mechanisms of Vision Transformers into generalization assessment, proposing two novel unsupervised metrics: Dependency Depth Bias and Circuit Shift Score. The former quantifies depth-wise biases in representational dependency structures, while the latter measures changes in circuit stability under distribution shifts. Extensive experiments across diverse tasks demonstrate that these metrics achieve substantially higher correlations with true generalization performance—improving by 13.4% and 34.1% on average over current approaches—thereby significantly enhancing the reliability of generalization prediction without requiring ground-truth labels.
This study addresses the issues of information leakage and optimistic bias in performance evaluation arising from improper data splitting during machine learning model validation. Focusing on biomedical contexts, it systematically reviews validation strategies and contrasts flawed versus leakage-free designs through eight controlled experiments. Employing techniques such as nested grouped cross-validation for reproducible simulations, this work proposes deployment-oriented, scenario-specific validation guidelines. Its primary contributions include practical tools—namely decision trees, checklists, and code templates—that operationalize the core principle of aligning independent units with deployment objectives within an auditable evaluation framework. Ultimately, these contributions significantly enhance the reliability and standardization of machine learning model validation practices.
This study addresses the attribution misalignment between evidence and component-level scientific claims in composite systems by proposing a topic-typed claim licensing mechanism. Methodologically, it introduces a “scientific topic” dimension to decouple claim strength from attribution, thereby enabling precise evidence mapping. Furthermore, the paper presents the SCOPE-Routing framework, which integrates preference-conditioned multi-graph routing with a declarative semantic reproduction mechanism. This work effectively distinguishes weak conclusions from non-substitutable credit, significantly reducing evaluation confusion and review bias. By revealing the fundamental differences between score-optimal and claim-qualified approaches, it provides a reliable credit-preservation solution for hybrid systems.
This study addresses the heavy reliance on manual effort and the difficulty of sustained optimization in task adaptation for large language models by proposing an autonomous post-training framework. Inspired by gradient-based optimization, the method precisely identifies model deficiencies through error attribution and instance-level analysis. It then automatically generates training data and configurations via strategy retrieval, achieving closed-loop iterative updates through small-scale validation. Experimental results demonstrate that this framework improves the performance of base and instruct models across eleven tasks by an average of 18.29 and 11.97 percentage points, respectively, with gains reaching up to 41.96 percentage points. These findings establish the proposed approach as an efficient and fully automated solution for the continuous improvement of large language models.
本文通过引入VTC-Bench和Validated Task Coverage方法,解决了多输出语言模型评估中忽视多样性和实用性的问题。
This work addresses the challenges in multi-subject personalized image generation—such as subject omission, appearance distortion, and interaction mismatch—and the absence of effective evaluation metrics. The authors propose MIBE, a unified framework comprising a Multi-subject Interaction Benchmark (MIB) and a lightweight diagnostic evaluator (MIE), establishing the first standardized benchmark for multi-subject interactions. By decoupling data mechanisms to span diverse relationships and scene complexities, and leveraging 60k vision-language model silver labels alongside 4k double-blind human gold-standard annotations, the evaluator is trained with a dual-head ranking and diagnostic objective. MIE achieves a 0.922 overall pairwise accuracy on the gold-standard set, with 0.982 and 0.884 for seen and unseen generators, respectively, significantly outperforming baselines like CLIP and DINO, while demonstrating high alignment with human preferences and strong generalization across generators.
Traditional language model evaluation often conflates the ability to produce assessable responses with the correctness of those responses, thereby masking execution-level failure modes under aggregate accuracy metrics. This work proposes a two-tiered evaluation framework that disentangles scorer-agnostic execution states—such as termination, answer exposure, parseability, and output length—from scorer-dependent correctness judgments. By enforcing a fixed output budget, tracking multidimensional execution trajectories, formulate a verification mechanism driven by coverage auditing, the study systematically uncovers divergent execution behaviors across models on MATH and ARC-Challenge benchmarks. The analysis reveals that extended output lengths can mitigate certain failure modes and demonstrates that verification strategies substantially influence comparative accuracy outcomes.