Score
Designs and implements evaluation protocols and downstream tasks to measure the quality, utility, and fidelity of policy representations or embeddings. Builds metrics and benchmark experiments to compare representation methods across predictive and decision-making tasks, analyze policy evaluation outcomes, and validate how representations support downstream algorithms or analyses.
Downstream probing only assesses task-relevant information in representations, failing to characterize critical properties—such as equivariance, invariance, and disentanglement—that govern interpretability and generalization; moreover, existing evaluation frameworks lack standardization, modularity, and cross-modal applicability. Method: We propose the first representation quality assessment framework that transcends downstream tasks, employing controlled factorial probe design to systematically quantify informativeness, equivariance, invariance, and disentanglement. The framework is modular, interpretable, and supports cross-modal analysis (e.g., image and speech). Contribution/Results: It establishes the first standardized, multi-dimensional semantic attribute disentanglement protocol. Experiments reveal substantial divergence in intrinsic representation properties—even among models with comparable downstream performance—enabling fine-grained representation understanding, diagnosis, and optimization. This work introduces a novel paradigm and practical toolkit for representation evaluation beyond task-specific metrics.
Existing probabilistic forecasting evaluation metrics primarily emphasize predictive accuracy while neglecting their practical utility in downstream decision-making tasks, leading to a misalignment between evaluation and application. To address this, we propose a data-driven evaluation alignment framework that formulates the learning of a surrogate evaluation function as an end-to-end optimization problem. Leveraging proper scoring rule theory, our approach employs a neural network-parameterized weighted scoring rule to automatically learn an evaluation function aligned with downstream objectives—without assuming any prior cost structure. This work is the first to formalize evaluation alignment as a learnable problem, combining theoretical rigor with engineering scalability. Experiments on synthetic and real-world regression tasks demonstrate its effectiveness: it significantly reduces the gap between evaluation scores and downstream decision utility, enabling rapid, task-adaptive model selection and hyperparameter tuning.
Existing attribution methods lack a unified, scalable, and reproducible evaluation framework, hindering systematic comparison. To address this gap, this work proposes the first modular benchmarking framework for attribution, decoupling the pipeline into five interoperable layers—data, preprocessing, model, attribution method, and evaluation—and enabling flexible integration through abstract interfaces and a dynamic registration mechanism. The framework introduces an innovative four-tier categorization system coupled with an automated testing protocol that rigorously validates whether each implemented method reproduces results from its original publication. An interactive web interface further supports multidimensional configuration and comparative analysis. Currently integrating 28 state-of-the-art attribution methods, the framework establishes the first automated, quantitative guarantee of method-level reproducibility in the field.
This work addresses the longstanding challenge of automatically translating legal texts into executable decision logic, which has traditionally relied on manual encoding and evaluation. The authors propose enhancing large language models by introducing an intermediate structured representation and present the first systematic assessment—based on real-world data from the Dutch Environmental Planning Act—of how input/output constraints and semantic role labeling influence the structural and functional equivalence of generated logic. Experimental results demonstrate that incorporating I/O constraints improves structural similarity by 37–54%, achieves functional equivalence in 51–53% of test cases, and automatically eliminates 45–55% of redundant logic nodes, thereby revealing a notable inconsistency between structural similarity and functional equivalence.
Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.
This study addresses the limitations of policy evaluation in capturing institutional dynamics and predicting second-order effects by proposing a Second-Order Effect State Transition Framework. We construct a source-link benchmark comprising 96 cases alongside a side-effect simulator, and introduce a transition channel auditing protocol to systematically validate model sensitivity to downstream institutional impacts. Experimental results demonstrate that the proposed simulator achieves a policy effect quality score of 0.945, significantly outperforming baseline methods. This research effectively enhances model sensitivity to complex institutional environments, providing novel methodological support and evaluation benchmarks for dynamic policy assessment.
本文介绍了GPS-Bench,一种基于证据的治理政策模拟基准,通过链接多种公开记录来重建政策相关方及其行动和影响,以评估不同方法在预测政策效果方面的表现。
This study addresses how AI evaluation influences multi-party decision-making, arguing that existing benchmark designs urgently require contextualization within ecosystem dynamics. We propose a simulation framework based on large language model-driven generative agents, integrating rule-based markets with strategic agents to model evaluative ecosystem interactions within a six-dimensional capability space. Our analysis reveals the nonlinear effects of private holdout sets on disparities between scores and satisfaction, demonstrating that these effects vary significantly across different weight distributions. By establishing a verifiable hypothesis sandbox, this work provides a systematic approach for investigating how evaluation policies reshape the broader AI ecosystem.
为解决AI代理实验理解问题,引入WhatWorkedBench基准测试方法,通过预测组件更改后的结果准确性来评估,使用高斯过程提高效果恢复精度。
论文针对代理基准测试中的双重测量混淆问题,通过将关键决策转移给模型、使用基于真实值的评分及报告更全面的可靠性指标来解决。