Score
Design, build, or evaluate methods that assign tasks or items to discrete or continuous difficulty levels — e.g., by constructing classifiers, regressors, scoring functions, or calibration pipelines that map task features and agent/human performance signals to difficulty strata. Validate and analyze the stratification by measuring consistency, calibration, and how well difficulty tiers predict performance or error rates across systems or populations.
Traditional item difficulty estimation relies on costly field testing and is constrained by classical test theory’s assumptions. Method: This study systematically reviews and empirically evaluates text-based automated difficulty prediction methods, synthesizing findings from 37 studies within a unified evaluation framework. It benchmarks classical machine learning models against Transformer architectures—ranging from small to large—using only item stems as input, without manual feature engineering. Contribution/Results: For the first time, cross-study model benchmarks are aggregated, revealing Transformers’ superior capacity to capture syntactic and semantic difficulty cues. The best-performing model achieves RMSE = 0.165, Pearson correlation = 0.87, and classification accuracy = 0.806. These results demonstrate the feasibility of purely text-driven difficulty prediction, offering substantial gains in efficiency, scalability, and fairness—establishing a novel paradigm for intelligent assessment design.
This work addresses the limitations of existing curriculum learning approaches, which rely on static or computationally expensive dynamic difficulty assessments and struggle to generate efficient, learner-specific training sequences. The authors propose a novel problem difficulty evaluation mechanism based on a relative measure of model capability, introducing and formally defining “transitional problems”—critical instances that shift from difficult to easy as the model’s competence improves. Leveraging this insight, they construct an adaptive curriculum that aligns dynamically with the learner’s evolving capacity, yielding a personalized, interpretable, and computationally efficient training trajectory. Experiments on chess and mathematical reasoning tasks demonstrate that the proposed strategy significantly outperforms current methods, effectively facilitating transitions to higher levels of model performance.
本文将链梯法的模式选择视为监督学习问题,通过定义损失函数、惩罚项和超参数来优化调整模式,以更客观地设定经验调整。
This study addresses the cold-start problem in item parameter estimation when newly developed test items lack empirical response data. The authors propose a prediction approach leveraging textual embeddings and regularized regression, accompanied by an evaluation framework integrating resampling-based cross-validation, reliability ceilings, and design ceilings. Innovatively employing a dual “ceiling” analysis, the work demonstrates that differences in parameter predictability stem primarily from measurement reliability rather than the strength of textual information, underscoring the necessity of repeated validation. In the EEDI mathematics item bank, predicted difficulty parameters achieved an R² of 0.53, representing 57% of the reliability ceiling, whereas pseudo-guessing parameters in the three-parameter logistic model proved largely unpredictable due to near-zero reliability ceilings. BEA benchmark experiments further reveal that relying solely on RMSE can obscure extremely low explained variance, highlighting the critical role of dimensionless metrics in model evaluation.
Existing machine learning frameworks suffer from insufficient formalization of objective functions and lack a unified, cross-domain behavioral design paradigm. Method: We propose an equation-constrained compositional function modeling approach for learners, constructing task graphs and compositional semantic graphs to enable model-agnostic behavioral specification and optimization. We introduce a novel task-oriented pattern language framework and the “manipulator” task paradigm, supporting end-to-end, architecture-agnostic, and adversarial-training-free minimal editing of data attributes. Contribution/Results: Theoretically, our work integrates formal methods and theoretical computer science principles. Empirically, we demonstrate precise, controllable, and interpretable behavioral editing on small-scale models under stable training—without stochastic sampling or data intervention—yielding significant improvements in deployment efficiency and formal verifiability.
Standard language modeling loss fails to accurately predict downstream task performance for pretrained language models under overtraining regimes. Method: This paper proposes a task-level two-stage scaling framework: first predicting task-specific loss as a function of model and data scale, then mapping loss to task accuracy. It introduces a lightweight “model ladder” strategy—requiring only 1% of the target model’s compute—to efficiently fit multi-task performance trends. Contribution/Results: The framework achieves precise performance forecasting for four downstream multiple-choice tasks, with mean absolute error ≤2 percentage points. Experiments demonstrate a strong correlation between the number of ladder models and prediction accuracy, significantly outperforming single-step power-law baselines. This work overcomes the fundamental limitation of conventional loss-based scaling laws, which cannot characterize task-specific behavior.
This work proposes a method to accurately predict agent task difficulty directly from task descriptions without requiring time-consuming simulations, thereby supporting benchmark calibration and curriculum design. Addressing the limitations of AUC-based metrics in difficulty assessment, the approach introduces token-level entropy as a predictive signal and incorporates difficulty residual modeling to effectively identify contaminated or infeasible tasks within environments. Through textual entropy analysis, fine-grained task feature extraction, and validation across multiple domains, the method significantly enhances the reliability of difficulty prediction across 17 diverse agent scenarios—spanning coding, mathematical reasoning, and web navigation—and further aids in uncovering flaws in environment design.
研究通过分析CoderForge-Preview数据集中的任务特征,使用集成方法、SHAP归因和效应量分析,探究了软件问题解决任务的难度,并发现难度主要受补丁碎片化和仓库规模影响。
This study addresses the inference bias that arises when organizations evaluate expert competence solely based on project success or failure, a distortion attributable to differences in task bundling architectures. Drawing on Bayesian inference and Blackwell’s partial order theory, this work compares the informational value of bundled, outsourced, and unbundled projects, derives reliability thresholds, and quantifies the statistical costs associated with coarse-grained aggregation. The primary contribution lies in establishing, for the first time, a precise threshold relationship between task architecture and learning efficiency, revealing that under fixed workloads, the advantage of bundling strengthens as task scope expands. Furthermore, it demonstrates that bundling dominates when external technologies are unreliable, and that interim auditing can significantly broaden its region of superiority.
This study addresses a key limitation in existing occupational exposure metrics, which fail to distinguish between human roles in task execution versus evaluation, thereby obscuring the nuanced impact of artificial intelligence (AI) on employment. Leveraging 19,265 task descriptions from O*NET, the authors propose a reproducible “execution share” measure that explicitly differentiates AI capabilities, routine task intensity, and human execution roles, enabling the construction of an occupation-level AI capability exposure index. Panel data regression analyses reveal persistently subdued employment growth since 2012 in white-collar occupations intensive in execution tasks. Moreover, a pronounced gradient effect of AI capability emerged after 2022, though causal inference remains tentative. This work offers a novel measurement framework and empirical evidence to better understand the heterogeneous labor market effects of AI.
This study addresses the capability degradation and goal-contract invalidation of AI agents caused by shifting conditions during model transfer, cross-domain deployment, or scaling. To mitigate these issues, we propose a three-layer interactive calibration methodology that decouples trainable policies from frozen configurations, establishing a multidimensional constraint system encompassing information preservation, execution framework adaptation, and user acceptance. Technically, the approach integrates semantic checkpoint repair, tool substitution, local replanning, and output contract enforcement mechanisms. For evaluation, we introduce a factorial testing holdout system to prevent aggregated gains from masking localized failures. This work provides a methodological framework enabling the coexistence of shared standards and market-specific adapters in global e-commerce scenarios, with empirical validation reserved for future research.