Score
Designs and formalizes behavioral and synthetic tasks — including auxiliary, proxy, probing, turn-taking, and action-sequence tasks — along with the sampling, scenario, and specification rules that generate task instances. Specifies evaluation and training objectives (for example expected-future-loss or other surrogate objectives), constructs proxy tasks and online or sampling-based selection procedures, and analyzes proxy-to-target relationships and task-based interventions to guide model training, hyperparameter selection, and evaluation without relying on explicit posterior estimation.
This work addresses a fundamental limitation in conventional Bayesian experimental design, which relies on prior-to-posterior uncertainty reduction and yields an intractable objective that is doubly hard to evaluate and poorly aligned with downstream tasks. By reframing the problem through decision theory, the authors formulate it as optimizing the expected future loss (EFL) of downstream actions, thereby reducing the objective to a singly intractable form that obviates explicit posterior or marginal likelihood computation. They introduce a stochastic gradient method that jointly optimizes both the experimental design and the action policy, requiring only samples from the joint parameter–data model and evaluations of the loss function. This approach naturally accommodates implicit modeling and task-specific customization, demonstrating marked improvements over existing methods in both optimization efficiency and task adaptability.
This work addresses the challenge of efficiently exploring universal models of human behavior in high-dimensional task spaces, where conventional random sampling proves inadequate. Focusing on binary sequence prediction tasks, the authors propose an adversarial construction strategy that leverages a hidden Markov model (HMM) to represent the task space and actively generates task instances most likely to elicit novel behavioral patterns. By concentrating experimental design on regions critical for behavioral diversity, this approach substantially outperforms random sampling, uncovering a greater number of qualitatively new phenomena with fewer experiments. The method thus establishes an efficient and practical paradigm for constructing generalizable models of human behavior.
Access to real-world Applied Behavior Analysis (ABA) session data is severely limited by privacy constraints, hindering the training of AI models in this domain. To address this challenge, this work proposes a deterministic synthetic data generation method grounded in authoritative ABA taxonomies, enabling—for the first time—the construction of fully traceable instruction-tuning datasets. The approach supports two core tasks: instructional program generation and multi-session behavioral trajectory interpretation, while integrating standard ABA paradigms such as Discrete Trial Teaching and Natural Environment Teaching. The resulting TRACE dataset comprises 2,999 structurally transparent and content-compliant samples, partitioned into training, validation, test, and reasonableness-check splits according to predefined ratios. Both code and data are publicly released under CC BY-NC 4.0 and MIT licenses.
Existing query-first data synthesis approaches struggle to generate valid and executable tool-use sequences. This work proposes SyntheticAgentTraceQA, a novel framework that introduces an "execution-first" paradigm: it first constructs high-level workflows, maps and validates feasible tool trajectories, and then synthesizes corresponding natural language tasks and reference answers. The method integrates dependency-aware tool assignment, trajectory validation in a controlled environment, and reasoning-augmented annotation generation, followed by fine-tuning and evaluation using the Qwen model. Experimental results demonstrate that this framework substantially improves large language model (LLM) agents’ tool execution accuracy, trajectory consistency, and answer quality. Furthermore, the study reveals that masked supervision outperforms full supervision for models at the 9B scale.
This work addresses the scarcity of high-quality, verifiable, and diverse task data that hinders large-scale training of terminal-based intelligent agents. Existing synthetic approaches often suffer from a disconnect between task generation and execution and rely heavily on pre-existing repositories, limiting diversity and scalability. To overcome these limitations, the authors propose modeling the task synthesis process itself as a Terminal-Bench–formatted terminal task, enabling closed-loop iterative generation, execution, and validation within real containerized environments. Their method enhances diversity and realism through multi-stage task specification, decoupling of task dimensions, and augmentation with external materials, while employing an LLM-as-Judge mechanism for quality filtering. Using only 3,221 synthesized trajectories for fine-tuning, Qwen3-14B and Qwen3-32B achieve Avg Pass@1 scores of 22.5% and 31.8%, respectively, on Terminal-Bench 2.0—significantly outperforming concurrent methods with substantially less training data.
Although proxy metrics are commonly employed to substitute for hard-to-observe primary outcomes, their systematic biases often undermine inference validity and distort confidence intervals. This work proposes the *proxymate* framework—a systematic, modular four-layer diagnostic and correction system encompassing representativeness, unit, estimation, and domain levels—that maps specific failure modes to targeted remediation strategies. Implemented through hierarchical diagnostics, bias-correction algorithms, and an open-source Python toolkit, the framework supports diverse applications including experimentation, monitoring, and prevalence estimation. Validated across thousands of Meta experiments and multiple product lines on millions of proxy–primary outcome pairs, *proxymate* significantly enhances inference reliability and accelerates decision-making.
This work addresses the frequent failure of multi-task Bayesian optimization due to inaccurate estimation of cross-task correlations—even under the simplified assumption of affine relationships between source and target tasks. The study systematically identifies two root causes: alignment errors induced by task normalization with finite samples and insufficient identifiability of marginal likelihood under non-overlapping experimental designs. To mitigate these issues, the authors propose modeling task-specific means and scales as learnable parameters and introduce three conservative remedies: enforcing non-negative task covariance constraints, adopting partially co-located experimental designs, and refining the correlation estimation mechanism. The method successfully recovers single-task performance on affine synthetic benchmarks and hyperparameter transfer tuning tasks, yet limitations persist in more complex settings, particularly those involving ranking-based objectives or latent contextual structures.
Current test-time training (TTT) approaches for large language models predominantly rely on proxy metrics such as perplexity, which inadequately capture the models’ true capabilities in deployment scenarios—particularly regarding memory retention, personalization, or sparse learning. This work proposes the first calibration-based behavioral evaluation framework explicitly designed to validate deployment-oriented memory claims. By integrating an evidence hierarchy and standardized protocols, the framework aligns model memory assertions with observable behaviors and clearly distinguishes between streaming/domain adaptation, bridging internalization, and deployment behavior learning. Employing explicit memory baselines, failure taxonomies, single-step LoRA updates, controlled nonce facts, multi-scale Qwen3 models, and free-recall generation experiments under sparse-fact settings, the study reveals a critical disconnect: despite improvements in support and answer loss, free-recall accuracy remains at zero, empirically exposing a significant gap between proxy metric gains and genuine behavioral competence.