Score
Designs and implements instrumentation, metrics, models, and analytical pipelines to record, decompose, and evaluate temporal trajectories of individual roles and the composition of populations, including trajectory logging, diagnostics, error metrics, and trajectory decomposition. Builds and applies trajectory models and bespoke metric designs to quantify role evolution and population turnover, detect and explain changes in role assignments or cohort composition over time, and produce diagnostics at both the individual-trajectory and population level.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
AI systems produce large volumes of logs as they interact with tools and users. Analysing these logs can help understand model capabilities, propensities, and behaviours, or assess whether an evaluation worked as intended. Researchers have started developing methods for log analysis, but a standardised approach is still missing. Here we suggest a pipeline based on current best practices. We illustrate it with concrete code examples in the Inspect Scout library, provide detailed guidance on each step, and highlight common pitfalls. Our framework provides researchers with a foundation for rigorous and reproducible log analysis.
Existing performance measurement frameworks struggle to simultaneously satisfy customizability, interpretability, and mathematical tractability in interdisciplinary contexts. Method: This paper proposes a goal-oriented, customizable metric construction framework featuring a novel “base metric–auxiliary metric” dichotomy. Integrating utility theory and multi-criteria decision analysis, it introduces an uncertainty-aware utility function and establishes a systematic metric decomposition–synthesis workflow. Contributions: (1) It reduces reliance on complex mathematical formalisms, enhancing applicability under resource constraints or high uncertainty; (2) it ensures metric transparency, traceability, and domain adaptability; and (3) it enables quantitative assessment of goal attainment, real-time progress monitoring, and downstream statistical modeling and decision optimization. The framework has been empirically validated across diverse disciplines, demonstrating generality and extensibility.
Estimating the rate of change in nonlinear trajectories under individually scheduled, unequally spaced longitudinal measurements remains challenging—existing models struggle to jointly estimate dynamic change parameters and theory-driven substantive parameters. To address this, we propose a novel framework that conceptualizes the rate of change as the area under a time-varying functional curve, approximating the average rate within each interval by the instantaneous rate at its midpoint. This enables simultaneous estimation of both change and substantive parameters. The method is implemented within a latent-variable structural equation modeling framework using OpenMx or Mplus 8, integrating numerical integration with interval-specific approximations. Simulation and empirical studies demonstrate high accuracy, robustness, and the ability to derive both baseline-level and interval-specific change metrics. Accompanying open-source code ensures flexibility and reproducibility. The approach substantially enhances theoretical interpretability and practical utility for modeling nonlinear longitudinal processes.
This work addresses the frequent failure of AI agents in production environments due to errors or omissions in contextual sources such as system prompts, knowledge bases, or tool descriptions—a problem exacerbated by the reliance on manual log inspection for maintenance, which does not scale. To overcome this, the authors propose an automated context engineering framework that operates without explicit user feedback by mining implicit dissatisfaction signals (e.g., corrections, rephrasings, or task abandonment) from historical interaction trajectories. The framework integrates multi-component causal attribution with an exploratory validation strategy to automatically diagnose and repair contextual defects. Key contributions include the first verifiable simulation benchmark for context debugging, a taxonomy of six failure types, and a causal attribution and active verification mechanism applicable across heterogeneous context sources. Experiments demonstrate 72.7% root-cause attribution accuracy and 82% end-to-end repair effectiveness over 60 dissatisfaction trajectories, confirming the approach’s capability for efficient self-repair of context-layer faults.