longitudinal risk evaluation

Designs and implements evaluation frameworks and experimental protocols that measure how risks change and appear over extended interactions, including methods to aggregate retrospective assessments across turns and detect risks that emerge only after prolonged exposure. Builds analyses and metrics to estimate the temporal horizon required for stable risk estimates and to compare model behavior across risk dimensions and development or interaction stages.

longitudinalriskevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.06
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing causal models struggle to distinguish between the immediate and persistent effects of interventions in time-dynamic systems, particularly when such interventions alter the system’s equilibrium behavior. This work proposes a novel paradigm grounded in system and state representations, integrating causal directed acyclic graphs, the potential outcomes framework, and dynamic systems theory. By introducing an equilibrium-state assumption and employing state-space modeling, the study reformulates the causal inference framework to better capture temporal dynamics. It innovatively defines an equilibrium-oriented “zero effect” concept and combines it with a strategic selection of time points to enable valid identification of time-varying causal parameters. The approach establishes clear criteria for categorizing causal effects under dynamic interventions, substantially enhancing the interpretability and practical utility of causal inference in equilibrium analysis.

causal effectsequilibrium behaviorlasting effects

This study addresses a critical gap in current safety evaluations of AI systems for mental health, which often overlook the temporal dynamics of human–AI interactions and thus fail to detect clinical risks arising from dialogue sequencing, cumulative effects, or delayed deterioration. To remedy this, the authors propose SCOPE-MH, a novel framework that formally defines “temporal safety unidentifiability” and introduces safety assessment principles grounded in the preservation of temporal evidence. Leveraging formal modeling, sequential dialogue analysis, and the expert-annotated AnnoMI dataset, SCOPE-MH establishes domain-specific evaluation criteria tailored to mental health applications. Empirical results demonstrate that conventional turn-level scoring methods miss key failure modes, whereas SCOPE-MH effectively captures temporally dependent risks, substantially enhancing the clinical relevance and reliability of safety assessments.

Mental Health AISafety EvaluationSequence Dependence

This study addresses the challenge of managing residual risks that cannot be fully eliminated, noting that existing qualitative analyses lack actionable dynamic management mechanisms. To bridge this gap, the authors propose formalizing the Bowtie risk diagram as a directed acyclic graph (DAG) capable of supporting Bayesian inference and causal intervention. By incorporating safety-state semantics and explicit intervention nodes, and integrating expert probability assessments with do-calculus, the framework enables risks to be observable, quantifiable, and intervenable. The work introduces Realtime Risk Studio—a modeling tool—and Probability Capture—a method for eliciting probabilistic judgments—to construct, for the first time, an executable real-time risk reasoning model. Validation in an instant payment gateway scenario demonstrates the efficacy of transforming Bowtie diagrams into DAGs, fusing noisy expert probabilities, and performing “What-if” causal intervention analyses.

causal analysisoperationalizationprobabilistic modeling

Assessing model error in counterfactual worlds

Nov 30, 2025
EH
Emily Howerton
🏛️ Princeton University | University of North Carolina at Chapel Hill

This work addresses the evaluation of model error under counterfactual scenarios, distinguishing between prediction bias and miscalibration to accurately quantify a model’s actual decision-making value. We propose and systematically compare three counterfactual error estimation methods, leveraging simulation studies and error decomposition to characterize the differential impacts of distinct error components on decision utility. Innovatively, we develop an error–decision-value mapping framework that identifies and validates key scenario-design factors governing counterfactual prediction credibility—namely, intervention strength and degree of covariate distribution shift. Results demonstrate that calibration error’s contribution to decision loss is frequently underestimated, while bias-dominated error becomes markedly more detrimental under strong interventions. The study provides a principled, actionable methodology and practical guidelines for assessing model reliability in high-stakes decision-making contexts.

Assessing model value for decision makingEstimating model error in counterfactual scenariosEvaluating scenario projections retrospectively

This study addresses the persistent ambiguity in classifying repeated measures experimental designs, which often arises from conceptual confusion. To resolve this issue, the authors systematically clarify the core characteristics of such designs and propose a novel classification framework grounded in experimental units and randomization strategies. For the first time in this context, Hasse diagrams are introduced to visually represent the hierarchical structure of these designs. This approach effectively distinguishes among various types of repeated measures designs, eliminates terminological ambiguities, and substantially enhances both the rigor and interpretability of experimental planning and reporting.

experimental designexperimental unitsHasse diagrams

Latest Papers

What's happening recently
View more

Current governance paradigms typically assume that the safety properties of foundation models are preserved after fine-tuning; however, this assumption lacks systematic validation in high-stakes domains such as healthcare and law. This study presents the first comprehensive, multidimensional safety evaluation of 100 fine-tuned models—including widely deployed domain-specific models and controlled experimental variants—across both general and domain-specific benchmarks. The findings reveal that fine-tuning frequently induces significant and heterogeneous shifts in safety: improvements along certain metrics often coincide with severe degradation in others. These results challenge the prevailing practice of relying solely on base model safety assessments and underscore the critical need for independent, thorough safety re-evaluation of fine-tuned models before deployment in high-risk applications.

fine-tuningfoundation modelshigh-stakes domains

This study addresses the limitations of current safety evaluations, which predominantly rely on isolated multiple-choice setups and overlook the real-world impact of agent scaffolding on model safety. Through a large-scale controlled experiment (N = 62,808), we systematically assess four scaffolding architectures—including Map-Reduce—across evaluation formats (open-ended vs. multiple-choice) on state-of-the-art language models. We find that evaluation format exerts a far stronger influence on safety scores than scaffolding effects themselves. Critically, strong model–scaffolding interactions lead to complete reversals in safety rankings across benchmarks (G = 0.000). Employing preregistration, evaluator blinding, TOST equivalence testing, and generalizability analyses, we demonstrate that safety must be evaluated for each specific model–deployment configuration: Map-Reduce significantly reduces safety (NNH = 14), whereas other architectures remain equivalent within ±2 percentage points.

benchmarkingevaluation formatlanguage models

The rapid expansion of AI deployments has put organizational leaders in a decision maker's dilemma: they must govern these technologies without systematic evidence of how systems behave in their own environments. Predominant evaluation methods generate scalable, abstract measures of model capabilities but smooth over the heterogeneity of real world use, while user focused testing reveals rich contextual detail yet remains small in scale and loosely coupled to the mechanisms that shape model behavior. The Forum for Real World AI Measurement and Evaluation (FRAME) addresses this gap by combining large scale trials of AI systems with structured observation of how they are used in context, the outcomes they generate, and how those outcomes arise. By tracing the path from an AI system's output through its practical use and downstream effects, FRAME turns the heterogeneity of AI in use into a measurable signal rather than a trade off for achieving scale. FRAME establishes two core assets to accomplish this: a Testing Sandbox that captures AI use under real workflows at scale and a Metrics Hub that translates those traces into actionable indicators.

Immune checkpoint inhibitors often induce time-varying, heterogeneous survival effects that cannot be adequately captured by conventional hazard ratios. To address this limitation, this study proposes a milestone-based analytical framework that disentangles early outcomes from long-term survival by integrating milestone survival probabilities with the tau statistic. This approach effectively characterizes dynamic treatment patterns under non-proportional hazards, where short-term risks and late-onset benefits coexist. Applied to three phase III clinical trials, the method successfully uncovers time-dependent therapeutic effects overlooked by traditional analyses, precisely identifying the onset of clinical benefit and distinguishing between short- and long-term efficacy profiles. The proposed framework thus offers a novel paradigm for the accurate evaluation of immunotherapies.

heterogeneous survival responsesimmunotherapy trialsmilestone survival

研究探讨了前沿模型在行动前是否寻求安全证据,通过SAFE基准测试不同模型的证据获取策略,发现检索成本和问题严重性对决策影响较大。

Decision MakingEvidence AcquisitionFrontier Models

Hot Scholars

PM

Pulak Mehta

New York University
Useable securityusable privacy
JM

James M. Hyman

Emeritus Professor of Mathematics, Tulane University
MathematicsApplied MathematicsNumerical AnalysisMathematical Biology
EZ

Evangelia Zve

Sorbonne University
topic modelingsocial network analysisdisinformation
DY

Dong-Yeun Koh

Associate Professor of Chemical and Biomolecular Engineering, KAIST
Separation ProcessMembranesAdsorptionDirect Air Capture