design diagnostic evaluations

Design and implement evaluation artifacts—targeted diagnostic tasks, controlled tests, probing tasks, and statistical indicators—that make specific model or system behaviors measurable and reveal failure modes. Select and calibrate diagnostic metrics, sensitivity thresholds, supervision signals, and experimental protocols to quantify, compare, and interpret performance and root causes.

designdiagnosticevaluations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.09
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$217K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of effectively monitoring early-stage agent systems, where structural flaws often obscure task-level errors. The authors propose a three-dimensional (quality, suitability, efficiency) and three-granularity (intra-run, inter-run, structural) monitoring and triaging framework tailored for low-maturity agent systems. They introduce a novel system maturity staging model based on the coefficient of variation and monitoring granularity, integrated with a severity classification adapted from FMEA to guide human review. The resulting transferable monitoring architecture supports document-driven, multi-stage workflows, enhanced by a synthetic testbed with controlled error injection. Experimental results demonstrate that structural defects significantly mask task-level signals; 97% of issues can be automatically traced, with only 2% requiring human intervention, and each granularity level precisely identifies its corresponding defect type (coefficients of variation: 0.02, 1.25, and 0.00, respectively).

Agentic SystemsMonitoringStructural Defects

Technique to Baseline QE Artefact Generation Aligned to Quality Metrics

Nov 18, 2025
EF
Eitan Farchi
🏛️ IBM Research | IBM Consulting

This study addresses the uncontrolled quality of quality engineering (QE) artifacts—such as requirements specifications, test cases, and Behavior-Driven Development (BDD) scenarios—automatically generated by large language models (LLMs). We propose an iterative optimization framework integrating forward generation, backward generation, and rubric-guided scoring to enhance artifact quality along four dimensions: clarity, completeness, consistency, and testability. Our approach enables automated, quantitative, and reproducible quality assessment and improvement. Evaluated across 12 real-world projects, the method significantly improves output stability: it preserves high quality under high-quality inputs and substantially outperforms baselines under low-quality inputs. The core contribution is the first integration of backward generation with structured rubric-based guidance, establishing a closed-loop, artifact-centric quality enhancement paradigm for QE.

Ensuring generated requirements and test cases meet quality metricsEstablishing baselines for automated QE artefact quality evaluationValidating LLM outputs through reverse generation and iterative refinement

This work addresses the challenge that large language models (LLMs) struggle to accurately calibrate trust across software artifacts—such as code, documentation, and tests—when inconsistencies arise, a nuance overlooked by existing evaluation methods. To systematically quantify LLMs’ trust mechanisms in multi-source software conflicts, we propose TRACE, a framework featuring blind perturbation generation, structured trust trajectory collection, and multidimensional assessment encompassing quality judgment, inconsistency detection, attribution, and prioritization. Experiments across seven models and 22,339 human-validated trajectories reveal that while models effectively identify explicit documentation errors (67–94% accuracy), their detection performance drops substantially—by 7 to 42 percentage points—when only implementation drift is present. Moreover, models consistently exhibit poor confidence calibration in such scenarios.

artifact-level reasoningcalibrationconflicting software artifacts

This work addresses the efficiency bottleneck of manual reproducibility reviews in safety-critical domains such as the Internet of Things and cyber-physical systems, which hampers research transparency and deployability. The paper presents the first systematic framework leveraging large language models (LLMs) to automate reproducibility assessment by integrating natural language understanding, code generation, sandboxed environment auto-configuration, and rule-guided flaw detection. This approach enables reproducibility scoring, automatic execution environment setup, and identification of methodological flaws. Experimental results demonstrate that the proposed method achieves over 72% accuracy in reproducibility judgment, automatically constructs executable environments for 28% of runnable artifacts, and attains F1 scores exceeding 92% across seven common categories of methodological defects, substantially enhancing both the efficiency and quality of reproducibility review.

Artifact EvaluationCPSCybersecurity

Latest Papers

What's happening recently
View more

Current AI-assisted scientific writing lacks auditable generation processes and mechanisms for accountability, undermining the verifiability of research credibility and compliance. This work proposes a novel auditing paradigm embedded directly within the production workflow, enforcing end-to-end traceability, immutability, and third-party reproducibility of AI involvement through preregistered blind-spot indicator cards, sealed execution environments, and automated gatekeeping intercepts. Core technical components include Git-sealed lineage anchoring, hash-bound provenance tracking, red-flag interception protocols, cross-model role isolation, and programmatic assembly. In experimental validation, one project was automatically terminated when preregistered confirmatory tests triggered a No-Go decision. An open-source toolkit is released to enable independent recomputation of all core audit metrics by third parties.

AI AccountabilityAuditable AIProvenance

This work addresses the limited interpretability and accountability of large language models (LLMs) in root cause analysis, which hinder their applicability in high-stakes operational settings requiring rigorous evidence chains, hypothesis comparison, and uncertainty handling. The authors propose JustDiag, a diagnostic argumentation engine that introduces, for the first time, an explicit modeling of the diagnostic reasoning process into root cause analysis. JustDiag structures and maintains states such as evidence, findings, competing hypotheses, conflicts, and follow-up checks to enable traceable and auditable inference, complemented by a calibration mechanism that explicitly accounts for uncertainty. Integrating LLMs with a structured reasoning framework, the approach employs a two-tier evaluation protocol to assess both outcome and reasoning quality. Experiments on 66 real-world incidents demonstrate that JustDiag significantly outperforms non-argumentative baselines in both outcome and process scores, exhibiting superior uncertainty retention despite a slightly lower completion rate.

accountabilitydiagnostic justificationincident response

Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.

auditable recordsevidence validationindustrial machine learning

Hot Scholars

PK

Pilsung Kang

Dankook University, Sunmoon University, Virginia Tech, Samsung
computational scienceparallel computing
JR

Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning
DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
ZW

Zihao Wu

University of Georgia
Brain-inspired AIArtificial General IntelligenceNLPMedical Image Analysis