artifact validation

Designs and implements methods and tools to detect, parse, analyze, verify, and reject artifacts, including deterministic verifiers and compliance-checking components, along with artifact-detection methods for flagging non‑conforming items. Builds systems for artifact packaging, versioning, simulation, tracking, and mitigation, and develops metrics and processes to quantify deviations and tradeoffs and to report or correct artifacts that violate declared constraints.

artifactvalidation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.62
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$195K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that large language models (LLMs) struggle to accurately calibrate trust across software artifacts—such as code, documentation, and tests—when inconsistencies arise, a nuance overlooked by existing evaluation methods. To systematically quantify LLMs’ trust mechanisms in multi-source software conflicts, we propose TRACE, a framework featuring blind perturbation generation, structured trust trajectory collection, and multidimensional assessment encompassing quality judgment, inconsistency detection, attribution, and prioritization. Experiments across seven models and 22,339 human-validated trajectories reveal that while models effectively identify explicit documentation errors (67–94% accuracy), their detection performance drops substantially—by 7 to 42 percentage points—when only implementation drift is present. Moreover, models consistently exhibit poor confidence calibration in such scenarios.

artifact-level reasoningcalibrationconflicting software artifacts

Current automated detection tools struggle to meet regulatory practice demands due to insufficient transparency, interpretability, and the inability to map findings to specific legal provisions, resulting in a disconnect between academic research and enforcement applications. Through in-depth interviews with nine regulatory practitioners and an analysis integrating regulatory workflows with technical feasibility, this study systematically uncovers, from a regulatory perspective, the practical barriers to deploying automated tools for identifying deceptive designs. The work proposes a human-in-the-loop compliance review framework that is user-need-driven, supports the entire investigative workflow, and aligns both research and regulatory objectives, offering critical guidance for developing automated detection systems that genuinely meet real-world enforcement requirements.

automated detectiondark patternsdeceptive design patterns

This work addresses the problem of global inconsistency in multi-component intelligent agent releases, where local validation passes but cross-component relational integrity fails due to the absence of holistic consistency guarantees. To tackle this, we propose the Schema-SIP Relational Consistency (SIP-RC) framework—the first systematic approach to formally define and mitigate relational inconsistency faults in multi-component deployments. SIP-RC models release packages as graph structures and integrates schema documentation with product contract principles to enable cross-component relational verification. Key mechanisms include declarative–evidential linkage, decision authority scoping, provenance tracking of derived components, and byte-level consistency checks. Preliminary experiments demonstrate the feasibility of the proposed framework, offering a practical and actionable paradigm for ensuring relational consistency in intelligent agent releases.

Agent SystemsMulti-Artifact ReleasesPackage Consistency

A Unit Proofing Framework for Code-level Verification: A Research Agenda

Oct 18, 2024
PC
Paschal C. Amusuo
🏛️ Purdue University | Michigan State University

Existing code-level formal verification tools scale poorly to large-scale software, while mainstream unit-level verification relies heavily on manual effort, often missing critical defects. This paper proposes the “Unit Proof Framework” research agenda—the first systematic definition of a unit verification paradigm supporting automated decoupling and independent verification of code units. Methodologically, it integrates formal verification, program analysis, modular verification, and automated toolchain design, with deep alignment to industrial development practices (e.g., AWS workflows). Its core contributions include: (1) establishing a scalable, engineering-friendly unit verification methodology; (2) characterizing a taxonomy of key technical challenges; (3) overcoming bottlenecks inherent in manual verification; and (4) significantly improving early detection of code-level defects. Collectively, this work lays the theoretical foundation and provides a practical technical pathway for building high-assurance, deployable automated verification infrastructure.

Automating unit proofing to reduce manual errorsEarly detection of implementation defects in verificationEnsuring code-level correctness in large-scale software

Latest Papers

What's happening recently
View more

This work addresses the problem of implementation drift in evolving distributed systems, where runtime behavior gradually deviates from the original design. To tackle this issue, the paper proposes a design conformance assessment method based on distributed tracing data. It introduces, for the first time in the domain of distributed systems, conformance checking techniques from process mining, leveraging runtime traces collected via the OpenTelemetry standard and automatically comparing them against behavioral models defined at design time to quantify their alignment. The key contribution lies in establishing persistent, monitorable conformance metrics that enable continuous, automated evaluation of deviations between system implementation and design. This approach is readily applicable to modern distributed systems widely adopting OpenTelemetry for observability.

design conformancedistributed systemsimplementation drift

This work addresses the limited reliability of structured security artifacts—such as KQL queries and MITRE ATT&CK mappings—generated by large language models, which often fall short of production-grade requirements. To bridge this gap, the authors propose a lightweight verification framework that shifts the focus of quality assurance from generation to validation. The core innovations include a hybrid testing strategy integrating test-driven generation, deterministic program verification, and semantic evaluation by large language models, alongside an interpretable judging mechanism distilled from expert decision distributions. Deployed in Microsoft Sentinel’s production environment, the framework significantly enhances the reliability of three critical types of security artifacts, establishing professional-grade and scalable validation standards.

artifact generationLarge Language Modelsproduction reliability

This study addresses environment inconsistencies, expanded supply chain attack surfaces, and weak compliance arising from redundant builds in cloud deployments by proposing an artifact promotion control model that establishes a “build once” principle. Methodologically, the work rigorously distinguishes artifact from environment identities, demonstrates that secret injection compromises artifact integrity, derives that release roles require no production credentials, and implements end-to-end autonomous governance on AWS in accordance with NIST SP 800-204D. Experimental results show that the model supports fully autonomous releases via a single command, completing 22 deployments in the first month with individual rollbacks requiring only 36 seconds. These findings indicate a significant reduction in operational complexity alongside strengthened integrity guarantees for cloud deployment pipelines.

Artifact PromotionBuild-onceCloud Deployment

This study addresses the challenges posed by rapid evolution in digital forensic systems and tools, which induces drift in evidentiary behaviors and tool outputs, thereby undermining result reproducibility and trustworthiness. To mitigate this, the authors propose a test-driven forensic methodology that introduces state-transition testing for causal attribution, encoding forensic expectations as executable specifications. The approach integrates virtual machine environments with computer vision–guided GUI automation to simulate authentic user interactions and verify system state changes. An open web platform is developed to facilitate sharing and replication of experiments. The method’s efficacy is demonstrated through five case studies, including a regression analysis across 25 versions of Autopsy, which uncovered numerous undocumented, substantial changes in its reporting output.

artifact driftdigital forensicsregression

This study addresses the accountability deficit in agent development arising from the misalignment between platform controls and service provider terms. By analyzing four categories of tools and policy documents, we map workflow responsibilities and propose a novel grid model distinguishing verification mandates from executors. This framework reveals structural deficiencies in approval mechanisms, demonstrating that responsibility gaps have evolved from human oversight to inherent product attributes. Empirical findings indicate conflicting accountabilities across layers, contradictory attribution logic, and insufficient efficacy of approval artifacts. To support further research, we release a comprehensive dataset and validation scripts as open-source resources. Collectively, this work provides both theoretical grounding and empirical evidence necessary for reconstructing accountability frameworks in agent-based software systems, highlighting the urgent need to address systemic rather than incidental failures in current governance architectures.

AccountabilityAgentic Software DevelopmentCode Review

Hot Scholars

OP

Oiwi Parker Jones

Applied Artificial Intelligence and Clinical Neurosciences, University of Oxford
AINeuroscienceDeep LearningSpeech Recognition
XJ

Xiaojun Jia

Nanyang Technological University
Explainable AIRobust AIEfficient AI
PL

Peng Liang

School of Computer Science, Wuhan University
Software EngineeringSoftware ArchitectureEmpirical Software Engineering
ZJ

Zhi Jin

Sun Yat-Sen University, Associate Professor
MS

Muhammad Shafique

Professor, ECE, New York University (AD-UAE, Tandon-USA), Director eBRAIN Lab
Embedded Machine LearningBrain-Inspired ComputingRobust & Energy-Efficient System DesignSmart