Score
Designs and implements structured, turn-level representations (trace data structures) for debate or multi-turn argumentation processes that record each turn's claims, evidence, reasoning steps, actor identities, and relevant metadata. Builds diagnostics and analysis tools that inspect these traces to detect and attribute compounding failures, sycophantic agreement patterns, evaluation-collapse triggers, and to enforce strict trace-level constraints for auditing, evaluation, and iterative correction.
This work addresses the lack of auditable reasoning mechanisms in large language models, which rely solely on statistical correlations and thus struggle to support trustworthy agent decision-making. To bridge this gap, the authors propose TRACE, a novel framework that integrates formal reasoning theory into a typed, versioned structure for recording reasoning trajectories. By introducing a record-consumer contract mechanism, TRACE transforms reasoning traces into actionable interfaces. The framework features a causally specialized TraceRecord structure, an eight-stage reference writer, and a gated priority metric, complemented by the TRACE-Bench protocol and counterfactual regret analysis. Experiments in music instruction and flood rescue scenarios demonstrate TRACE’s ability to effectively disentangle associative, interventional, and normative reasoning, thereby establishing a comprehensive, auditable decision-support architecture.
论文针对LLM数据代理在结构化任务中可能产生无效推理路径的问题,提出通过执行合约来实现可审计的Trace Integrity方法,以评估输出背后的计算是否可靠。
This work addresses the challenge of deploying generative AI in high-stakes decision-making, where hallucinated reasoning, unsupported claims, and weak traceability often preclude compliance with certification-grade accountability requirements. To bridge this gap, the authors propose a “compliance-by-construction” architecture that uniquely integrates typed argumentation graphs, retrieval-augmented generation (RAG), formal verification kernels, and W3C PROV-based provenance tracking. This framework ensures that every AI-generated claim is grounded in authoritative evidence and subjected to rigorous inference constraints before being admitted into official decision records. Empirical evaluation demonstrates that the architecture effectively blocks unsubstantiated assertions from entering the decision pipeline and substantially improves the efficiency of constructing compliant, auditable arguments, thereby enabling controlled, verifiable, and accountable use of generative AI in high-assurance settings.
Current large language model (LLM) agents lack verifiability, debuggability, and auditability, and relying solely on the accuracy of final answers fails to reveal their underlying reasoning. To address this, this work proposes the first unified provenance framework for LLM agents, systematically modeling causal relationships in tool usage, memory access, and environmental interactions. It introduces a comprehensive provenance taxonomy encompassing source, granularity, representation format, and trust functions. By integrating provenance-aware representation modeling, evidence attribution, runtime safeguards, provenance-informed memory management, and trajectory observability analysis, the study shifts the evaluation paradigm from outcome correctness to process accountability. The framework consolidates existing benchmarks to define a clear pathway for process-level trustworthiness assessment and highlights key challenges, including standardized trajectory schemas, semantic-level provenance, and privacy-preserving auditing.
This work addresses the lack of traceable and tamper-resistant transparency mechanisms in large language models (LLMs) deployed in high-stakes decision-making contexts, which undermines accountability. To bridge this gap, the paper introduces the first LLM lifecycle auditing framework that integrates technical provenance with governance records. It proposes a reference architecture enabling cross-organizational traceability and implements a lightweight, open-source Python-based auditing layer. By leveraging append-only logs, event emitters, structured metadata, and an auditor interface, the system seamlessly integrates into existing LLM workflows with minimal intrusiveness. This design ensures complete, tamper-evident traceability across critical stages—including training, deployment, and monitoring—thereby facilitating robust accountability and responsibility attribution throughout the model’s lifecycle.
Industrial research agents often generate experimental trajectories containing invalid or incomplete information, rendering them unreliable for direct decision-making. This work proposes an evidence-oriented framework that automatically transforms such trajectories into structured evidence through a context-isolated generate–verify–repair pipeline. The approach introduces intervention-level claim categorization—distinguishing actionable repairs, diagnostic safeguards, and retained discoveries—and incorporates end-to-end provenance tracking to enable claim scoping and auditability. Experimental results demonstrate that the resulting candidate solutions outperform existing baselines. Audits further reveal that trajectory evolution is non-monotonic, and that applicability assessment constitutes a key performance bottleneck for the controller.
为解决大语言模型代理输出的审核问题,提出LEDGER系统,通过构建分层追踪图来连接声明与支持性操作、工件和检查,以提高审计效率。
This study addresses the failure of regulatory accountability caused by accumulating “interpretability debt” in production AI systems. To this end, it proposes the TRACE governance framework, which introduces a novel Explainability Debt Score (EDS) and a seven-tool architecture. Through DART trajectory analysis, SHIV validation, and FDE feature drift evaluation, the framework enables systematic quantification, tracking, and remediation of interpretability debt. Empirical results demonstrate that TRACE accurately predicts debt evolution trends, reveals explanation asymmetries in high-risk decision-making, and comprehensively aligns with EU AI Act compliance requirements. This work establishes a new paradigm for AI explainability governance, providing methodological support for sustained regulatory compliance in production-grade systems.
This study addresses cross-object tampering and audit failures in tool-use agents caused by the separation of declarations, evidence, and execution logs. We propose a declaration-anchored execution contract mechanism that jointly binds declarations, source snippets, and execution prefixes to enable replayable integrity auditing. A core innovation is the first architecture decoupling structural from semantic support, defining seven independently testable properties to ensure evidence immutability while integrating domain-isolated commitments, hash verification, deterministic validation, and conflict-aware guards. Experimental results demonstrate that this approach achieves a 99.61% detection rate across 1,280 attacks, with a guard F1-score of 0.8865 and a false acceptance rate of only 7.29%.
Chain-of-Thought (CoT) is frequently regarded as a faithful record of model reasoning, yet its natural language form resists mechanical verification. This study leverages the iGSM benchmark to decouple "correct answers" from "valid reasoning" by programmatically inspecting generated trajectories step-by-step. The findings reveal that even when final answers are correct, 31.6% of the most challenging instances are accompanied by semantically invalid CoTs. Furthermore, accuracy remains robust under non-minimal or perturbed training data. These results expose critical blind spots in CoT monitoring and challenge prevailing assumptions within AI safety that rely on chain-of-thought reasoning for model interpretability and alignment.