runtime trace gating

Designs, builds, or analyzes systems that monitor and intercept runtime execution traces and tool calls to enforce policies by checking trace prefixes and applying online gates that block, warn, or otherwise alter execution when constraints are violated. Implements interception/harness layers and prefix-based execution checks to provide real-time constraint gating, procedural-compliance warnings, and measurable enforcement behavior.

runtimetracegating

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the semantic gap faced by current AI agents in enforcing natural language policies: the intended policy semantics are difficult to enforce precisely and interpretably at the system level. To bridge this gap, the authors propose a novel approach that integrates agent-side context with kernel-level enforcement mechanisms. For the first time, policy context is preserved on the agent side, while a domain-specific language (DSL) for information flow control (IFC), implemented via eBPF, enables comprehensive, action-level policy enforcement within the operating system kernel. This framework supports cross-event data-flow and ordering constraints, significantly improving policy compliance rates by covering indirect execution paths invisible to conventional tool-call interception. The system incurs only 1.9%–8.4% runtime overhead and provides semantically clear feedback instead of ambiguous errors.

AI agentsinformation-flow controlpolicy enforcement

Existing automated tool-calling systems often suffer from insufficient generalization due to model-centric designs and heavy reliance on prompting, leading to recurrent failures such as unsafe side effects, invalid parameters, uncontrolled retries, and sensitive data leakage. This work proposes a model-agnostic, policy-first framework for tool orchestration that enforces permission control prior to invocation, enhancing safety through explicit constraints, risk-aware gating, recovery mechanisms, and auditable explanations. Key contributions include a policy-first paradigm for tool workflows, a lightweight domain-specific language (DSL) for policies, a runtime execution engine, and a reproducible safety benchmark based on trajectory replay. In 225 controlled experiments, the strictest policy configuration achieved a violation prevention rate of 0.681, reduced retry amplification to 1.378, and attained a sensitive information leakage recall of 0.875, effectively quantifying the trade-off between safety and utility.

model-agnostic safetysensitive data leakagetool-using automation

This work addresses the lack of runtime-enforceable mechanisms for natural language behavioral policies in AI agents when invoking third-party skills, which hinders fine-grained monitoring and intervention against policy violations across actions. The authors propose VIGIL, an end-to-end runtime enforcement framework that translates skill specifications, operator constraints, and global rules into executable policies through a context-aware policy language and symbolic evaluation. VIGIL dynamically detects violations along agent execution traces by enabling fine-grained reasoning over event orderings, parameter relationships, and cross-call data flows, thereby overcoming limitations of fixed event models. Leveraging SMT-based symbolic execution and bounded trajectory constraints, VIGIL achieves efficient runtime monitoring and intervention. Empirical evaluation demonstrates that the approach attains over 95% violation detection accuracy with a false positive rate below 10% in real-world LLM agent tasks.

agentic systemsbehavioral specificationscontextual granularity

Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.

autonomous enterprise agentscompliance-aware behaviorgeneralist agents

This study addresses the challenge of providing certifiable runtime safety guarantees prior to tool invocation, focusing on three core issues: the representability of policy states, the observability of monitoring evidence, and the impact of interventions on future behavior. To this end, we propose the first formal theoretical framework for runtime safety-executable boundaries, distinguishing among static policy executability, statistical calibration under exogenous legal constraints, and closed-loop intervention effects. Building upon finitely controlled models, we develop a method for closed-loop safety certification that integrates register model identification, Neyman–Pearson hypothesis testing, conformal calibration, and occupancy planning. Empirical validation through static diagnosis, model enumeration, representation rewriting, and closed-loop re-execution experiments demonstrates the efficacy of our approach and exposes the fundamental limitations of static calibration under representation attacks.

certified safetyenforceable policiesguardrails

Latest Papers

What's happening recently
View more

This work addresses the security threat posed by external tools that, while appearing benign at their interfaces, may embed malicious behaviors when invoked by large language model agents. To mitigate this risk, the authors propose ToolGuardian, a novel framework featuring a declarative policy layer based on Answer Set Programming. This layer explicitly models tool capabilities, effects, task context, and compositional relationships to enable auditable, deterministic pre-execution screening and runtime task-aware authorization. By integrating progressive tool representations—including natural-language descriptions, system call traces, simulated executions, and source code analysis—with logical reasoning, ToolGuardian achieves an F1 score of 0.86 and 88% accuracy in admission control across 16 MCP tools (including eight malicious variants) and 20 diverse scenarios, and attains 100% classification accuracy during runtime authorization.

AI agent securitydeclarative policymalicious tools

This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.

deterministic workflowsLLM agent tracestool-use

This work addresses the long-standing lack of systematic validation for processor specifications, which can lead to distorted program behavior and security vulnerabilities. It presents the first automated differential testing framework tailored for open-source SLEIGH specifications, automatically generating decodable instructions and initial execution states by parsing specification structures, and systematically validating them against multiple hardware reference implementations across architectures. Applied to x86-64 and AArch64, the approach uncovered 38,920 semantic discrepancies, identified 125 unique defects—many of which were subsequently fixed—and significantly improved specification fidelity. Furthermore, it exposed inconsistencies across vendor implementations and led to eight concrete recommendations, establishing a new paradigm for ensuring the reliability of instruction set architecture specifications.

disassembleremulatorprocessor specification

Hot Scholars

SF

Stefano Forti

Department of Computer Science, University of Pisa
cloud-edge continuumdistributed systemsgreen computingautomated reasoning
LL

Liang Luo

University of Washington
Systems for Machine LearningComputer SystemsComputer ArchitectureMachine Learning for Systems
JG

Jinyu Gu

Shanghai Jiao Tong University
Operating SystemSystem SecurityVirtualization