controlled text editing

Applying minimal, precise textual edits under governance to correct errors or simplify content while preserving coherence of reasoning chains and providing traceable decision records. This includes designing multi-agent workflows and edit policies that limit scope, document rationale, and leverage external knowledge for safe corrections.

controlledtextediting

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing evaluation methods, which focus narrowly on task completion and fail to ensure trustworthy deployment of embodied agents in multi-step, externally impactful scenarios, while also lacking coordination among evaluation, governance, orchestration, and runtime assurance. To bridge this gap, the paper proposes an integrated four-layer framework that establishes, for the first time, a closed-loop mechanism linking governance obligations to verifiable execution. Guided by the ODTA principles—Observability, Decidability, Timeliness, and Attestability—the framework introduces runtime localization testing and minimal action evidence bundles. Through a human-in-the-loop evidence synthesis approach, it formally connects policy requirements to concrete agent behaviors, exposing critical gaps such as the inability of static permissions and prompts to govern path-dependent actions. Validation via an enterprise procurement agent demonstrates the framework’s capacity to unify safety, robustness, and trajectory-level evaluation.

Agentic AIcompliance verificationevidence synthesis

This work addresses the challenge of effectively governing side effects—such as memory accesses, external calls, and large model queries—in AI workflows without compromising their computational expressiveness. The authors propose an Effect-Transparent Governance framework that interposes a governance operator \( G \) to mediate all effectful operations, enforcing constraints at effect boundaries while preserving internal semantics. Leveraging interactive trees formalized in Rocq 8.19, the framework enables fully verified development without additional axiomatic assumptions. The implementation comprises 36 modules, approximately 12,000 lines of code, and 454 formally verified theorems, establishing seven core properties: orthogonality between governance and computational expressiveness, decidability of governance predicates, superiority of structural over content-based filtering, Turing completeness, semantic transparency, and minimal capability expression.

AI governancecomputational expressivitydecidability boundaries

To address low efficiency, subjectivity, and poor auditability in manual review of enterprise structured documents, this paper proposes a modular, multi-agent collaborative review framework. Leveraging an integrated toolchain—including LangChain, CrewAI, TruLens, and Guidance—the framework implements a configurable, auditable AI agent system that automates assessment across four dimensions: accuracy, consistency, completeness, and clarity, while incorporating continuous feedback to mitigate bias and enhance iterative refinement. Experimental results demonstrate a 99% accuracy rate in consistency evaluation, reduction of per-document review time from 30 to 2.5 minutes, a 50% decrease in errors and bias incidence, and a 95% agreement rate between AI judgments and domain experts. The framework significantly improves objectivity, reproducibility, and engineering deployability of document quality assessment.

AI-driven accuracy, consistency, completeness, clarity assessmentAutomated evaluation of enterprise document quality metricsMulti-agent system for structured document review

This study addresses the lack of a systematic evaluation framework for AI governance prompts, which undermines their structural integrity as enforceable norms. To bridge this gap, the work proposes an integrative five-principle assessment framework grounded in computability theory, proof theory, and Bayesian epistemology. The authors conduct a static analysis of 34 AGENTS.md files from GitHub, revealing that 37% of file–model pairs fail to meet the defined threshold for structural integrity. Common deficiencies include missing data categorization and absent evaluation criteria, exposing undocumented gaps in artifact classification. These findings provide both theoretical grounding and empirical evidence for developing automated tools capable of detecting and repairing such deficiencies, thereby advancing the formalization and operationalization of AI governance prompts.

AGENTS.mdAI governanceexecutable specifications

Latest Papers

What's happening recently
View more

This work addresses the surge in AI agent–generated contributions to open-source projects and the consequent challenges faced by maintainers due to the absence of coordinated governance mechanisms for risk assessment, evidence provision, and review. The paper proposes a project-level governance infrastructure that conceptualizes AI-mediated contributions as governable boundary objects. Central to this framework is the Agent Governance Manifesto (AGM), a bilateral contract linking contributors’ evidence preparation with maintainers’ verification authority. Evaluated through GitHub audits, user studies, and structured validation, the AGM significantly improves risk-label recovery rates (37/38 versus 15/37) and perceived reviewer support (6.14 versus 3.27), while enabling high-fidelity expression of governance state and structural compliance.

agent-mediated contributionsAI governanceboundary objects

Enterprise-scale general-purpose agents lack built-in, reusable governance mechanisms for autonomous cross-tool operation, making it difficult to satisfy requirements for compliance, auditability, and behavioral controllability. This work proposes the CUGA policy system, which embeds runtime governance capabilities into five critical checkpoints of the agent execution pipeline—intent protection, playbook guidance, tool invocation control, human approval gating, and output formatting—through a modular “policy-as-code” architecture. Without requiring model fine-tuning, CUGA enables proactive, continuous, and structured behavior control. By integrating typed governance primitives, dynamic playbook injection, and human-in-the-loop approval, the system effectively blocks malicious requests, enforces structured tool sequences, and triggers manual review for high-risk operations in healthcare scenarios, significantly enhancing policy adherence, execution consistency, and deployment safety.

autonomous enterprise agentscompliance-aware behaviorgeneralist agents

Existing autonomous agents struggle to maintain global consistency in complex document operations and lack a verifiable evaluation framework. This work proposes DocOps—the first verifiable benchmark specifically designed for evaluating agents on intricate document manipulation tasks. By modeling real-world document operations through hierarchical task decomposition and a deterministic verification mechanism, DocOps characterizes complexity along atomic dimensions and progressive workflow structures. Leveraging this framework, the study systematically evaluates multiple state-of-the-art agents and uncovers three critical failure modes in long-range coupled tasks: collapse in long-term state tracking, superficial semantic validation, and destructive editing of structural metadata. These findings delineate clear boundaries in current methods’ capacity to preserve global consistency during complex document operations.

autonomous agentscomplex workflowsdocument operations

This study addresses the challenges of collaboration and quality control in open-source deep learning projects stemming from inadequate governance mechanisms. Drawing on the Institutional Analysis and Development (IAD) framework, it employs a mixed-methods empirical approach combining document content analysis and code commit tracking across PyTorch, TensorFlow, and PaddlePaddle. The analysis encompasses 109 governance documents and over 1,700 code commits, systematically uncovering the structure, temporal evolution, and functional dimensions of governance rules. The research identifies 17 rule themes and 7 rule types, revealing a distinct evolutionary pattern wherein operational rules emerge early and undergo frequent revisions, while structural rules appear later and evolve more steadily. Four core governance functions are distilled, culminating in 33 actionable recommendations for effective open-source AI project governance.

coordinationgovernanceopen source software

This work addresses the critical issue that large language models in financial compliance often produce decisions that appear compliant on the surface yet violate regulations in substance, while existing evaluation frameworks neglect governance constraints on the reasoning behind such decisions. To tackle this, the authors propose a mechanistic governance mechanism that decouples governance from task execution through four external enforcement primitives, ensuring auditability and genuine compliance of the model’s rationales. The study reveals, for the first time, a decoupling between governance quality and task accuracy, demonstrating that high accuracy does not guarantee effective governance. Experiments show that the proposed approach reduces uninformative delayed decisions by 73%, more than doubles the information content of delayed responses, and improves the Matthews Correlation Coefficient (MCC) from 0.43 to 0.88, with causal ablation studies confirming the necessity of each primitive.

decision rationalefinancial decision systemsgovernance-task decoupling

Hot Scholars

TY

Tingting Yu

Associate Professor, University of Connecticut
Software EngineeringSoftware Testing
ZL

Zexin Li

University of California, Riverside
Real-time Embedded SystemsMachine Learning SystemsEfficient Machine Learning
JC

Jinghan Cao

San Francisco State University
Deep LearningLarge Language ModelCloud Software Computating
GL

Guandong Li

hfut
Hyspectral image,Computer vision,AIGC
TH

Ting Hua

University of Notre Dame
Efficient learningCompressionReasoning