🤖 AI Summary
This work addresses the lack of system-level support in existing large language model (LLM) agents for long-running execution, state persistence, access control, and auditability, as well as the common misconception of treating tool invocations as trust boundaries. Inspired by library operating systems, the paper proposes a novel runtime environment that models LLM agents as AgentProcesses endowed with identity, lifecycle management, and capability-based access control. Authorization and isolation are enforced through runtime primitives rather than tool dispatching. Key innovations include a capability-based security model, a human approval queue, checkpointing mechanisms, and comprehensive audit logging. The system features asynchronous scheduling, namespace-isolated object memory, just-in-time tool registration, and an injectable resource layer. A prototype implementation passes 123 regression tests and real-model evaluations, supporting working directories, shell primitives, and one-time permission grants to ensure secure and controllable long-term operation.
📝 Abstract
Large language model (LLM) agents are evolving from request-response assistants into long-running software actors: they maintain state across model calls, fork subtasks, wait for external events, request human authority, generate tools, and perform side effects that must be resumed and audited. This paper presents Agent libOS, a library-OS-inspired runtime substrate for LLM agents. Agent libOS runs above a conventional host operating system; it does not implement hardware drivers, kernel-mode isolation, or a POSIX-compatible operating system. Instead, it treats an agent as an AgentProcess: a schedulable execution subject with process identity, parent-child lineage, lifecycle state, a tool table derived from an AgentImage, typed Object Memory, explicit capabilities, human queues, checkpoints, events, and audit records. Its central design rule is tools are libc-like wrappers; runtime primitives are the authority boundary. Filesystem access, object access, sleeps, human approval, JIT tool registration, and external side effects are checked at primitive boundaries under explicit capabilities and policy.
We describe the design, threat model, Python prototype, and safety-oriented evaluation. The current prototype implements async scheduling, namespace-local Object Memory, runtime-integrated human approval, one-shot permission grants, per-process working directories, shell and image-registration primitives, Deno/TypeScript JIT tools over a libOS syscall broker, filesystem/object bridge tools, an injectable Resource Provider Substrate, deterministic demos, real-model smoke scripts, and 123 regression tests at the time of writing. Rather than improving planner accuracy, Agent libOS demonstrates a runtime substrate in which long-running LLM agents can be scheduled, authorized, resumed, and audited without treating tool dispatch as the trust boundary.