SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of explicit organizational, coordinative, and auditing mechanisms for safety-critical decision-making in small language model deployments. To this end, it proposes SAFESHIELD, a responsibility-oriented framework that decouples and coordinates four core functions—admission review, dynamic routing, evidence retrieval, and publication filtering—while recording auditable decision trajectories. Crucially, this paradigm emphasizes inter-mechanism dependencies over isolated defensive capabilities. Ablation studies demonstrate that removing the coordination mechanism causes a substantial surge in unsafe content leakage, with publication accuracy deteriorating sharply from 96.0% to 69.5%. These findings empirically validate the critical role of organized synergy in ensuring holistic runtime safety.
📝 Abstract
Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capable mechanisms for these functions, but offer limited guidance on how the safety decisions they produce should be explicitly organized, coordinated, and audited. We formulate deployment-time safety as a decision-organization problem with two elements: responsibility-oriented decomposition of safety decisions and explicit coordination among them. We instantiate this formulation in SAFESHIELD, a deployment-time safety system for small language models that organizes four recurring decision responsibilities (admission, routing, evidence, and release) and records committed decisions in auditable Decision Traces. We evaluate SAFESHIELD through mechanism-level experiments, aggregate stage ablations, controlled coordination ablations, and a deployment-oriented stress suite. Mechanism-level results show that the instantiated safeguards provide the capabilities required by the decision process, while aggregate ablations show substantial degradation in end-to-end safety as the surrounding safety organization is removed. More importantly, dedicated coordination ablations preserve the participating safeguard mechanisms while selectively severing their dependencies: removing admission gating substantially increases false release, and withholding upstream evidence from the release decision reduces release accuracy from 96.0% to 69.5%. These results provide system-level evidence that deployment-time safety depends not only on the capability of individual guardrails, but also on how their decisions are organized and coordinated.
Problem

Research questions and friction points this paper is trying to address.

deployment-time safety
small language models
decision organization
runtime guardrails
safety coordination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deployment-time Safety
Decision-Organization Framework
Small Language Models
Decision Traces
Runtime Guardrails
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.