🤖 AI Summary
This work addresses the absence of structured, auditable, and jurisdiction-aware behavioral governance mechanisms in large language models during inference. We propose the Dynamic Behavioral Constraints (DBC) benchmark, introducing a taxonomy-based hierarchical governance framework that deploys 150 model-agnostic controls at the system prompt layer across 30 risk domains. The framework integrates six risk clusters and five adversarial attack strategies—including role-playing and authority impersonation—to establish a causally attributable, three-tier comparative evaluation protocol. Experimental results demonstrate that DBC reduces overall risk exposure from 7.19% to 4.55% (a 36.8% relative reduction), substantially outperforming standard safety prompts. The model achieves an MDBC compliance score of 8.7/10 and an EU AI Act alignment score of 8.5/10. All code and evaluation artifacts are open-sourced to enable automated compliance assessment and longitudinal tracking.
📝 Abstract
We introduce the Dynamic Behavioral Constraint (DBC) benchmark, the first empirical framework for evaluating the efficacy of a structured, 150-control behavioral governance layer, the MDBC (Madan DBC) system, applied at inference time to large language models (LLMs). Unlike training time alignment methods (RLHF, DPO) or post-hoc content moderation APIs, DBCs constitute a system prompt level governance layer that is model-agnostic, jurisdiction-mappable, and auditable. We evaluate the DBC Framework across a 30 domain risk taxonomy organized into six clusters (Hallucination and Calibration, Bias and Fairness, Malicious Use, Privacy and Data Protection, Robustness and Reliability, and Misalignment Agency) using an agentic red-team protocol with five adversarial attack strategies (Direct, Roleplay, Few-Shot, Hypothetical, Authority Spoof) across 3 model families. Our three-arm controlled design (Base, Base plus Moderation, Base plus DBC) enables causal attribution of risk reduction. Key findings: the DBC layer reduces the aggregate Risk Exposure Rate (RER) from 7.19 percent (Base) to 4.55 percent (Base plus DBC), representing a 36.8 percent relative risk reduction, compared with 0.6 percent for a standard safety moderation prompt. MDBC Adherence Scores improve from 8.6 by 10 (Base) to 8.7 by 10 (Base plus DBC). EU AI Act compliance (automated scoring) reaches 8.5by 10 under the DBC layer. A three judge evaluation ensemble yields Fleiss kappa greater than 0.70 (substantial agreement), validating our automated pipeline. Cluster ablation identifies the Integrity Protection cluster (MDBC 081 099) as delivering the highest per domain risk reduction, while graybox adversarial attacks achieve a DBC Bypass Rate of 4.83 percent . We release the benchmark code, prompt database, and all evaluation artefacts to enable reproducibility and longitudinal tracking as models evolve.