🤖 AI Summary
This study addresses the disconnect between LLM compliance system judgments and regulatory rules, demonstrating that high accuracy does not necessarily indicate genuine rule adherence. To investigate this, the authors introduce perturbation techniques—including deletion, swapping, and negation of regulatory rules—combined with internal representation intervention and prompt engineering optimization. Furthermore, two metrics, OCS and ICS-delta, are proposed to directly quantify and audit the model’s reliance on specific rules. The findings reveal a fundamental insensitivity of internal representations to regulatory constraints. Experimental results demonstrate that while general-purpose models achieve high accuracy, they exhibit low rule sensitivity; notably, dedicated guardrail models perform worst, achieving results only marginally better than random chance.
📝 Abstract
Large language model compliance systems are deployed on the assumption that a verdict depends on the regulatory rule it is given. We test this directly across five models and 20 regulatory and platform-policy domains: delete, swap, or negate the governing rule while holding the case fixed, and check whether the verdict changes (OCS) or the model's internal representation of compliance shifts at all (ICS-delta). Neither moves much: models' verdicts are often invariant to substantial perturbations of the supplied rule, and the guard model, evaluated here under a custom-rule adaptation of its native taxonomy, is the least rule-sensitive and least accurate of the five, barely above chance (51%, versus 90-92% for general-purpose models). This reflects easy cases more than blanket neglect: on cases where deleting the rule changes a previously correct model prediction, models do track it closely. Neither better prompting nor direct intervention on the model's internal representations closes this gap. Accuracy alone does not establish that a compliance verdict is grounded in the supplied rule.