🤖 AI Summary
This study addresses the limitation of existing reasoning systems that focus solely on premise relevance while neglecting policy authorization constraints. To this end, this work proposes a controlled deduction framework, constructs an RBAC-enhanced Spider benchmark, defines transfer local admission predicates, and introduces matching authorization pair construction, linear controllers, and context-local role permutation methods. Experimental results demonstrate that linear models fail to recover policy relations, with accuracy degrading to 50%, and exhibit limitations such as role-name shortcuts and frozen representations, whereas symbolic oracles maintain 100% accuracy. This research reveals the inherent deficiencies of linear models in permission reasoning and underscores the necessity of conducting leakage auditing for reasoning systems.
📝 Abstract
Reasoning systems usually treat premise use as a question of relevance: if a fact is available and useful, it may be selected for inference. Authorization imposes a different constraint: a premise may be represented and logically usable but not permitted for a particular local transition. We formalize this distinction as Governed Deduction (GD), with a transition-local admission predicate admit(p, tau, S). From an independently produced RBAC-augmented Spider benchmark, we construct 4,461 matched authorization pairs in which the same query premise and policy state support permitted and denied consuming transitions. An initial joint controller reaches 99.19% held-out accuracy, but a transition-only control reaches 100%, exposing a role-name shortcut. After a frozen, label-independent context-local role permutation removes that shortcut, premise/state-only, transition-only, and joint linear controllers all score exactly 50% on 1,856 held-out edges, while a symbolic policy oracle remains at 100%. The result is a controlled negative finding: the benchmark instantiates policy-grounded authorization beyond relevance, but the frozen linear representation does not recover the relation. Matched one-sided controls and leakage audits are therefore essential for evaluating learned policy-sensitive reasoning.