🤖 AI Summary
This study addresses the limitations of policy evaluation in capturing institutional dynamics and predicting second-order effects by proposing a Second-Order Effect State Transition Framework. We construct a source-link benchmark comprising 96 cases alongside a side-effect simulator, and introduce a transition channel auditing protocol to systematically validate model sensitivity to downstream institutional impacts. Experimental results demonstrate that the proposed simulator achieves a policy effect quality score of 0.945, significantly outperforming baseline methods. This research effectively enhances model sensitivity to complex institutional environments, providing novel methodological support and evaluation benchmarks for dynamic policy assessment.
📝 Abstract
Policy evaluation often estimates direct benefits and costs while treating the institutional environment as fixed. In practice, a policy changes the system it enters: actors adapt, enforcement capacity shifts, burdens move, and new equilibria form around capture, gaming, compliance theater, irreversibility, and repair costs. We formalize this as second-order policy-effect prediction and present a source-linked benchmark for policy simulation. The benchmark contains 96 named public-policy cases across eight domains and four balanced action classes: implement, modify, pilot, and block. Each case includes source locators and state variables for benefit, capture, gaming, burden shift, instability, uncertainty, irreversibility, distributional risk, and implementation capacity. The runner regenerates method outputs and aggregate results from the case table, and the simulator never reads the expert action target. We report a protocol-based transition-channel audit with recall, precision, F1-style efficiency, and selective top-channel stress diagnostics, so universal channel coverage is not mistaken for field validation. The side-effect simulator achieves mean policy-effect quality of 0.945, compared with 0.838 for the risk-register baseline and 0.879 for the causal-loop baseline. Its advantage is concentrated in side-effect recall and aggregate transition scoring; it does not dominate the best structured baselines on exact policy-action choice. The evidence remains benchmark-based, but supports a bounded claim: transition-state variables make policy simulators more sensitive to downstream institutional effects.