🤖 AI Summary
This study addresses the opacity and non-auditable nature of large language model (LLM) decision-making. Methodologically, it introduces a modular, interpretable LLM agent framework that integrates deterministic analyzers—including Vester’s sensitivity analysis, normal-form and sequential game modeling, matrix classification, and backward induction—with an LLM (default: GPT-5) to jointly generate explicit, traceable intermediate reasoning artifacts. The framework supports dynamic switching among analytical paradigms and role-conditioned agency. Its key contribution is a dual-track “LLM + deterministic analyzer” architecture that preserves reasoning flexibility while enabling end-to-end auditability. Evaluated on a real-world logistics decision-making case, the approach achieves a factor alignment rate of 55.5% (full dataset) and 62.9% (core subset), a role-matching accuracy of 57%, and LLM-generated assessments statistically comparable to human expert baselines.
📝 Abstract
We present a modular, explainable LLM-agent pipeline for decision support that externalizes reasoning into auditable artifacts. The system instantiates three frameworks: Vester's Sensitivity Model (factor set, signed impact matrix, systemic roles, feedback loops); normal-form games (strategies, payoff matrix, equilibria); and sequential games (role-conditioned agents, tree construction, backward induction), with swappable modules at every step. LLM components (default: GPT-5) are paired with deterministic analyzers for equilibria and matrix-based role classification, yielding traceable intermediates rather than opaque outputs. In a real-world logistics case (100 runs), mean factor alignment with a human baseline was 55.5% over 26 factors and 62.9% on the transport-core subset; role agreement over matches was 57%. An LLM judge using an eight-criterion rubric (max 100) scored runs on par with a reconstructed human baseline. Configurable LLM pipelines can thus mimic expert workflows with transparent, inspectable steps.