Score
Designs, builds, and evaluates decision systems and frameworks that select actions by retrieving and applying explicit decision rules conditioned on current context and history; this includes creating decision models, online decision-rule mechanisms, rule-augmented action-selection policies, and interpretable decision-support components that enable rule-aware corrections under sparse feedback.
Existing decision support systems treat analytical frameworks (e.g., 6C) and heuristic strategies (e.g., Thirty-Six Stratagems) as disjoint entities, lacking semantic-level integration. Method: We propose a semantics-driven strategy recommendation system that (i) establishes the first semantic alignment between analytical frameworks and heuristics via a cross-paradigm semantic mapping mechanism; (ii) designs a multimodal language representation to uniformly encode heterogeneous strategic knowledge—including text, matrices, and diagrams; and (iii) adopts an LLM-constrained computational architecture to ensure interpretable and controllable reasoning. Our approach integrates deep semantic NLP, vector-space modeling, cross-framework similarity computation, and lightweight collaborative inference. Contribution/Results: Experiments on multiple corporate strategy cases demonstrate its effectiveness. The system supports plug-and-play recommendation for arbitrary framework–heuristic combinations and generates strategy proposals that balance theoretical rigor with practical feasibility.
This work addresses a critical limitation in existing large language model systems, where control decisions—such as whether to answer, seek clarification, or invoke external tools—are tightly coupled with text generation, hindering failure diagnosis. To resolve this, the authors propose a decision-centric framework that explicitly models control decisions as a standalone module, decoupling evaluation from action through a signal–policy separation architecture. This design enables precise failure attribution and modular system improvement. The framework unifies handling of both single-step and sequential decisions and integrates mechanisms such as routing and adaptive reasoning to achieve interpretable and intervenable control. Experimental results demonstrate that the approach significantly reduces ineffective actions and improves task success rates across three benchmarks, while also revealing distinct and diagnosable failure patterns.
To address the “garbage-in, garbage-out” problem in reinforcement learning (RL) that degrades decision-making quality, this paper proposes an enhanced RL framework integrated with external supervision. Methodologically, it introduces a dual-external-agent architecture: one agent delivers real-time human feedback and action correction, while the other performs dynamic data filtering and constructs high-quality trajectory loops. The framework unifies RL, reinforcement learning from human feedback (RLHF), online evaluation, and data cleansing into a human–machine collaborative hybrid intelligence training paradigm. Its key innovation lies in the first integration of dual-agent supervision into the RL decision pipeline, enabling simultaneous dynamic policy correction and training data quality enhancement. Experiments on banking document recognition and information extraction demonstrate a 12.7% accuracy improvement over baselines, alongside significantly enhanced robustness and cross-scenario adaptability.
This study addresses the challenge of designing an optimal recommendation mechanism in a finite-horizon discrete-time dynamic system where a system designer cannot directly control the actions of two strategic agents. The designer aims to maximize their own objective by recommending actions based on shared historical information, while ensuring that the agents find it sequentially rational to follow these recommendations, thereby forming a sequential rationality equilibrium. To this end, the paper proposes a novel recommendation mechanism that explicitly satisfies sequential rationality constraints and develops a computationally tractable solution framework combining backward induction with linear programming. This approach achieves, for the first time, the optimization of the designer’s objective under strict sequential rationality conditions, demonstrating both the effectiveness and computational feasibility of the proposed mechanism.
To address suboptimal decision-making in real-world project selection arising from human cognitive limitations, this paper proposes a teachable strategy discovery method explicitly designed for human cognitive constraints and develops an interactive intelligent tutoring system to enhance practical decision-making competence. Methodologically, it pioneers the application of automated strategy discovery to authentic project selection tasks, introducing the MGPS (Model-Guided Policy Search) algorithm and an interpretable, pedagogically grounded strategy generation framework that integrates cognitive-model-informed policy optimization with computationally rigorous benchmark evaluation. Results demonstrate that MGPS consistently outperforms state-of-the-art methods in both solution quality and computational efficiency. Moreover, human participants trained via the intelligent tutor exhibit statistically significant improvements in strategy quality, empirically validating the method’s effectiveness and practical utility in improving human decision-making under naturalistic conditions.
This work addresses the limitations of human-selected interpretable concepts in reinforcement learning—namely their reliance on domain expertise, high cost, and lack of performance guarantees. The authors formulate concept selection as a state abstraction problem and introduce a “decision relevance” criterion: retaining only those concepts that distinguish states requiring different actions, thereby ensuring that states sharing the same concept also share an optimal action. Building on this principle, they propose DRS, the first algorithm for automatic concept selection in sequential decision-making, which integrates state abstraction theory to define a decision relevance metric and jointly optimizes candidate concept filtering with policy learning. Theoretical analysis provides error bounds linking selected concepts to policy performance, and experiments demonstrate that DRS automatically recovers or even surpasses handcrafted concept sets across multiple RL benchmarks and a real-world clinical environment, significantly improving the efficacy of concept-based interventions at test time.
This study addresses the multifaceted challenges confronting decision-makers in high-stakes environments—namely, uncertainty, resource constraints, time pressure, and accountability risks. To navigate these complexities, the paper proposes an agent-based metadata governance mechanism that synergistically integrates machine intelligence with human cognition. By dynamically managing contextual metadata, this approach enhances situational awareness, establishes an adaptive decision-making framework, and balances risk tolerance with conditional accountability. Moving beyond conventional decision-support paradigms, the proposed method significantly improves contextual understanding, decision coherence, and adaptability in complex, time-critical scenarios, thereby offering a practical pathway toward responsible and effective decision-making in high-consequence settings.
This work addresses a critical limitation in current automated decision-making systems, which prioritize predictive accuracy while neglecting their systemic impact on organizational workflows and decision processes, thereby compromising downstream societal outcomes. To overcome this “prediction-first” paradigm, the study proposes an intervention-oriented integrative framework that systematically incorporates organizational dynamics and intervention mechanisms into the design and evaluation of automated decision systems. By synthesizing insights from social science theory and decision systems analysis, the research establishes an interdisciplinary pathway tailored to real-world deployment contexts. This approach offers a novel paradigm for understanding and optimizing the societal consequences of automated decision-making in high-stakes domains such as criminal justice, medical triage, and educational support.