Score
Designs, builds, and evaluates systems that use large language models to support human decision-making by ingesting and synthesizing heterogeneous natural-language inputs. Produces operator-facing action recommendations, concise incident or event summaries, and interpretable, evidence-backed rationales that justify and explain suggested decisions.
Existing research on large language models (LLMs) as autonomous agents and tool users remains fragmented and limited in architecture design, multi-agent coordination, tool integration, cognitive mechanism modeling, and evaluation frameworks. Method: This survey systematically analyzes 2023–2025 top-tier conference and journal publications using structured literature analysis, integrating prompt engineering and fine-tuning techniques to dissect LLM implementations of core cognitive capabilities—reasoning, planning, and memory. Contribution/Results: We identify three breakthrough directions—verifiable reasoning, self-improvement, and personalized customization—and distill ten concrete future research pathways. Further, we propose a unified evaluation framework covering 68 publicly available datasets, exposing critical gaps in current benchmarks regarding task generalization, dynamic adaptability, and causal attribution capability.
To address the imbalance between predictive accuracy and mechanistic interpretability in cognitive modeling, this paper proposes a reinforcement learning (RL)-guided large language model (LLM) framework that jointly achieves high-accuracy prediction and cognitively interpretable attribution of human risk-sensitive choices. Innovatively, we employ outcome-oriented Proximal Policy Optimization (PPO) to train LLaMA/GPT-series models to generate natural-language reasoning traces, enabling the first explicit computational modeling of underlying cognitive mechanisms in risky decision-making. Evaluated on multiple Risky Choice Tasks benchmarks, our method achieves state-of-the-art predictive performance (average AUC > 0.89). Moreover, expert evaluations confirm that its generated explanations significantly outperform baselines in interpretability—yielding a 37% improvement in explainability scores. This work effectively bridges the long-standing gap between behavioral prediction and cognitive mechanism explanation in computational cognitive science.
Group decision-making scenarios—such as conference scheduling—face challenges including heterogeneous preferences, asymmetric power dynamics, and inefficient interpersonal interaction. Method: This paper proposes a large language model (LLM)-based collaborative decision-making framework. It introduces a novel large-scale synthetic dialogue simulation paradigm for evaluation and implements a two-stage mechanism: (1) heterogeneous preference modeling via dialogue state tracking and employee profiling; and (2) multi-objective iterative optimization balancing fairness and satisfaction. The framework integrates LLM-based reasoning, preference-balanced aggregation, and dynamic solution generation. Contribution/Results: Experiments demonstrate that the framework significantly reduces required interaction rounds while achieving over 92% accuracy in preference aggregation and over 92% plausibility in logical reasoning—both in synthetic simulations and human-subject studies. It establishes a verifiable, LLM-augmented paradigm for collective intelligence in group decision-making.
This work identifies a human-like “descriptive–normative” dual mechanism in large language model (LLM) autonomous decision sampling: descriptive components reflect statistical regularities, while normative components encode implicit ideal prototypes. Through behavioral analysis, conceptual prototype modeling, cross-domain case studies (e.g., public health, economic forecasting), and comparative validation against human cognitive experiments, we demonstrate that LLM-generated samples systematically deviate from statistical means and converge toward latent ideal values—inducing consistent, significant decision biases. This is the first study to both propose and empirically validate this dual-component structure in LLM sampling. We introduce a novel theoretical framework—“normative normality shaping conceptual prototypes”—that explicates the intrinsic mechanism by which such biases propagate systemic ethical risks. Our findings provide critical theoretical warnings and actionable intervention levers for developing trustworthy AI decision-making systems. (149 words)
To address the high computational overhead, suboptimal prompt engineering, and prohibitive fine-tuning costs of large language models (LLMs) in sequential decision-making tasks, this paper proposes the first lightweight LLM-based decision framework grounded in online model selection. Our method requires no gradient-based updates; instead, it dynamically orchestrates multiple lightweight LLM agents, selecting the optimal one per decision step via real-time scheduling, while integrating efficient prompting interfaces and sequence-aware modeling. Crucially, we formulate model selection as an online decision problem that jointly optimizes statistical performance and inference efficiency. Evaluated on a large-scale Amazon dataset, our approach achieves over 6× improvement in decision quality compared to strong baselines, with only a 1.5% LLM invocation rate—significantly outperforming both conventional decision algorithms and full-scale LLM agent ensembles.
This work addresses the misalignment between large language models’ (LLMs) reasoning logic and human cognition in legal domains. We propose the first fine-grained evaluation framework explicitly designed for cognitive alignment—moving beyond mere output correctness to interrogate internal reasoning processes. Grounded in interactive attribution, the framework explicitly models LLMs’ raw decision logic as quantifiable objects and establishes multi-dimensional logical consistency metrics grounded in mathematical fidelity theory. Empirical evaluation on legal reasoning tasks reveals that over 60% of correctly answered instances exhibit internal reasoning that significantly violates domain-specific human knowledge and established legal inference patterns. This study is the first to systematically expose the pervasive “superficially correct but logically flawed” behavior of LLMs, empirically validating the necessity of logic-level assessment. Our framework introduces a novel paradigm for enhancing model trustworthiness and enabling effective human-AI collaboration in high-stakes legal applications.
Existing intrinsic NLG evaluation metrics—such as n-gram overlap and sentence fluency—exhibit weak correlation with real-world decision outcomes in high-stakes domains. Method: This paper introduces the first decision-oriented text evaluation framework, specifically designed for market briefing generation in finance. It pioneers the use of human investors’ and autonomous LLM agents’ trading performance as primary evaluation criteria, moving beyond superficial textual features. Experiments employ morning summaries and closing commentary, quantifying how generated texts influence human–AI collaborative decision-making in live trading tasks. Contribution/Results: When relying solely on summary texts, neither humans nor LLMs significantly outperform random baselines. However, texts exhibiting analytical depth substantially improve joint human–LLM decision accuracy and financial returns. These findings empirically validate the effectiveness and practical utility of decision-utility–driven evaluation.
This work addresses the prevailing limitation in large language model (LLM) development, wherein human values are typically incorporated only post-training, lacking systematic integration across the model’s entire lifecycle. To bridge this gap, the paper introduces the Human-Centric Large Language Model (HCLLM) framework, which for the first time deeply integrates natural language processing, human-computer interaction, and responsible AI methodologies throughout all stages—from system design and data collection to training, evaluation, and deployment. The framework harmonizes ethical, economic, and technical objectives, offering developers actionable, principle-based guidance. Its forward-looking applicability and practical utility are demonstrated through a case study situated in future workplace scenarios, thereby advancing LLM development toward a genuinely human-centered paradigm.
Large language models (LLMs) suffer from opaque and non-auditable decision-making processes, hindering trust and regulatory compliance. Method: This paper proposes a hierarchical, explainable AI architecture that integrates LLMs with structured decision frameworks—including QOC (Question-Options-Criteria), sensitivity analysis, game-theoretic modeling, and risk management—thereby decoupling reasoning and explanation spaces. Unlike post-hoc interpretability methods, it enables prospective modeling of inference paths through standardized analytical workflows. Contribution/Results: It is the first work to systematically co-model classical decision science paradigms with LLMs, supporting end-to-end traceability and formal verification of decision logic. Experiments demonstrate that the system replicates expert-level reasoning in complex domains—including decentralized governance, systems analysis, and strategic planning—while significantly enhancing transparency, auditability, and trustworthiness of AI-driven decisions.
Large language models (LLMs) as cybersecurity decision aids may exert a double-edged effect on human judgment—enhancing accuracy while undermining independent reasoning, amplifying automation bias, and fostering decision homogenization. Method: We conducted a focus group experiment comparing user performance in security tasks with and without LLM support, measuring accuracy, behavioral resilience, and dependency dynamics; participants were stratified by cognitive resilience level. Contribution/Results: LLMs significantly improved accuracy and consistency on routine tasks but suppressed cognitive diversity. High-resilience users actively calibrated LLM outputs and mitigated bias, whereas low-resilience users exhibited heightened overreliance. This study provides the first empirical evidence that cognitive resilience is a critical individual moderator of LLM–human collaborative efficacy in cybersecurity. It underscores the necessity of resilience-centered design for human–AI collaboration frameworks in safety-critical domains.
This study addresses the fundamental question of whether large language models (LLMs) exhibit human-like rationality. To this end, it introduces the first comprehensive evaluation benchmark covering both theoretical rationality (e.g., logical consistency) and practical rationality (e.g., preference coherence). Methodologically, it integrates principles from cognitive science and behavioral economics to design multi-domain, context-sensitive assessment tasks, and develops an open-source, extensible automated evaluation toolkit. The contributions are threefold: (1) a systematic, theory-grounded framework for assessing LLM rationality; (2) empirical evaluations across mainstream LLMs, revealing critical boundaries and cross-model disparities in rational behavior; and (3) a reproducible benchmark and analytical foundation to guide model refinement, trustworthy AI development, and rationality alignment research. (136 words)