A 3D Characterization Framework for Intelligent Sequential Decision Making

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a unified evaluation benchmark for cross-paradigm sequential decision-making methods, including graph-based approaches, reinforcement learning (RL), and large language models (LLMs). We propose a three-dimensional feature framework integrating Markov Decision Process (MDP) formal projection, autonomy ranking, and skill cost. Using this framework, we conduct quantitative comparisons of algorithms such as Neurosolver, FBRL, and AutoToS on the Tower of Hanoi task. This work bridges the gap in fair cross-paradigm evaluation, revealing that LLMs exhibit exponentially increased reasoning verification complexity due to weak action-space constraints, resulting in significantly higher resource overhead than traditional RL and graph-based methods. These findings provide a theoretical foundation for algorithm selection in sequential decision-making.
📝 Abstract
Puzzles are widely used to evaluate the reasoning capabilities of artificial intelligence (AI) systems for sequential decision making, yet approaches originating from different paradigms are rarely compared under unified conditions. To address this gap, we introduce a three-dimensional characterization framework that enables the analysts of AI methods by 1) projecting them to the Markov decision process (MDP) sequential decision making formalism, 2) degree of autonomy through human prior ranking of their designs and, 3) skill and computational cost. Using this framework, we analyze how representative graph-based, reinforcement learning, and large language model (LLM)-based approaches differ in their design choices and performance characteristics, instantiated respectively by Neurosolver, forward-backward reinforcement learning (FBRL), and automated thought-of-search (AutoToS), including a double-agent extension of thought-of-search (DA-ToS). The analysis relies on the Tower of Hanoi puzzle that provides a controlled benchmark with well-defined rules and scalable complexity, enabling consistent comparison across increasing problem sizes. The 3D characterization reveals that LLM-based methods, due to their weakly constrained action-space design, shift complexity from architecture to inference-time verification, leading to substantially higher memory and runtime costs than Neurosolver and FBRL.
Problem

Research questions and friction points this paper is trying to address.

Sequential Decision Making
Characterization Framework
Markov Decision Process
Large Language Models
Benchmark Comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D Characterization Framework
Sequential Decision Making
Markov Decision Process
Large Language Models
Tower of Hanoi
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sadig Gojayev
Department of Communication Systems, Jožef Stefan Institute, SI-1000 Ljubljana, Slovenia
Carolina Fortuna
Carolina Fortuna
Jozef Stefan Institute
artificial intelligencecyber-physical systems