Score
Designs and implements systems that generate recommendations subject to explicit constraints and bounded candidate option sets, ensuring outputs conform to heuristics or pre-defined choices. Builds decision-aware generation pipelines that align models to offline reference behaviors and analyzes deployment trade-offs such as latency and compute or monetary cost.
Existing research on generative models for decision-making lacks a unified conceptual framework and systematic taxonomy, hindering comparative analysis and practical deployment. Method: We propose the first comprehensive taxonomy unifying seven major generative model families—Energy-Based Models (EBMs), Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Normalizing Flows, Diffusion Models, Generative Flow Networks (GFlowNets), and Autoregressive Models. We introduce a functional “Controller–Modeler–Optimizer” framework for generative decision-making, instantiated across five real-world domains including autonomous driving and medical diagnosis. We identify and formalize three key evolutionary trajectories: high-performance algorithms, large-scale generalizable models, and self-evolving, adaptive models. Contribution/Results: This work establishes the first complete classification system and standardized evaluation benchmark for generative decision-making. It clarifies methodological strengths, fundamental bottlenecks, and viable pathways for breakthroughs—providing both theoretical foundations and practical guidelines for trustworthy AI-driven decision systems.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
This study addresses how organizations adopting commercial AI decision-support systems often passively accept vendors’ embedded and non-negotiable value judgments, thereby constraining their own decision flexibility. The paper introduces the concept of the “behaviorally feasible set” to formally characterize the range of recommendations an AI system can generate under value-alignment constraints and identifies critical conditions under which organizational needs exceed the system’s adaptive capacity. Through controlled experiments comparing binary decisions and multi-stakeholder preference rankings, the research demonstrates that value alignment substantially shrinks the behaviorally feasible set, diminishing the system’s responsiveness to legitimate contextual variation. Commercial models exhibit heightened rigidity, and the alignment process systematically shifts—rather than neutralizes—stakeholder priorities, revealing that value alignment functions as a structural mechanism embedding vendor values and narrowing organizational negotiation space.
Small and medium-sized enterprises (SMEs) face high deployment costs, poor model reusability, and heavy reliance on expert domain knowledge when adopting combinatorial optimization decision-support systems. Method: This paper proposes the first fully automated, LLM-driven end-to-end paradigm that directly generates executable optimization code from natural language problem descriptions. Our approach integrates multi-stage prompt engineering, domain-knowledge injection, constraint-modeling guidance, and a cross-problem generalization evaluation framework—unifying problem understanding, mathematical modeling, solver integration, and code generation. Results: Evaluated on four canonical combinatorial optimization problem classes, the best-performing generator achieves over 70% syntactic correctness and 45% semantic functional correctness—substantially outperforming baseline methods. The core contribution lies in eliminating manual modeling bottlenecks, thereby enabling SMEs to construct optimization systems with minimal expertise, low entry barriers, and high model reusability.
This work proposes a novel paradigm that integrates discrete-event simulation with a large language model (Gemini-1.5-Pro) to overcome the limitations of traditional simulation-based optimization, which treats simulators as black boxes and offers little insight into policy failure. By leveraging event-level trajectory replay, the method automatically identifies bottlenecks from low-scoring simulation runs and generates interpretable, traceable, code-level policy revisions in parallel. It pioneers the use of simulation trajectories to guide the LLM in targeted heuristic rule modification, combined with rolling evaluation and an elite retention mechanism for iterative policy improvement. Evaluated on dynamic production and AGV scheduling tasks, the approach achieves an average policy score of 77.51 (out of 100), improving the best run from 62.49 to 78.61, and significantly outperforms MILP, handcrafted rules, and metaheuristic baselines across 100 random seeds and fault perturbations.
This work addresses the challenge of maximizing end-to-end success probability in structured agent workflows under hard constraints on budget and deadline. The authors propose Monte Carlo Combinatorial Planning (MCPP), a lightweight closed-loop planner that dynamically replans during execution in response to observations. MCPP employs a finite-horizon stochastic online allocation model with parallel sampling and leverages Monte Carlo simulation to estimate, in real time, the probability of successful task completion under the given constraints. Experimental results demonstrate that MCPP significantly outperforms strong baseline methods on the CodeFlow and ProofFlow benchmarks, consistently achieving higher task completion rates across diverse budget–deadline configurations. These findings validate MCPP’s effectiveness and robustness in resource-constrained scenarios.
This work addresses the limitations of large language models in generating programming competition problems—specifically, their lack of explicit algorithmic planning, difficulty in robustly handling boundary conditions, and inefficient use of execution feedback—by proposing a blackboard-driven Monte Carlo Tree Search (MCTS) framework. The approach formulates program synthesis as a sequential decision-making process that jointly optimizes five stages: strategy selection, code generation, test case generation, quality evaluation, and repair. By integrating a blackboard architecture with MCTS for the first time, the method enables continuous accumulation of structured evidence and coordinated multi-stage decision-making, substantially enhancing the reliability of program generation under constrained settings. Evaluated on four benchmarks—including APPS and CodeContests—the framework achieves state-of-the-art Pass@1 performance, outperforming the strongest baseline, CodeSim, by 26.06 percentage points when using GPT-4o.
This work addresses the semantic gap between tactical Domain-Driven Design (DDD) patterns and general-purpose modeling languages, which often leads to persistent misalignment between design intent and code implementation. To bridge this gap, the authors propose a DDD-native metamodel that treats tactical DDD constructs as first-class modeling primitives and embeds expert architectural knowledge as executable constraints. Integrated with a real-time constraint validation engine and a bidirectional round-trip engineering mechanism, the approach ensures continuous consistency between models and code. By doing so, it substantially lowers the barrier to adopting tactical DDD, transforming it from an expert-dependent, elite practice into a tool-supported, widely reusable engineering methodology.
This work addresses the limited transparency and controllability of large language models (LLMs) in task planning, which often hinder effective incorporation of user intent and real-world constraints. The authors propose an interactive planning framework that enables users to specify constraints in natural language as either hard rules or soft preferences. Hard rules are verified through formal model checking, while soft preferences are evaluated using an LLM-as-judge mechanism. By abstracting constraints into high-level types and applying differentiated validation strategies, the approach significantly enhances the reliability of generated plans and user control over the planning process. User studies demonstrate that the system maintains strong usability while substantially improving user ratings of usefulness, performance, and overall satisfaction.
This work addresses the inefficiencies and cost waste in cloud virtual machines caused by over-provisioning, particularly under shifting workloads that hinder dynamic tuning. The authors propose an engineer-centric, interactive instance tuning system that, for the first time, integrates zero-shot time-series forecasting models—such as Chronos-2—into non-stationary cloud environments. By leveraging large language models to generate offline recommendations, structured decision contexts, and multi-timescale suggestions, the system aligns resource allocation decisions with latency and cost constraints without requiring per-tenant model training, thereby substantially reducing operational overhead. Evaluation on seven production VMs demonstrates a 52.9% average monthly cost reduction (saving \$795) with only a 1.5% resource violation rate, achieving recommendation quality comparable to supervised baselines.