Score
Designs, builds, or analyzes algorithms, models, or processes that draw conclusions from premises and evidence, chain multi-step inferences, apply formal or informal logical rules, and generate or evaluate justified arguments or plans under uncertainty.
This study addresses a critical limitation in traditional reproducible research, where sharing only code and results fails to expose the implicit assumptions, expectations, and premises underlying an analyst’s reasoning—thereby hindering thorough evaluation of analytical quality. To overcome this, the paper proposes a formal modeling framework that explicitly translates the analyst’s tacit reasoning process into structured logical representations, statically capturing the construction logic of the analysis. This approach enables systematic scrutiny of the analytical chain of reasoning, assumption sensitivity, and conclusion robustness—even in the absence of the original data. Empirical validation on representative data analysis tasks demonstrates the framework’s effectiveness, achieving both logical visualization and data-free static assessment of analytical integrity.
Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.
Existing legal AI systems predominantly target isolated subtasks, failing to support end-to-end, high-stakes decision-making in real-world scenarios. To address this gap, we introduce LawFlow—the first dynamic, modular, and iterative legal workflow dataset—grounded in law students’ authentic reasoning during corporate formation tasks. LawFlow uniquely captures three hallmarks of legal practice: ambiguity, iterative revision, and client-adaptive reasoning. Through comparative human-AI process tracing, we identify systematic deficiencies in large language models regarding closed-loop decision-making and execution flexibility. We propose a collaborative AI paradigm centered on hybrid planning, adaptive execution, and decision-point support. Empirical findings show practitioners prefer AI in auxiliary roles—e.g., ideation, blind-spot identification, and solution generation—over autonomous decision-making. All data and code are open-sourced to advance explainable, human-in-the-loop legal AI. (149 words)
This paper addresses the lack of a unified metatheoretic characterization for program logics handling multi-branching effects—such as nondeterminism and probabilism. We propose a novel program logic framework centered on algebraic choice structures. Methodologically, we are the first to embed algebraic effects modeling directly into the core of Hoare logic, integrating modal semantics with a relatively complete proof system that supports general loops and uniform reasoning across effect types (e.g., nondeterministic and probabilistic). Our main contributions are: (1) the first relatively complete proof system for Hoare logic strictly extending it to cover multiple branching effects; (2) a unified metatheoretic account of multi-result programs; and (3) formal support for cross-model reuse of proof fragments—enabling verification transfer between distinct semantic models (e.g., relational, probabilistic, or game-based interpretations).
Existing datasets of scientific ideation trajectories struggle to comprehensively capture the full research process—from literature exploration and tool utilization to the evolution of intermediate artifacts and final proposals. This work proposes a reverse-to-forward synthesis mechanism that emulates the uncertainty, evidence integration, and phased convergence characteristic of real scientific inquiry through a Generator–Advisor architecture. By leveraging action–observation–editing sequence modeling, context-aware verification, and process-level supervision, the approach generates multi-turn trajectories aligned with authentic research practices, starting from high-quality papers and proposals. The study yields the first trajectory dataset spanning the complete scientific workflow and establishes a generalizable paradigm for synthesizing process-supervised data for scientific agents.
This study demonstrates that AI agents exhibit systematic decision biases when processing identical evidence, depending on how a problem is framed—particularly in high-stakes domains such as medicine, election forensics, and geopolitics. Through controlled experiments that hold evidence constant while varying contextual framing, and by integrating Bayesian inference modeling with cross-domain agent behavior analysis, the research reveals for the first time that AI systems display human-like motivated reasoning: when task framing aligns with their prior beliefs, agents are significantly more likely to endorse corresponding conclusions and adapt their search strategies, analytical norms, and evidence evaluation criteria accordingly. These findings underscore the profound influence of prior beliefs on AI reasoning and offer critical insights into the reliability and transparency of AI systems in high-risk applications.
This work addresses the opacity of existing large language model (LLM)-driven simulation-based decision systems, which treat scientific simulators as black boxes and lack explicit reasoning about their underlying mechanisms and assumptions. To overcome this limitation, the authors propose MechSim, a novel framework that introduces mechanism-level reasoning into the interaction between LLMs and scientific simulators. MechSim employs structured mechanistic representations to model a simulator’s assumptions, variable dependencies, and execution traces, integrating neural-symbolic reasoning with a constraint engine to enable LLMs to perform explainable, traceable, and constraint-aware inference. Experiments across multiple high-stakes domains demonstrate that MechSim significantly enhances the quality of mechanistic explanations, deepens simulation analysis, and improves the reliability of downstream decisions, thereby transcending the traditional limitation of neural-symbolic systems that operate only on static symbolic representations.
This work addresses the challenge of deploying generative AI in high-stakes decision-making, where hallucinated reasoning, unsupported claims, and weak traceability often preclude compliance with certification-grade accountability requirements. To bridge this gap, the authors propose a “compliance-by-construction” architecture that uniquely integrates typed argumentation graphs, retrieval-augmented generation (RAG), formal verification kernels, and W3C PROV-based provenance tracking. This framework ensures that every AI-generated claim is grounded in authoritative evidence and subjected to rigorous inference constraints before being admitted into official decision records. Empirical evaluation demonstrates that the architecture effectively blocks unsubstantiated assertions from entering the decision pipeline and substantially improves the efficiency of constructing compliant, auditable arguments, thereby enabling controlled, verifiable, and accountable use of generative AI in high-assurance settings.