Score
Designs and implements systems that match algorithms to specific execution or data contexts by encoding algorithm requirements and context features as semantic descriptions, filtering out incompatible options and enforcing data/resource constraints, and ranking or prioritizing candidate algorithms according to predicted suitability or performance.
This paper addresses the challenges of prior knowledge deficiency and fidelity assessment in executing algorithm specifications—expressed in natural language or pseudocode. To this end, it proposes the “Algorithm Specification Domain” framework, which systematically characterizes the requisite knowledge体系, comprising four structured components: syntactic and semantic rules, domain-specific entities, causal constraints, and operational instructions. Methodologically, it introduces the first integration of large language models with reusable documentation techniques to automatically extract and formalize implicit knowledge, yielding a lightweight, persistable, and cross-system-applicable domain model. Key contributions include: (1) a rigorous definition of “execution fidelity” (distinct from correctness), establishing a novel evaluation paradigm for specifications lacking canonical interpretations; (2) partial automation of knowledge modeling and extraction; and (3) enabling implementation-agnostic algorithm execution and mechanized verification, thereby advancing methodological unification and practical deployment of algorithm specifications.
Current AI evaluation methods often operate in abstraction from real-world deployment contexts, failing to assess an AI system’s capacity to sustainably generate value within specific organizations. This work proposes a novel “contextual specification” framework that leverages qualitative modeling and collaborative stakeholder analysis to transform ambiguous, context-dependent elements into clearly defined, nameable constructs. By explicitly delineating the attributes, behaviors, and outcomes that warrant evaluation, the framework establishes a set of observable and measurable context-sensitive metrics. This approach provides organizations with an actionable evaluation roadmap, effectively bridging the gap between technical performance and business value, thereby substantially enhancing the relevance and efficacy of AI deployment decisions.
In software design, paradigm-implied semantic expectations—such as data abstraction consistency and feedback-control closed-loop behavior—are often left implicit, leading to design deviations and verification challenges. To address this, we introduce the concept of *design obligations*: explicit, logically formalizable, and verifiable specifications that codify such implicit constraints inherent to design paradigms. Leveraging formal modeling and paradigm semantics analysis, we establish two obligation frameworks—one for data-abstraction-based systems and another for feedback-driven adaptive systems—precisely capturing their core semantic requirements. We demonstrate that common design flaws stem from obligation violations and show how these obligations enable rigorous compliance verification and pedagogical application. This work bridges the semantic gap between design intent and implementation, providing both theoretical foundations and a methodological framework for paradigm-driven design assurance.
Algorithm engineering has long lacked a unified methodology, resulting in fragmented knowledge across subfields and poor reproducibility. To address this, this paper introduces Karl Popper’s “Three Worlds” theory—comprising ontology (clarifying problems, tasks, design, and implementation), epistemology (distinguishing descriptive from prescriptive knowledge), and methodology (systematizing knowledge evolution)—to establish the first integrated three-dimensional framework for the field. By synthesizing philosophical methodology, ontological modeling, and empirical paradigm analysis, we propose the first formal research framework for algorithm engineering, explicitly defining validity criteria for diverse scholarly contributions. This framework enhances systematicity, rigor, and cross-domain comparability in algorithm design, implementation, and evaluation. It provides foundational methodological support for disciplinary integration and advances algorithm engineering toward a mature, theory-grounded science.
This work addresses the often-overlooked optimization potential in existing published algorithms, where manual refinement is typically costly and inefficient. The authors propose a two-stage AI-assisted pipeline: first, a research-capable large language model identifies recently published algorithms that meet predefined experimental criteria; second, a Claude Code agent automatically reproduces baseline implementations and iteratively optimizes the code. This study presents the first systematic application of embodied coding agents to automate performance improvements across diverse domains of published algorithms, while underscoring the indispensable human role in defining objectives, validating outcomes, and ensuring ethical transparency. Evaluated on eleven cross-domain tasks, the approach consistently achieves performance gains, with each optimization cycle completed within a single day.
This work addresses the tightly coupled decisions in "person-to-goods" warehouse order fulfillment—such as item allocation, order batching, and picker routing—and the absence of a general mechanism to automatically compose and evaluate optimization algorithms tailored to specific operational contexts. To bridge this gap, the authors propose CASOP, a novel framework that enables, for the first time, context-aware automatic synthesis of end-to-end warehouse optimization pipelines. CASOP integrates a modular algorithm library, semantically annotated algorithm cards, a problem taxonomy, a pipeline synthesizer, and an automated evaluator to construct and validate customized solutions for given warehouse settings. Empirical evaluation across seven benchmark datasets yields 1,063,044 valid pipelines, and the accompanying open-source toolkit provides researchers and practitioners with robust support for high-performance pipeline design and selection.
This work addresses the prevalent issue in large language models (LLMs) where code generation errors often stem from misinterpretations of user requirements. Existing approaches typically assume that LLMs accurately comprehend user intent, thereby overlooking the misalignment between actual requirements and the model’s understanding. To bridge this gap, we propose REA-Coder, the first framework to explicitly integrate a requirement alignment mechanism into the code generation pipeline. REA-Coder operates within a closed-loop iterative process that parses user requirements, detects semantic discrepancies, dynamically refines prompts, and incorporates validation feedback to progressively enhance code output. Extensive experiments across four state-of-the-art LLMs and five standard programming benchmarks demonstrate that REA-Coder consistently outperforms strong baselines, achieving average performance gains ranging from 7.93% to 30.25%, thereby underscoring the critical role of requirement alignment in improving code generation accuracy.