Score
Design, build, and analyze iterative refinement workflows that generate explanatory principles or corrective rules from model behavior, test those principles against model outputs (often using simulation-in-the-loop), diagnose failures, and apply mixed-precision or other iterative-refinement steps to update model components or supervision signals until the principles and model outputs meet desired consistency and performance criteria.
In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.
This study addresses the common omission in current large language model evaluations of code generation—the iterative refinement process inherent in real-world programming and the models’ capacity for self-correction using feedback. The authors propose a novel framework that leverages execution-based feedback, such as compilation errors and test failures, to systematically investigate how reasoning and non-reasoning models utilize such signals across multiple programming languages. Through multidimensional categorization of code failures and extensive cross-model, cross-language experiments, they demonstrate that reasoning models consistently improve over iterations and significantly outperform non-reasoning counterparts. While syntactic and runtime errors prove relatively amenable to correction, logical and algorithmic errors remain challenging, thereby delineating the current limits of feedback-driven repair mechanisms.
This study addresses the limitation of traditional autoformalization methods that rely on fine-grained feedback, which often triggers cross-level verification conflicts. To this end, we propose ProGS, a novel approach that introduces a tree-structured proof sketch to guide the construction and repair of formal models using large language models. By precisely mapping verification failures to specific nodes within the proof tree, ProGS enables structured iterative refinement while synergizing with formal tools to enhance automation. Experimental evaluations across 27 benchmark systems demonstrate that our method significantly improves syntactic validity, deductive verifiability, and behavioral correctness compared to existing baselines.
Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.
This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.
研究通过不同模型大小在自精炼管道各阶段的效果,发现生成和修正阶段需较大模型,而批评阶段对模型大小不敏感,为设计更高效的语言模型系统提供指导。
为解决从正交视图生成CAD代码时出现的几何不一致问题,提出IterCAD框架,通过迭代修正方法逐步改进代码准确性。
This work addresses the lack of structural correctness verification during the design phase in existing AI agent workflow platforms, which typically rely on runtime safeguards. The authors propose a workflow modeling approach centered on reusable building blocks and introduce, for the first time, a set of twelve structural rules. By leveraging graph-based representations and a rule engine, the method enables static, formal checks for compatibility and logical consistency at design time. Experimental evaluation demonstrates that the prototype system efficiently detects design violations on a dataset comprising 48 defective workflows and 168 structural variants, maintaining high detection accuracy even when tasks are split across multiple agents. This significantly enhances the reliability and maintainability of workflow designs.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
This study addresses the lack of systematic understanding regarding the impact of repair loop iteration counts in large language model (LLM)-based software engineering tasks, where prior work often relies on arbitrarily defined repair budgets. Through a cross-task (code generation, test generation, code translation) and cross-model empirical analysis, this work reveals—for the first time—a pronounced diminishing marginal returns phenomenon in iterative repair: performance gains are concentrated within the first 3–4 iterations, with negligible improvements thereafter. The findings underscore that the design of the repair workflow and feedback mechanisms exerts a far greater influence on repair efficacy than the choice of LLM itself. The authors advocate for treating repair budget as a critical experimental variable to ensure reliable, computationally efficient, and reproducible evaluation outcomes.