🤖 AI Summary
This study addresses the challenge of skill acquisition for LLM-based agents in industrial planning, which is hindered by heterogeneous and incomplete multi-source knowledge. To this end, we propose a cross-source skill induction and execution verification framework. The method extracts coordination patterns from discrepancies between documentation and practice, compiles them into standardized skill packages anchored by COM specifications, and achieves iterative refinement through unlabeled compliance screening and signal-conditioned trajectory attribution. Evaluated on the PIMS-Bench benchmark across four LLM backbones, the proposed approach yields absolute improvements of 14%–30% in component matching F1 scores, with particularly significant gains observed on complex tasks.
📝 Abstract
Operating industrial planning software such as AspenTech PIMS (Process Industry Modeling System) requires an LLM agent to combine structural knowledge, procedural knowledge from expert records, and constraints revealed only during execution. These evidence sources are heterogeneous and individually incomplete: documentation describes tables and interfaces but omits task-level coordination, expert CASE records expose multi-table modification patterns without explicit schema grounding, and execution feedback reveals latent constraints only when a plan is executed. We present \textsc{SkillRefine}, a framework that exploits the Documentation--Practice Gap between documented structure and expert practice to extract candidate coordination patterns, ground them against table definitions and COM specifications, and compile them into progressively disclosed skill packages with provenance. The library is then refined through label-free compliance screening, oracle-based match decomposition over table, row, column, and value dimensions, and signal-conditioned trajectory attribution for localized repair. We evaluate \textsc{SkillRefine} on PIMS-Bench, a benchmark built from two AspenTech PIMS demonstration models, using disjoint construction and held-out test tasks. On the held-out test set, \textsc{SkillRefine} achieves absolute component match F1 gains of 14\%--30\% across four LLM backbones, with the largest gains on complex multi-table coordination tasks.