IndustryLLM: Failure-Driven LLM Training for Industrial Procurement

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the semantic gap between informal terminology, sparse attributes, and stringent engineering standards in industrial procurement, along with associated factual fragility. Building upon Qwen3.5, we propose a failure-driven domain adaptation framework that integrates national standards with technical archives. Methodologically, we design multi-genre rewriting, confidence-routed editing, and error-guided QA synthesis mechanisms, while introducing an evidence-gated three-valued logic interface to ensure output reliability. The training pipeline unifies continual pre-training, supervised fine-tuning, and a Mixture-of-Experts architecture. Experimental results demonstrate that the proposed approach improves structured query accuracy by nearly 3%, increases online Gross Merchandise Volume (GMV) by 4.25%, enhances user satisfaction by 8.3%, and reduces inference latency to 1.5 seconds, validating its effectiveness for real-world industrial applications.
📝 Abstract
Industrial procurement requires language models to bridge informal buyer jargon, sparse marketplace attributes, and authoritative engineering standards under strict safety tolerances. We present IndustryLLM, an open-weight industrial language model trained from Qwen3.5-35B-A3B-Base (35B total parameters with ~3B activated per token, with the vision encoder frozen). Rather than relying on generic text scaling, we introduce a failure-driven adaptation recipe spanning continued pre-training (CPT) and supervised fine-tuning (SFT). CPT leverages a curated ~100B-token corpus integrating 5B tokens of national standards (e.g., GB/T) and technical archives, 10B tokens of de-identified real-world industrial transaction and inquiry records, and 60B tokens of general replay. To overcome register mismatch and factual brittleness, we systematically reconstruct an estimated 20B-token domain subset via multi-register rewriting across 10 genres and 8 writing styles, confidence-routed minimal factual editing, and error-targeted QA synthesis (resolving colloquial typos like'42-luo-mu'->42CrMo, expanding ambiguous codes like'16674'->GB/T 16674, and clarifying conflicting dimensional specs). For downstream deployment, we formalize an evidence-gated constraint-evaluation interface enforcing three-valued logic where unverified product evidence remains unknown rather than satisfied. Offline evaluations demonstrate consistent gains on procurement-query structuring (+2.97 percentage points in exact match, 95% CI [2.11, 3.86] in No-Think mode), while randomized online A/B experiments in production yield substantial improvements (+4.25% GMV, +8.3% satisfied inquiries) alongside a latency reduction from 6-7 s to 1.5 s. Model weights and configs are released at https://huggingface.co/alibaba-multimodal-industrial-ai/IndustryLLM.
Problem

Research questions and friction points this paper is trying to address.

Industrial Procurement
Language Model Adaptation
Register Mismatch
Factual Brittleness
Constraint Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Failure-Driven Adaptation
Industrial Procurement
Multi-Register Rewriting
Evidence-Gated Constraint Evaluation
Continued Pre-Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.