🤖 AI Summary
Low-resource semantic frame detection in logistics text poses significant challenges for large language models (LLMs) due to sparse annotated data and domain-specific semantics.
Method: This paper proposes an end-to-end automatic prompt optimization framework that enhances LLM inference accuracy and annotation efficiency without fine-tuning. It introduces a novel LLM-driven, self-iterative prompt optimization agent integrating retrieval-augmented generation (RAG), few-shot prompting, chain-of-thought (CoT), and automatic CoT synthesis (Auto-CoT), augmented with retrieval guidance and self-evaluation mechanisms.
Contribution/Results: To our knowledge, this is the first work to deeply couple Auto-CoT with RAG for industrial-scale frame detection—replacing conventional fine-tuning paradigms. Experiments on real-world logistics data demonstrate a 15% absolute improvement in inference accuracy over zero-shot and static-prompt baselines. The framework exhibits strong cross-model generalization across GPT-4o, Qwen2.5 (72B), and LLaMA3.1 (70B), confirming its robustness and scalability.
📝 Abstract
Prompt engineering plays a critical role in adapting large language models (LLMs) to complex reasoning and labeling tasks without the need for extensive fine-tuning. In this paper, we propose a novel prompt optimization pipeline for frame detection in logistics texts, combining retrieval-augmented generation (RAG), few-shot prompting, chain-of-thought (CoT) reasoning, and automatic CoT synthesis (Auto-CoT) to generate highly effective task-specific prompts. Central to our approach is an LLM-based prompt optimizer agent that iteratively refines the prompts using retrieved examples, performance feedback, and internal self-evaluation. Our framework is evaluated on a real-world logistics text annotation task, where reasoning accuracy and labeling efficiency are critical. Experimental results show that the optimized prompts - particularly those enhanced via Auto-CoT and RAG - improve real-world inference accuracy by up to 15% compared to baseline zero-shot or static prompts. The system demonstrates consistent improvements across multiple LLMs, including GPT-4o, Qwen 2.5 (72B), and LLaMA 3.1 (70B), validating its generalizability and practical value. These findings suggest that structured prompt optimization is a viable alternative to full fine-tuning, offering scalable solutions for deploying LLMs in domain-specific NLP applications such as logistics.