Scaling Textual Gradients via Sampling-Based Momentum

📅 2025-05-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the performance degradation and escalating computational cost of Textual Gradient Descent (TGD) in prompt optimization—where performance first improves then declines with increasing training samples—this paper proposes Textual Stochastic Gradient Descent with Momentum (TSGD-M). TSGD-M is the first method to incorporate momentum into textual gradient optimization, explicitly modeling and reweighting historical prompt distributions across batches to achieve low-variance, computationally efficient in-context learning. Evaluated on nine benchmark tasks—including BBH and diverse natural language understanding and reasoning benchmarks—TSGD-M consistently outperforms the TGD baseline: it achieves higher average accuracy and reduces prediction variance on 8 out of 9 tasks. These results demonstrate that TSGD-M effectively mitigates both the performance deterioration and excessive computational overhead induced by data scaling, advancing scalable and robust prompt optimization.

Technology Category

Search and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLPMachine Learning: Optimization

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Practical large-scale studies of user experience
📝 Abstract
As prompts play an increasingly critical role in large language models (LLMs), optimizing textual prompts has become a crucial challenge. The Textual Gradient Descent (TGD) framework has emerged as a promising data-driven approach that iteratively refines textual prompts using LLM - suggested updates (or textual gradients) over minibatches of training samples. In this paper, we empirically demonstrate that scaling the number of training examples initially improves but later degrades TGD's performance across multiple downstream NLP tasks. However, while data scaling improves results for most tasks, it also significantly increases the computational cost when leveraging LLMs. To address this, we draw inspiration from numerical gradient descent and propose Textual Stochastic Gradient Descent with Momentum (TSGD-M) - a method that facilitates scalable in-context learning by reweighting prompt sampling based on past batch distributions. Across nine NLP tasks spanning three domains - including BIG-Bench Hard (BBH), natural language understanding tasks, and reasoning tasks - TSGD-M significantly outperforms TGD baselines that do not incorporate reweighted sampling, while also reducing variance in most tasks.
Problem

Research questions and friction points this paper is trying to address.

Optimizing textual prompts for large language models
Scaling training examples degrades Textual Gradient Descent performance
Reducing computational cost in prompt optimization with sampling-based momentum
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scaling prompts via sampling-based momentum
Reweighting prompt sampling for efficiency
Reducing variance in NLP tasks
🔎 Similar Papers
Z
Zixin Ding
University of Chicago
J
Junyuan Hong
The University of Texas at Austin
Jiachen T. Wang
Jiachen T. Wang
Princeton University
data-centric machine learning
Zinan Lin
Zinan Lin
Microsoft Research (Redmond), Carnegie Mellon University
machine learningprivacy
Z
Zhangyang Wang
The University of Texas at Austin
Y
Yuxin Chen
University of Chicago