Can Gradient Descent Simulate Prompting?

📅 2025-06-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether parameter fine-tuning can replicate the few-shot generalization and logical reasoning advantages of prompting. To this end, we propose a meta-learning framework wherein a single gradient step dynamically approximates the model’s own prompt-based outputs—enabling self-supervised objective construction without ground-truth labels. The method integrates gradient-based meta-training, prompt-driven parameter-space optimization, and a one-step update mechanism. Its core innovation lies in using the model’s own prompt responses as the target function for gradient updates—the first such formulation—thereby revealing the inherent learnability of prompting behavior via fine-tuning. Experiments demonstrate substantial improvements on logical reasoning tasks (e.g., “reverse curse”), where our approach achieves or matches prompting performance with only one gradient update. This establishes a new paradigm for low-overhead, highly generalizable model adaptation.

Technology Category

Natural Language Processing: Prompt Engineering / PromptingSearch and Optimization: Metareasoning and MetaheuristicsMachine Learning: Large Multimodal Models (LMMs)

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
There are two primary ways of incorporating new information into a language model (LM): changing its prompt or changing its parameters, e.g. via fine-tuning. Parameter updates incur no long-term storage cost for model changes. However, for many model updates, prompting is significantly more effective: prompted models can generalize robustly from single examples and draw logical inferences that do not occur under standard fine-tuning. Can models be modified so that fine-tuning does emulate prompting? This paper describes a method for meta-training LMs such that gradient updates emulate the effects of conditioning on new information. Our approach uses tools from gradient-based meta-learning but uses an LM's own prompted predictions as targets, eliminating the need for ground-truth labels. Subsequent gradient descent training recovers some (and occasionally all) of prompted model performance -- showing improvement on the ``reversal curse'' tasks, and answering questions about text passages after a single gradient update. These results suggest that, with appropriate initialization, gradient descent can be surprisingly expressive. Our results suggest new avenues for long-context modeling and offer insight into the generalization capabilities of gradient-based learning.
Problem

Research questions and friction points this paper is trying to address.

Can gradient descent simulate prompting in language models?
Improving fine-tuning to match prompting effectiveness
Meta-training LMs for gradient updates mimicking prompting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Meta-training LMs to emulate prompting via gradients
Using LM's own predictions as targets, no labels needed
Gradient descent recovers prompted performance post-update