When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges

📅 2026-05-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges in multi-objective prompt optimization, where text gradient methods often fail due to gradient conflicts and instruction interference. It identifies and distinguishes two distinct failure modes: gradient dilution during optimization and instruction interference during inference, thereby clarifying the design boundaries for multi-objective judge customization. Drawing on multi-task learning principles, the work proposes five information-sharing architectures that decouple interactions among losses, gradients, and the language model at different levels of abstraction. Experimental results reveal that, among ten configurations, six fail to outperform the initial prompt; gradient specificity drops by 59%; and naively merging task instructions reduces the Spearman correlation coefficient by 5.3%, collectively highlighting the inherent difficulties and limitations of multi-objective text gradient optimization.
📝 Abstract
Customizing an LLM judge to a specific task or domain often involves optimizing its prompt across multiple evaluation criteria simultaneously. Textual gradient methods automate this for a single judge criterion, however they produce natural-language critiques, not numerical vectors. Thus, the conflict-resolution toolkit of multi-task learning (PCGrad, MGDA) doesn't apply to the multi-objective textual gradient setting. We test five decomposition modes of textual gradient optimizers by varying how much cross-task information the loss, gradient and optimizer LLMs share. In 6 of 10 configurations, we observe that optimization never improves over the initial prompt. Gradient specificity drops by 59% (from 9.0 to 3.7) when the gradient LLM processes multiple criteria jointly. Separately, we observe that naively combining per-task instructions into a single prompt degrades Spearman's rho by -5.3%. These results identify two separable failure modes: optimization-time gradient dilution and inference-time instruction interference, which together constrain the design space for multi-objective judge customization using textual feedback.
Problem

Research questions and friction points this paper is trying to address.

multi-objective optimization
textual gradients
LLM judges
prompt optimization
instruction interference
Innovation

Methods, ideas, or system contributions that make the work stand out.

textual gradients
multi-objective optimization
LLM judges
gradient dilution
instruction interference
💼 Related Jobs
No related jobs found.
P
Parth Darshan
IIT Jodhpur
A
Abhishek Divekar
Amazon