Two Vectors Replace In-Context Demos: Structured Task Adaptation via Embeddings

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost of repetitive encoding in multimodal in-context learning and the limitations of existing demonstration-free methods, whose parameters scale with model depth and lack attention modulation. To overcome these challenges, this work proposes STAVE, a method that replaces demonstration examples with two task-specific vectors to update answer generation and structural token embeddings, respectively, enabling efficient adaptation. The effectiveness of this dual-vector design is theoretically justified through first-order loss analysis and margin bound theory. Experimental results demonstrate that STAVE matches or surpasses state-of-the-art performance across both multimodal and text-only tasks. With minimal additional parameters, it outperforms 15-shot in-context learning while maintaining zero inference overhead.
📝 Abstract
In-context learning (ICL) adapts frozen large multimodal models (LMMs) to new tasks from a few demonstrations (demos), but re-encodes them at every query, where each demo image adds up to hundreds of visual tokens. Demo-free methods remove this cost with a compact task state. However, they add it at locations searched per task or at every decoder layer, where task parameters grow with depth. Moreover, inserted tokens or keys cannot change how the original prompt divides its attention within a layer. To address these issues, we propose Structured Task Adaptation via Embeddings (STAVE), which replaces demos with two task-specific vectors added to existing input embeddings. Specifically, a readout vector updates the answer-producing tokens and a context vector updates the other structural token groups. Both are trained with answer labels on prompts with and without demos. We justify these design choices theoretically using a first-order analysis of the loss and a margin bound. Extensive experiments on six LMMs and five large language models show that STAVE matches or outperforms state-of-the-art methods on multimodal tasks with far fewer task parameters and surpasses 15-shot ICL and prior task vectors on 18 text tasks, all at zero-shot inference cost.
Problem

Research questions and friction points this paper is trying to address.

In-context learning
Large multimodal models
Task adaptation
Computational overhead
Attention allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Task Vectors
Large Multimodal Models
Structured Task Adaptation
Zero-Shot Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xi Ding
University of Wisconsin–Madison
Naichen Shi
Naichen Shi
University of Michigan
J
Jiawei Zhang
University of Wisconsin–Madison