Visual Jev: Accurate and Efficient Decisions from Shared Visual Context

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究通过Visual Jev方法,利用共享视觉上下文一次性编码图像并批量处理问题,提高了决策准确性和效率。
📝 Abstract
Many vision applications ask several independent, forced-choice questions about the same image. Visual Jev encodes the image and public context once, executes isolated question suffixes as a batch, and reads candidate probabilities from the backbone's language-model head. Across four benchmarks, answer-supervised post-training raises equal-weight macro accuracy from 70.6% to 76.1%, with the gain concentrated on the two task families represented in training. At N=32 questions per image, shared batched execution is 8.9x faster in warm amortized time than independent serial execution and remains 3.4x faster than an already-batched baseline that recomputes the prefix, at the cost of higher peak memory. A matched typed-head control offers no consistent accuracy advantage over the language-model-head readout. The supported design is therefore simple: adapt the backbone for quality, retain the existing readout, and share execution for efficiency.
Problem

Research questions and friction points this paper is trying to address.

visual context
forced-choice questions
batch execution
accuracy
efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Shared Visual Context
Batch Execution
Language-Model Head
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Guanxu Yu
Independent Research Carnegie Mellon University
Yuhang Yao
Yuhang Yao
Carnegie Mellon University
Federated Graph LearningSuper Agent System