Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对PCBA视觉问答中域迁移及输出空间异质性问题,提出了一种多模态推理框架及任务感知的GRPO方法,通过设计特定奖励机制与鲁棒推断策略有效解决了上述挑战。
📝 Abstract
In automated Printed Circuit Board Assembly (PCBA) inspection, standards-guided decisions require systems to jointly reason over fine-grained visual cues, component semantics, and manufacturing knowledge. Although large vision-language models (VLMs) provide a promising foundation, their deployment is hindered by the domain shift between standards-derived samples and real-world production-line imagery, together with heterogeneous output spaces spanning choice-based and numerical counting tasks. To address these challenges, we propose a multimodal reasoning framework for cross-domain PCBA visual question answering. The framework converts standards-derived, real-world, and auxiliary PCB-domain data into a unified instruction format and constructs verified reasoning traces aligned with visual evidence, question semantics, candidate options, and ground-truth answers. We further introduce Task-Aware Group Relative Policy Optimization (GRPO), which moves beyond exact-match supervision by integrating multi-component semantic rewards for choice-based questions, distance-aware rewards for counting questions, and an auxiliary format reward for valid outputs. During inference, answer-option semantic consistency correction, self-consistency voting, and multi-model arbitration are combined to improve prediction robustness. The proposed system achieves an Overall Score of 83.24 on the official PCBA Standard-to-Real Grand Challenge leaderboard, demonstrating the effectiveness of task-aware reward design and robust inference for cross-domain PCBA visual question answering.
Problem

Research questions and friction points this paper is trying to address.

automated PCBA inspection
domain shift
heterogeneous output spaces
Innovation

Methods, ideas, or system contributions that make the work stand out.

Task-Aware GRPO
Cross-Domain PCBA VQA
Unified Instruction Format
Semantic Rewards
Robust Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.