Correcting to Predict: Pseudo-Value Correction for Multimodal Attribute Value Extraction

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of implicit attribute ambiguity and the underutilization of visual-textual cues in multimodal product attribute extraction by proposing the C2P framework. This approach reformulates attribute extraction as an error-correction process, introducing a novel pseudo-value guidance mechanism that drives the model to refine predictions based on multimodal evidence. Furthermore, a self-consistency optimization strategy is incorporated to enable efficient output generation within a single inference pass, eliminating the need for online retrieval or iterative refinement. Experimental results demonstrate that C2P significantly outperforms existing baselines on both public benchmarks and a large-scale AliExpress dataset, effectively enhancing seller adoption rates, attribute completeness, and user engagement.
📝 Abstract
Product attribute value extraction (AVE) is a fundamental task in e-commerce, aiming to identify specific values of predefined attributes from multimodal product profiles such as text and images. While multimodal large language models (MLLMs) have shown promise for AVE, they face challenges in extracting implicit attributes that require joint reasoning over visual and textual cues, often confusing semantically similar values. However, existing methods often fail to resolve such ambiguities because the correct value often depends on subtle multimodal cues that are easy to miss or override. To address this challenge, we propose Correcting to Predict (C2P), a framework that treats attribute extraction as a correction process. Given an initial pseudo-value such as a retrieved candidate or placeholder, the model learns to correct it using multimodal evidence. During training, diverse pseudo-values help the model learn evidence-based correction behavior, and a self-consistency refinement stage further reduces sensitivity to pseudo-value perturbations. At inference, a fixed placeholder triggers the learned correction behavior, enabling efficient single-pass prediction without online retrieval or iterative refinement. We evaluate C2P on a public benchmark and a large-scale industrial dataset. Offline results show that C2P outperforms strong baselines, with notable gains on ambiguous attributes. Online A/B tests on AliExpress further show consistent improvements in seller adoption, attribute completeness, and user engagement, validating C2P's effectiveness and efficiency in real-world deployment.
Problem

Research questions and friction points this paper is trying to address.

Attribute Value Extraction
Multimodal Large Language Models
Implicit Attributes
Semantic Ambiguity
E-commerce
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Attribute Value Extraction
Pseudo-Value Correction
Correcting to Predict
Self-Consistency Refinement
Multimodal Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Junhao Zhang
Junhao Zhang
National University of Singapore; Shandong University
Computer Vision
F
Feiran Hu
Alibaba International Digital Commerce Group, Hangzhou, China
X
Xiao Hu
Alibaba International Digital Commerce Group, Hangzhou, China
B
Baoliang Cui
Alibaba International Digital Commerce Group, Hangzhou, China
X
Xiaoyi Zeng
Alibaba International Digital Commerce Group, Hangzhou, China