MOVE: Multimodal Open-world Verification and Expansion for Graph Learning

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of novel class emergence after deployment and the inability of fixed label spaces to identify unknown nodes in multimodal graph learning by proposing an open-world learning framework. Methodologically, it integrates visual, textual, and graph contextual features for unknown node detection. Furthermore, it leverages multimodal large language models to generate candidate categories and introduces a reliability verification mechanism based on multimodal evidential consistency to filter novel classes, thereby overcoming single-modality limitations and preventing the introduction of redundant categories. Experimental results demonstrate that the proposed method achieves an average improvement of 11.87% across unknown node identification, open-domain annotation, and downstream tasks, significantly enhancing model generalization capability.
πŸ“ Abstract
Multimodal graph learning faces a fundamental challenge: new classes may emerge after deployment, while models are trained with a fixed label space. Existing approaches typically detect unknown nodes and use LLMs to generate candidate class descriptions, but they do not determine whether existing classes are insufficient to cover these nodes or whether a generated class is reliable enough to expand the class space. Our empirical study reveals three challenges: multimodal information beyond individual modalities is required for unknown-node identification, LLM-generated class descriptions may not fully capture multimodal class characteristics, and directly adding candidate classes can introduce redundant categories. Based on these observations, we propose MOVE, a multimodal open-world class verification and expansion framework. MOVE identifies nodes that cannot be assigned to existing classes by jointly considering visual tokens, textual attributes, and graph context, leverages a multimodal LLM to generate candidate classes, and selectively expands the class space only when candidates are consistently supported by multimodal evidence without introducing unnecessary categories. Experiments demonstrate that MOVE achieves an average improvement of 11.87\% across unknown recognition, open-domain annotation, and downstream graph learning tasks.
Problem

Research questions and friction points this paper is trying to address.

Multimodal graph learning
Open-world learning
Unknown node identification
Class expansion
Large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Graph Learning
Open-world Class Expansion
Unknown Node Detection
Multimodal Large Language Model
Class Verification
πŸ”Ž Similar Papers
No similar papers found.