GraphSelect for Budgeted Representation Selection in Multimodal Graph Inference

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of selecting representative subsets under budget constraints in multimodal graph reasoning by proposing GraphSelect. This method optimizes text and image subset selection through joint evaluation of swap operations, dynamically updating scores via class stability to perform input exchanges, thereby effectively overcoming node contribution dependencies and information propagation effects. Technically, it integrates multimodal graph neural networks, greedy search, and joint loss optimization mechanisms. Experimental results demonstrate that retaining only 20% of candidate representations incurs a mere 0.1 percentage point drop in model accuracy. The proposed approach significantly outperforms existing attribution methods, achieving an excellent balance between computational efficiency and predictive precision.
📝 Abstract
Multimodal graph predictors combine text, images, and relations to classify connected entities. How much of this input is needed to preserve their predictions? We study budgeted representation selection, which chooses a subset of candidate text and image vectors under a separate capacity for each modality. Predictions from the complete candidate input define the classes to preserve. The challenge is that a representation's contribution depends on the other selected inputs, while graph propagation extends its effects across nodes. Our empirical study shows that candidate rankings change with the selected input, while predicted probabilities remain informative after the class stops changing. Updating scores improves selection, and exchanging inputs can improve a subset whose capacity is already filled. These findings lead to GraphSelect, which starts from individual candidate gains and refines the subset through jointly evaluated exchanges. It screens promising removals and additions, accepts an exchange when it reduces the prediction loss, and updates the scores. Experiments on six graphs show higher mean objective recovery than six attribution and explanation methods adapted to the selection task. Across nine trained architectures on two graphs, retaining 20% of the candidate representations per modality gives a mean accuracy drop of 0.10 percentage points relative to full candidate input, preserving classification performance with substantially fewer text and image representations.
Problem

Research questions and friction points this paper is trying to address.

multimodal graph inference
budgeted representation selection
subset selection
representation contribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Budgeted Representation Selection
Multimodal Graph Inference
GraphSelect
Jointly Evaluated Exchanges
Capacity Constraint
🔎 Similar Papers
No similar papers found.