Efficient Multimodal Inference through Adaptive Acquisition and Sequential Fusion

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational redundancy and sequential dependency issues caused by exhaustive encoding in multimodal systems by proposing SemARC, a framework for adaptive acquisition and efficient inference. SemARC introduces a novel decoupling of order-independent marginal utility priors from residual Q-learning. By integrating a Sequential Aggregator (SeMA) for fixed-state updates, randomized subset supervision, and a Runtime Controller (ARC) for dynamic scheduling, the framework ensures predictive consistency while eliminating redundant computation. Experimental results demonstrate that SemARC improves Macro-F1 by 3.2%, reduces GFLOPs by 61.4%, and decreases end-to-end latency by 44%. Furthermore, it exhibits superior on-demand encoding and early-stopping decision-making capabilities across heterogeneous device environments.
📝 Abstract
Multimodal systems often encode every available input, even when a subset suffices for prediction. Adaptive acquisition can reduce this cost by using predictions from incrementally fused evidence to decide which modality to encode next and when to stop. However, sequential fusion makes these predictions order-dependent, so decisions based on them may need to distinguish factorially many histories of the same acquired set. We introduce SemARC, which couples a Sequential Modality Aggregator (SeMA) with an Adaptive Runtime Controller (ARC) and uses acquired evidence to select each modality before its encoder runs. SeMA executes only selected encoder and fusion branches, updates a fixed-size state, and predicts after each acquisition without recomputing earlier branches. We supervise every acquisition prefix under randomized modality subsets and orders to encourage consistent predictions across acquisition orders. ARC combines a set-dependent marginal-utility prior with residual fitted-Q learning to select the next available modality or stop, without inspecting unacquired inputs or retaining acquisition order. Across six multimodal classification datasets and eleven baselines, SemARC achieves 3.2% higher macro-F1 and 61.4% lower total inference GFLOPs on average relative to each dataset's most accurate baseline. End-to-end latency falls by 44.0% across GPU and CPU and by 47.2% on Android INT8 relative to the fastest measured baseline, on average. Under varying runtime modality missingness, SemARC still skips available modalities, matching or exceeding the best baseline macro-F1 in 21 of 24 conditions with 14.8% lower total GFLOPs on average. SemARC thus offers a practical path toward efficient multimodal inference across heterogeneous devices.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Inference
Adaptive Acquisition
Sequential Fusion
Computational Efficiency
Order Dependence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Acquisition
Sequential Fusion
Multimodal Inference
Fitted-Q Learning
Efficient Computing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Payal Mohapatra
Payal Mohapatra
Northwestern University | Analog Devices Inc. | IIT Madras
Time SeriesWearablesMachine Learning
H
Haodong Yang
Northwestern University, USA
Y
Yueyuan Sui
Northwestern University, USA
Stephen Xia
Stephen Xia
Northwestern University
Embedded IntelligenceMobile and Embedded SystemsCyber Physical SystemsSmart Environments
B
Benjamin Lundell
Arm Inc., USA
Q
Qi Zhu
Northwestern University, USA