MetaEncoder: Exploring the Limit of Bi-Encoders for Multimodal System One Decision Making with Natural Language Interface

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited flexibility of structured interfaces and the inefficiency of large-scale candidate matching in rapid decision-making for multimodal systems by introducing a fully natural language-driven System-1 decision paradigm. Methodologically, we propose a dual-encoder architecture fine-tuned from the Muse-Glimmer 30B decoder, which employs bidirectional contrastive learning to project user queries and candidates into a unified semantic space. By integrating image-text and video inputs, this approach overcomes the scalability limitations of closed-set understanding and million-scale open-set retrieval. Extensive evaluation across 11 benchmarks and 190 tasks demonstrates that our method surpasses state-of-the-art models in multimodal understanding and retrieval, while clearly delineating its performance boundaries in complex reasoning tasks.
📝 Abstract
System One models output constrained decisions and probability distributions rather than free-form text generation. While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface. In this framework, both the user request and each candidate option are expressed in natural language, supported by multimodal (image and video) auxiliary inputs. We introduce MetaEncoder, which fine-tunes a pre-trained Muse-Glimmer 30B decoder into an instruction-following decision-making encoder. To scale effectively across both small closed-set (< 256) and massive open-set (millions) candidate spaces, MetaEncoder employs a bi-encoder architecture trained via unidirectional contrastive learning for request-candidate alignment. We conduct extensive evaluations across 11 benchmark suites and 190 tasks spanning multimodal decision-making, understanding (closed-set) and retrieval (open-set), highlighting where MetaEncoder beats SOTA multimodal encoders, as well as its current limits on reasoning-intensive tasks.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Decision Making
System One
Bi-Encoder
Natural Language Interface
Contrastive Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

MetaEncoder
Bi-encoder
Multimodal System One
Contrastive Learning
Natural Language Interface
🔎 Similar Papers
No similar papers found.