🤖 AI Summary
This study addresses the limited flexibility of structured interfaces and the inefficiency of large-scale candidate matching in rapid decision-making for multimodal systems by introducing a fully natural language-driven System-1 decision paradigm. Methodologically, we propose a dual-encoder architecture fine-tuned from the Muse-Glimmer 30B decoder, which employs bidirectional contrastive learning to project user queries and candidates into a unified semantic space. By integrating image-text and video inputs, this approach overcomes the scalability limitations of closed-set understanding and million-scale open-set retrieval. Extensive evaluation across 11 benchmarks and 190 tasks demonstrates that our method surpasses state-of-the-art models in multimodal understanding and retrieval, while clearly delineating its performance boundaries in complex reasoning tasks.
📝 Abstract
System One models output constrained decisions and probability distributions rather than free-form text generation. While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface. In this framework, both the user request and each candidate option are expressed in natural language, supported by multimodal (image and video) auxiliary inputs. We introduce MetaEncoder, which fine-tunes a pre-trained Muse-Glimmer 30B decoder into an instruction-following decision-making encoder. To scale effectively across both small closed-set (< 256) and massive open-set (millions) candidate spaces, MetaEncoder employs a bi-encoder architecture trained via unidirectional contrastive learning for request-candidate alignment. We conduct extensive evaluations across 11 benchmark suites and 190 tasks spanning multimodal decision-making, understanding (closed-set) and retrieval (open-set), highlighting where MetaEncoder beats SOTA multimodal encoders, as well as its current limits on reasoning-intensive tasks.