ADAPTIVE TOKEN BOUNDARIES: INTEGRATING HUMAN CHUNKING MECHANISMS INTO MULTIMODAL LLMS

๐Ÿ“… 2025-05-03
๐Ÿ›๏ธ Social Science Research Network
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing multimodal large language models (MLLMs) employ static cross-modal tokenization, limiting their ability to emulate human-like, context-sensitive integration of multimodal information. To address this, we introduceโ€” for the first time in MLLMsโ€”the cognitive science principle of *chunking* into tokenizer design, proposing an adaptive cross-modal tokenization framework. Our method comprises three core components: differentiable dynamic boundary learning, hierarchical multi-granularity representation, and vision-language alignment-guided attention. This framework departs from conventional fixed-tokenization paradigms by enabling semantic-driven, context-aware token segmentation. Evaluated on visual question answering (VQA) and complex scene description tasks, our approach achieves absolute improvements of 7.8% and 5.3%, respectively. Moreover, error patterns and attention distributions align significantly more closely with human cognitive behavior. Our work establishes a novel paradigm for developing human-inspired multimodal understanding models.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
๐Ÿ“ Abstract
Recent advancements in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in processing diverse data types, yet significant disparities persist between human cognitive processes and computational approaches to multimodal information integration. This research presents a systematic investigation into the parallels between human cross-modal chunking mechanisms and token representation methodologies in MLLMs. Through empirical studies comparing human performance patterns with model behaviors across visual-linguistic tasks, we demonstrate that conventional static tokenization schemes fundamentally constrain current models' capacity to simulate the dynamic, context-sensitive nature of human information processing. We propose a novel framework for dynamic cross-modal tokenization that incorporates adaptive boundaries, hierarchical representations, and alignment mechanisms grounded in cognitive science principles. Quantitative evaluations demonstrate that our approach yields statistically significant improvements over state-of-the-art models on benchmark tasks (+7.8% on Visual Question Answering, +5.3% on Complex Scene Description) while exhibiting more human-aligned error patterns and attention distributions. These findings contribute to the theoretical understanding of the relationship between human cognition and artificial intelligence, while providing empirical evidence for developing more cognitively plausible AI systems.
Problem

Research questions and friction points this paper is trying to address.

Bridging human cognitive chunking and MLLM token representation gaps
Overcoming static tokenization limits in multimodal human-like processing
Enhancing MLLMs with dynamic cross-modal tokenization for cognitive alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic cross-modal tokenization with adaptive boundaries
Hierarchical representations for multimodal data integration
Cognitive science-based alignment mechanisms for human-like processing
๐Ÿ”Ž Similar Papers
Sanda University
D
Dongxing Yu
School of Education, Sanda University, Shanghai, China