🤖 AI Summary
This study addresses the unclear landscape of user intents and capability requirements for multimodal large language models in real-world scenarios. By analyzing 40,000 Copilot image-based conversation logs, this work proposes a ten-tier capability taxonomy encompassing perception, cognition, and generation, systematically quantifying the diversity and distributional characteristics of cross-modal tasks. The analysis reveals a significant misalignment between existing evaluation benchmarks and actual user demands. These findings provide empirical evidence and methodological support for constructing application-oriented multimodal evaluation benchmarks grounded in authentic usage patterns.
📝 Abstract
Multimodal large language models (LLMs) increasingly integrate vision and text, yet how people use them in natural settings remains underexplored. We seek to answer the question: when users upload images, what tasks are they trying to accomplish? Analyzing over 40,000 de-identified image-upload conversations from Microsoft Copilot, we characterize real-world multimodal use through a hierarchical framework of ten capabilities, spanning perception, cognition, and generation. First, we characterize the distribution and composition of these capabilities, finding that the majority of image-upload tasks involve multiple capabilities. Second, we find that multimodal use spans a broader and more diverse task space than text-only interactions, with asymmetric coverage and task classes that rely on cross-modal grounding. Third, mapping observed capability demand onto 253 existing benchmarks reveals uneven alignment between benchmark coverage and real-world use: benchmarks concentrate on perception and reasoning toward fixed answers, while common workflows involving text, code, and data generation remain comparatively undertested. We validate our taxonomy and findings on an independent ChatGPT dataset. Our results provide a large-scale empirical characterization of what users seek to accomplish with multimodal LLMs and highlight opportunities for benchmark design grounded in observed user demand.