Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation

📅 2024-03-04
🏛️ arXiv.org
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the generalization capability of multimodal large language models (MLLMs) on fine-grained visual attribute estimation—specifically, human age and gender prediction. We benchmark general-purpose MLLMs (e.g., GPT-4V, LLaVA-Next, ShareGPT4V) against the task-specialized model MiVOLO, revealing for the first time that zero-shot MLLMs achieve up to 70% of MiVOLO’s accuracy. We propose a lightweight, annotation-scenario-oriented fine-tuning paradigm: instruction-tuning ShareGPT4V alone reduces AgeMAE to 2.81 years—nearly matching MiVOLO’s 2.67 years. Through prompt engineering and cross-model benchmarking, we identify canonical bottlenecks in MLLMs’ logical reasoning and numerical precision. Our core contributions are threefold: (1) establishing MLLMs as high-quality visual annotation tools; (2) introducing the first MLLM-specific fine-tuning framework tailored to fine-grained attribute estimation; and (3) publicly releasing a standardized evaluation benchmark and analytical protocol.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision ModelsNatural Language Processing: (Large) Language Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Multimodal Large Language Models (MLLMs) have recently gained immense popularity. Powerful commercial models like ChatGPT-4V and Gemini, as well as open-source ones such as LLaVA, are essentially general-purpose models and are applied to solve a wide variety of tasks, including those in computer vision. These neural networks possess such strong general knowledge and reasoning abilities that they have proven capable of working even on tasks for which they were not specifically trained. We compared the capabilities of the most powerful MLLMs to date: ShareGPT4V, ChatGPT, LLaVA-Next in a specialized task of age and gender estimation with our state-of-the-art specialized model, MiVOLO. We also updated MiVOLO and provide details and new metrics in this article. This comparison has yielded some interesting results and insights about the strengths and weaknesses of the participating models. Furthermore, we attempted various ways to fine-tune the ShareGPT4V model for this specific task, aiming to achieve state-of-the-art results in this particular challenge. Although such a model would not be practical in production, as it is incredibly expensive compared to a specialized model like MiVOLO, it could be very useful in some tasks, like data annotation.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Age and Gender Recognition
Specialized Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Fine-tuning ShareGPT4V
Enhancements to MiVOLO
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
SaluteDev