MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

๐Ÿ“… 2026-07-10
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Micro-actions pose significant challenges for standardized modeling and evaluation due to their short duration, subtle motion dynamics, and fine-grained semantic distinctions. To address this, this work introduces the Third Micro-Action Challenge (MAC 2026), which pioneers a fine-grained micro-action understanding task, advancing research beyond detection toward deep semantic interpretation. The challenge establishes a robust evaluation framework grounded in publicly available datasets, standardized protocols, and state-of-the-art video understanding techniques. Notably, it innovatively leverages multimodal large language models to assess participantsโ€™ ability to capture and generate nuanced micro-action semantics. By aggregating contributions from leading global teams, MAC 2026 substantially expands the research frontier in human-centric video understanding.
๐Ÿ“ Abstract
Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communication. However, their short duration, weak motion patterns, and fine-grained semantic differences make them difficult to annotate, model, and evaluate in a standardized manner. To promote academic research on micro-action analysis, we proposed and have annually organized the Micro-Action Analysis Grand Challenge (MAC) as a public benchmark platform for this emerging field. The first two editions of MAC established standardized evaluation settings for micro-action recognition and detection, providing publicly accessible datasets and protocols. Building upon these editions, this paper presents the 3rd MAC, held in conjunction with ACM Multimedia 2026. Under the theme of moving from recognition to fine-grained micro-action understanding, this edition further expands the scope of the challenge beyond conventional recognition and detection. In particular, we introduce a new task named fine-grained micro-action understanding, evaluated with the assistance of multimodal large language models, aiming to assess models'ability to capture fine-grained semantic cues and interpret subtle human micro-actions at a deeper level. We summarize the datasets, task settings, evaluation protocols, competition results, and representative solutions from top-performing teams. Finally, we discuss future directions for micro-action analysis and its broader role in human-centric video understanding.
Problem

Research questions and friction points this paper is trying to address.

Micro-Actions
fine-grained understanding
non-verbal cues
standardized evaluation
semantic differences
Innovation

Methods, ideas, or system contributions that make the work stand out.

fine-grained micro-action understanding
multimodal large language models
micro-action analysis
non-verbal cues
human-centric video understanding
๐Ÿ”Ž Similar Papers