Can Motion-Language Models Ground Structure? STRIDE for Evaluating the Evaluators

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of existing motion-language model evaluators to overestimate their grounding capabilities for structural semantics such as temporal ordering and mirroring, often relying excessively on textual priors while neglecting actual motion structures. To this end, we propose STRIDE, a benchmark that systematically quantifies the structural blind spots of these evaluators, revealing that conventional metrics fail to expose deficiencies in mirror sensitivity. Furthermore, we introduce a likelihood-balanced debiasing technique, validated through structured hard negative optimization strategies derived from contrastive learning. Experimental results demonstrate that our approach effectively mitigates the structural shortcomings of mainstream evaluators, yielding significant performance improvements on both temporal and mirroring tasks.
📝 Abstract
Motion-language models are typically scored by motion-language evaluators, but how well these evaluators ground language structure remains unclear. Here, we introduce the Structure grounding via Temporal-order, Reflection, and Identity Diagnostic Evaluation (STRIDE) benchmark to systematically evaluate the ability of evaluators to track temporal order, mirror reflection, and action identity. STRIDE comprises $5{,}869$ triples, each consisting of a motion, its original caption, and a perturbed caption, spanning both short and long descriptions. We likelihood-balance caption pairs to reduce text-only bias and estimate each evaluator's caption preference under unrelated motions to measure the discrimination gain from matched motions relative to this baseline. Our experiments reveal weak structural grounding and severe deficits in mirror sensitivity among the audited evaluators, which commonly used evaluation protocols fail to expose. To understand why these limitations go undetected in standard tests, we examine the evaluators more closely. We find that text-only priors alone can solve naive perturbation tests on existing datasets, while common retrieval and distributional metrics barely respond to structural corruption introduced by mirroring ground-truth motions. These findings suggest a natural intervention: structural hard negatives. Our experiments show that a simple modification to contrastive learning substantially improves performance on temporal order and mirror reflection. The benchmark and code will be released.
Problem

Research questions and friction points this paper is trying to address.

Motion-Language Models
Evaluation Benchmark
Structural Grounding
Temporal Order
Mirror Reflection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Motion-Language Models
STRIDE Benchmark
Structural Grounding
Likelihood Balancing
Hard Negatives
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lixing Tan
Beihang University
Qing Xia
Qing Xia
Wenzhou Kean University
Numerical AnalysisScientific ComputingApplied Mathematics
Y
Yuting Guo
Beijing Information Science and Technology University
Shuai Li
Shuai Li
Beihang University
A
Aimin Hao
Beihang University