STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts

๐Ÿ“… 2026-09-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the absence of scalable evaluation frameworks for text-to-trajectory generation and its reliance on costly human annotation by proposing the first automated context-alignment evaluation framework. Methodologically, it leverages sociological theories to define the evaluation space and employs natural language processing to decompose contextual information into behavioral questions matched with deterministic measurement instruments. Furthermore, a VRDST protocol is introduced to enable verifiable, cross-scenario adaptive evaluation without requiring human-annotated data. The authors construct a benchmark comprising 1,000 scenarios that achieves 80% agreement with human judgments. Experimental results reveal significant deficiencies in existing models regarding fine-grained contextual conditioning, highlighting critical directions for future trajectory generation research.
๐Ÿ“ Abstract
Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult. We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data. We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation.
Problem

Research questions and friction points this paper is trying to address.

text-to-trajectory generation
trajectory evaluation
context alignment
pedestrian trajectory
scalable evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-to-Trajectory Generation
Automated Evaluation
Context Alignment
Pedestrian Trajectory
Benchmark
๐Ÿ”Ž Similar Papers
No similar papers found.