parse actions

Designs and implements methods and pipelines to segment temporal or sequential inputs into discrete action steps and to map observations into action tokens, including mechanisms to separate observable (visible) from latent (hidden) action labels. Includes techniques for extracting and recording action trajectories produced by fixed prompts or models for analysis, evaluation, and debugging.

parseactions

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Looking into the Unknown: Exploring Action Discovery for Segmentation of Known and Unknown Actions

Aug 07, 2025
FS
Federico Spurio
🏛️ University of Bonn | Birzeit University | Toyota Motor Europe

This work addresses the challenges of partial annotation—where only “known actions” are labeled while numerous “unknown actions” remain unlabeled—and ambiguous action boundaries in temporal action segmentation. To this end, we formally introduce a novel task termed “action discovery”: jointly identifying, segmenting, and clustering both known and unknown actions under weak supervision. Methodologically, we propose a Granularity-Guided Segmentation Module (GGSM) to capture multi-granularity temporal structures, and an Unknown Action Segment Assignment module (UASA) that jointly leverages temporal modeling and embedding similarity learning, using known actions as semantic anchors to guide the temporal localization and semantic clustering of unknown actions. Extensive experiments on Breakfast, 50Salads, and Desktop Assembly demonstrate significant improvements over existing state-of-the-art methods, validating the approach’s effectiveness and generalizability in realistic, boundary-ambiguous behavioral scenarios.

Classifying unknown actions using learned semantic similaritiesIdentifying unknown actions in partially labeled datasetsSegmenting ambiguous and unannotated actions temporally

From Observation to Action: Latent Action-based Primitive Segmentation for VLA Pre-training in Industrial Settings

Nov 26, 2025
JZ
Jiajie Zhang
🏛️ ShanghaiTech University | Hangzhou Dianzi University

This work addresses the problem of unsupervised extraction of semantically consistent action primitives from continuous industrial video streams to support Vision-Language-Action (VLA) model pretraining. Methodologically, we propose an end-to-end automated framework featuring a lightweight motion tokenizer to encode dynamic behaviors, introduce a novel unsupervised metric—“latent action energy”—for action segment boundary detection, and leverage vision-language models for semantic clustering and consistency evaluation of action primitives. Key contributions include: (1) the first fully automated pipeline converting unstructured industrial videos into VLA pretraining data; (2) the proposal of “latent action energy,” an interpretable and scalable measure of action saliency; and (3) empirical validation on public benchmarks and a newly curated motor assembly dataset, demonstrating effective segmentation and high semantic consistency of generated action primitives—establishing a scalable data foundation for embodied AI in manufacturing.

Automating VLA pre-training data extraction from unlabeled human demonstrationsEnabling scalable embodied AI integration in manufacturing settingsSegmenting continuous industrial videos into coherent action primitives

Towards Generalizing Temporal Action Segmentation to Unseen Views

Apr 03, 2025
EB
Emad Bahrami
🏛️ University of Bonn | Lamarr Institute for Machine Learning and Artificial Intelligence | Toyota Motor Europe

This work addresses the poor generalization of temporal action segmentation models to unseen camera viewpoints (e.g., third-person frontal → lateral, exocentric → egocentric). To systematically evaluate cross-view generalization, we introduce the first dedicated cross-view action segmentation benchmark protocol. We propose a multi-granularity shared representation framework that jointly models sequences and segments, enforcing cross-view semantic consistency via sequence alignment loss and contrastive action-level loss. Further, we incorporate cross-view consistency regularization and hierarchical temporal modeling to decouple and align video representations with action semantics. Evaluated on Assembly101, IKEA ASM, and EgoExoLearn, our method achieves +12.8% F1@50 improvement on unseen exocentric views and +54.0% on unseen egocentric views—marking substantial progress toward viewpoint-agnostic action segmentation.

Addressing view changes from top-frontal to side or exocentric to egocentricGeneralizing temporal action segmentation to unseen camera viewsImproving cross-view consistency in video and action representations

What Do Latent Action Models Actually Learn?

May 27, 2025
CZ
Chuheng Zhang
🏛️ Microsoft Research | Tsinghua University | Independent Researcher

This paper addresses the fundamental question of whether latent action models (LAMs) genuinely learn action-driven inter-frame dynamics or merely capture exogenous noise. Method: We develop an analytically tractable linear system model to theoretically characterize the learning mechanism of LAMs, uncovering their intrinsic relationship with principal component analysis (PCA) and rigorously analyzing how structural coupling among observations, actions, and noise governs model performance. Leveraging controllability theory, we derive principled guidelines for designing data generation strategies. Contribution/Results: These guidelines inform video data augmentation, noise denoising, and auxiliary action prediction. Numerical simulations demonstrate that our strategy significantly enhances learning of action-relevant features, thereby advancing the interpretability and reliability of unsupervised action representation learning.

Analyzing what latent action models actually learn from videosDetermining if latents capture actions or irrelevant noiseProviding insights on data structure influencing LAM learning

In unsupervised temporal action segmentation, existing methods rely solely on frame-level features while neglecting segment-level semantic modeling. To address this, we propose a joint frame-segment modeling framework: a Transformer encoder-decoder architecture jointly predicts frame-wise action labels and generates video action transcripts; a frame-to-segment permutation-aware alignment module is introduced, and—novelly—temporal optimal transport (OT) is employed to construct segment-level pseudo-labels for end-to-end unsupervised training. Our core contributions are: (1) the first segment-level semantic-guided paradigm for unsupervised action segmentation; and (2) an OT-based frame-segment alignment and pseudo-label generation mechanism. Extensive experiments demonstrate state-of-the-art performance across four major benchmarks—50 Salads, YouTube Instructions, Breakfast, and Desktop Assembly—outperforming all prior unsupervised approaches.

Aligns frame-level features with segment-level features for permutation-aware resultsIntroduces pseudo labels via temporal optimal transport for unsupervised trainingUnsupervised temporal activity segmentation using frame and segment cues

Latest Papers

What's happening recently
View more

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of latent action supervision from two perspectives: (i) regularizing the trajectory via image-based latent actions, and (ii) unifying the target space with action-based latent actions. Under a unified VLA baseline, we instantiate and compare four representative integration strategies. Our results reveal a formulation-task correspondence: image-based latent actions benefit long-horizon reasoning and scene-level generalization, whereas action-based latent actions excel at complex motor coordination. Furthermore, we find that directly supervising the VLM with discrete latent action tokens yields the most effective performance. Finally, our experiments offer initial insights into the benefits of latent action supervision in mixed-data, suggesting a promising direction for VLA training. Code is available at https://github.com/RUCKBReasoning/From_Pixels_to_Tokens.

action supervisionheterogeneous datasetsintermediate representation

Existing action tokenization methods struggle to simultaneously achieve compact sequence length, structured representation, and compatibility with downstream policies. This work proposes Ordered Action Tokenization (OAT), the first approach to jointly attain high compression ratio, full invertibility, and an ordered discrete token space. OAT integrates a register-augmented Transformer, Finite Scalar Quantization (FSQ), and a ranking-aware training mechanism to compress continuous action chunks into progressively decodable, ordered token sequences. This design enables flexible trade-offs during inference between computational cost and action fidelity. Evaluated across more than 60 simulated and real-world tasks, OAT consistently enhances both performance and inference flexibility across diverse policy architectures.

action tokenizationdiscrete tokenspolicy compatibility

This work addresses the challenges of temporal action segmentation—particularly abrupt action transitions, ambiguous boundaries, and high annotation costs—especially under low-resource or novel-domain settings. To this end, the authors propose a lightweight constraint-aware decoding framework that requires neither retraining nor increased model complexity during inference. By integrating structural priors derived from annotated data, including transition confidence, action boundary distributions, and class duration statistics, the framework employs an enhanced Viterbi algorithm to efficiently rectify structured prediction errors. Compatible with both fully and semi-supervised models, the approach significantly improves segmentation accuracy while maintaining high computational efficiency.

Action VariabilityAmbiguous BoundariesAnnotation Costs

Current text-to-motion generation methods struggle to precisely control the timing of motion strokes—such as punches—often resulting in blurred or merged actions. This work proposes explicitly modeling each stroke via Action Units (AUs), which encode the involved body parts, action category, temporal window, and precise impact moment. To achieve this without retraining, the authors introduce a lightweight gated adapter and a dual-stream injection mechanism that incorporates frame-level detector gradients into a frozen backbone model at no additional training cost. This approach is the first to decompose motion into explicit, type- and timing-constrained conditional signals. Evaluated on the StrokeBench benchmark, it substantially improves temporal accuracy for individual strokes while maintaining or even surpassing the motion quality of existing methods, further demonstrating that keyframes can serve as an effective controllable dimension.

action unitsmotion generationstroke timing

Hot Scholars

YD

Yilun Du

Harvard University
Artificial IntelligenceMachine LearningRoboticsComputer Vision
MD

Mingyu Ding

Assistant Professor, UNC Chapel Hill
RoboticsEmbodied AIComputer Vision
JZ

Jiazhao Zhang

Peking University
Embodied AINavigation3D Vision
CZ

Ce Zhang

University of North Carolina at Chapel Hill
Computer VisionVideo UnderstandingMultimodal LearningRobotics
PA

Pieter Abbeel

UC Berkeley | Covariant
RoboticsMachine LearningAI