sequence modeling

Modeling temporal or sequential data (e.g., forecasting, sequence-to-sequence, encoding/segmentation) to predict future outputs, generate realistic trajectories or action sequences, and train agents on simulator-generated trajectories for planning.

sequencemodeling

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.

AI trainingdata qualitydata volume

Human Motion Prediction, Reconstruction, and Generation

Feb 21, 2025
CG
Canxuan Gang
🏛️ AI Geeks

This paper presents a systematic review of recent advances in human motion prediction, reconstruction, and generation. Addressing key challenges—including instability in long-horizon prediction, limited reconstruction accuracy, and insufficient physical plausibility and diversity in motion generation—we propose a unified “prediction–reconstruction–generation” co-evolutionary framework. Our method integrates diffusion models with physics-informed dynamical constraints in the loss function to enhance motion realism and biomechanical consistency. Furthermore, we introduce multimodal alignment and fine-grained contextual modeling to improve text-to-motion generation and human-object interaction synthesis. Extensive experiments demonstrate significant improvements over state-of-the-art methods: a 23% reduction in average prediction error for long-horizon motion forecasting, an 18% decrease in MPJPE for 3D pose reconstruction, and a 31% reduction in FID score for motion generation. The framework supports applications in digital avatars, embodied AI, and real-time AR interaction.

Forecasting future human poses from historical dataRecovering 3D human movements from visual inputsSynthesizing realistic motions from textual descriptions

PPT: Pre-Training with Pseudo-Labeled Trajectories for Motion Forecasting

Dec 09, 2024
YX
Yihong Xu
🏛️ Valeo.ai | Sorbonne Université

To address the poor generalizability of motion prediction in autonomous driving—stemming from reliance on costly, non-scalable, and non-reproducible manually annotated data—this paper proposes a novel self-supervised pretraining paradigm leveraging pseudo-labeled trajectories. Unlike prior approaches that treat automatically generated detection-and-tracking trajectories as noisy artifacts to be filtered out, our core innovation lies in deliberately harnessing these inherently diverse, uncurated pseudo-trajectories as valuable supervisory signals, thereby eliminating dependence on clean, single-label ground truth. The method jointly integrates 3D object detection, multi-object tracking, and contrastive/reconstructive self-supervised learning, enabling efficient lightweight fine-tuning. Evaluated on standard benchmarks, it achieves state-of-the-art performance, with substantial improvements in low-label-data regimes, cross-domain transfer, and multi-class end-to-end motion forecasting—while demonstrating strong robustness and practical applicability.

Addressing domain gaps and improving generalization in dynamic scene predictionLeveraging noisy, diverse pseudo-labels for robust representation learningReducing reliance on costly manual trajectory annotations for motion forecasting

AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials

Dec 12, 2024
YX
Yiheng Xu
🏛️ University of Hong Kong | Salesforce Research

High-quality multi-step GUI interaction trajectories for training GUI agents are scarce and prohibitively expensive to annotate manually. Method: This paper proposes a web-tutorial-based automated trajectory synthesis framework: (1) crawling open-source online tutorials and parsing them into structured, multi-step task specifications; (2) orchestrating a vision-language model (VLM) agent to execute tasks and generate trajectories in real GUI environments; and (3) employing a VLM-based evaluator for end-to-end automatic trajectory validation. We introduce “guided replay”—the first paradigm enabling fully automated conversion of unstructured textual tutorials into executable, verifiable GUI trajectories without human annotation. Contribution/Results: Experiments demonstrate that synthesized trajectories significantly improve agent performance in GUI element localization and multi-step planning, outperforming prior methods across multiple benchmarks. Moreover, the per-trajectory data cost is reduced by over an order of magnitude, enabling scalable, low-cost GUI agent training.

Enhances GUI agent performance with multimodal data.Generates web agent trajectories using web tutorials.Reduces data collection costs without human annotation.

This study addresses the limitations of traditional physics simulators in robotics—such as restricted expressiveness due to simplifying assumptions, high data costs, and difficulties in modeling complex physical interactions—by systematically reviewing video generation models as embodied world models. Integrating high-fidelity, multimodal-conditioned video synthesis with imitation learning, reinforcement learning, and visual planning frameworks, this work provides the first comprehensive analysis of their potential and limitations in tasks including action prediction, dynamics modeling, and policy evaluation. The review highlights breakthroughs in high-fidelity modeling of physical interactions while identifying key challenges in instruction following, physical consistency, and safety. These insights lay a theoretical foundation and outline future directions for replacing conventional simulators and enabling deployment in safety-critical scenarios.

hallucinationphysics violationrobotics

Latest Papers

What's happening recently
View more

Traditional time series forecasting is constrained by static, one-shot, model-centric paradigms that struggle to support dynamic reasoning and continual learning. This work proposes Agent-based Time Series Forecasting (ATSF), introducing an agent framework into the field for the first time and reconceptualizing prediction as a multi-round workflow encompassing perception, planning, action, reflection, and memory. ATSF enables tool invocation, feedback integration, and experiential evolution. Through three implementation pathways—workflow-based design, agent reinforcement learning, and hybrid agent architectures—ATSF establishes a novel forecasting paradigm that is interactive, evolvable, and supports iterative refinement. This study not only opens an agent-oriented research direction for time series forecasting but also systematically articulates its technical pathways, key challenges, and future opportunities.

adaptive forecastingagentic forecastingmodel-centric prediction

Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models

Oct 04, 2025
HW
Hao Wu
🏛️ Tsinghua University | OpenAI | Tencent Hunyuan | SLAI | CUHK | Tencent Jarvis Lab | University of Wisconsin | Nanyang Technological University

Addressing the dual challenges of inherent stochasticity and non-differentiable evaluation metrics in physical spatiotemporal forecasting, this paper proposes a novel model-based reinforcement learning paradigm that reformulates prediction as sequential planning. Methodologically, we construct a generative world model to simulate high-fidelity, diverse future states and employ domain-specific non-differentiable metrics—such as extreme-event hit rate—as sparse reward signals. We design a beam-search–guided, reward-driven imagination mechanism and introduce an iterative pseudo-labeling self-training strategy. Crucially, our framework enables end-to-end optimization of non-differentiable objectives without gradient approximation. Experiments demonstrate substantial reductions in overall prediction error alongside marked improvements in long-tail event detection. This work establishes a new pathway toward interpretable and robust forecasting for complex physical systems.

Addresses spatiotemporal forecasting challenges with stochasticity and non-differentiable metricsOptimizes forecasting through planning algorithms using non-differentiable reward signalsProposes model-based reinforcement learning with generative world simulation

What Happens Next? Anticipating Future Motion by Generating Point Trajectories

Sep 25, 2025
GB
Gabrijel Boduljak
🏛️ University of Oxford

This paper addresses the problem of single-image motion trajectory forecasting: predicting dense future trajectories of scene objects directly from a static image, without requiring auxiliary physical parameters such as velocity or force. We propose a conditional generative model based on a *trajectory grid*, which bypasses redundant pixel-level modeling typical in video generation and instead performs end-to-end synthesis of structured motion fields. Our approach explicitly captures global dynamic patterns and motion uncertainty. Built upon modern video generation architectures, the model is trained jointly on synthetic physics-based simulations and real-world scenes. Experimental results demonstrate significant improvements over state-of-the-art regression- and generation-based methods on both simulated and real-world physical benchmarks. Furthermore, we validate the practical utility and generalization capability of our method in downstream robotic navigation tasks.

Anticipating object motion from single imagesGenerating dense trajectory grids instead of pixelsOvercoming limitations of video generators in motion forecasting

This work investigates whether flow matching in temporal generation learns a universal dynamical structure or merely reproduces historical trajectories. By analyzing the empirical flow matching objective under Gaussian conditional paths, we derive—for the first time—a closed-form expression for its optimal velocity field, revealing it to be a similarity-weighted mixture of historical instantaneous velocities. This formulation constitutes a nonparametric, memory-augmented continuous-time dynamical system. Building on this insight, we propose a training-free closed-form sampler that directly generates high-quality probabilistic forecasts from historical transitions. Evaluated on nonlinear dynamical system benchmarks, our method substantially improves sampling efficiency and numerical stability while offering an explicit, interpretable mechanism for data-dependent dynamics.

dynamical structureFlow Matchingsequential data

This work addresses the absence of a unified closed-loop learning environment that enables agents to continuously learn from real-world events and forecast future outcomes. To bridge this gap, we propose FutureWorld—the first framework that formulates real-time future prediction as a reinforcement learning environment. By integrating a closed-loop mechanism of prediction, outcome realization, and parameter update, FutureWorld effectively prevents answer leakage and supports continual learning. Built upon open-source large language models and grounded in real-world event feedback, the framework establishes a daily-updated benchmark for training and evaluation. Experimental results over consecutive days demonstrate the efficacy of our approach, setting a new state-of-the-art baseline and significantly advancing agents’ predictive capabilities.

agent traininglive future predictionpredictive agents

Hot Scholars

SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing