cross-temporal alignment

Designs algorithms and models that align representations, correspondences, or distributions across time (e.g., pre/post) and across hierarchical temporal scales, producing temporally matched features or trajectories; this includes methods for temporal domain alignment, temporal matching, and temporal distribution alignment. Builds coarse-to-fine flow decoding and multi-scale alignment mechanisms and associated regularizers that enforce trajectory consistency, spatial/temporal smoothness, and coherence across scales.

cross-temporalalignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception

Jul 24, 2025
CT
Chengchang Tian
🏛️ Southeast University | Washington State University

In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.

Address domain gaps from hardware diversity and deployment conditionsEnhance semantic feature quality for collaborative perception fusionMitigate temporal misalignment caused by transmission delays

This paper addresses the semantic alignment challenge in non-visual time-series-to-high-fidelity-image generation. We propose TimeArtist, a novel framework that introduces two key innovations: (1) a “warm-up–alignment” two-stage training paradigm—first performing self-supervised pretraining via dual autoencoders with a shared vector quantizer, then freezing encoders and introducing a representation-layer projection to achieve time-vision cross-modal semantic alignment; and (2) a transferable spatial prior unified architecture that maps temporal dynamics into controllable visual styles. Experiments demonstrate state-of-the-art performance on image quality metrics (e.g., FID, LPIPS) and achieve SOTA results on downstream zero-shot time-series forecasting tasks. These findings validate the effectiveness of both cross-modal semantic alignment and generative representation transfer.

Enabling high-fidelity image generation directly from temporal dataEstablishing semantic alignment between time series and visual conceptsTransferring spatial priors from vision models to temporal tasks

Time series exhibit nonlinear temporal misalignments, hindering accurate alignment and averaging—key bottlenecks for conventional analysis methods. To address this, we propose the Diffeomorphic Temporal Alignment Network (DTAN), a framework enabling unsupervised or weakly supervised joint alignment and averaging of variable-length time series collections. Our contributions are threefold: (1) We introduce Inverse-Consistent Averaging Error (ICAE) regularization—a novel, registration-free constraint ensuring diffeomorphic consistency without explicit correspondence supervision; (2) We design Multi-Task DTAN (MT-DTAN), unifying alignment and classification in an end-to-end trainable architecture; (3) We extend Principal Component Analysis (PCA) to unaligned time series for the first time. DTAN leverages input-dependent diffeomorphic deformation prediction, warp regularization, and UCR-compatible backbone designs. Evaluated on all 128 UCR time series classification benchmarks, DTAN consistently outperforms state-of-the-art alignment and averaging methods, achieving significant gains in alignment accuracy and downstream task performance.

Addressing nonlinear temporal misalignment in time-seriesEnabling PCA for misaligned time-series dataJoint alignment and averaging of time-series ensembles

Diffusion models often deviate from the underlying data manifold under arbitrary guidance, degrading sample fidelity. To address this, we propose Temporal Alignment Guidance (TAG), a lightweight, plug-and-play mechanism that dynamically estimates the temporal deviation between the current sampling step and the target data manifold via an auxiliary time predictor. TAG then applies gradient-based correction to realign the sample trajectory at each step, enforcing fine-grained manifold constraints without modifying the diffusion model or requiring additional training. Empirically, TAG significantly improves generation quality across text-to-image synthesis and class-conditional generation: samples remain consistently closer to the true data manifold throughout sampling, exhibit enhanced robustness to guidance scale, achieve up to a 12.3% reduction in FID, and preserve both sampling efficiency and diversity.

Addressing off-manifold errors in diffusion model generationCorrecting sample fidelity degradation from arbitrary guidance mechanismsImproving generation quality through temporal alignment at each timestep

Unsupervised Multi-modal Feature Alignment for Time Series Representation Learning

Dec 09, 2023
CL
Chen Liang
🏛️ Harbin Institute of Technology

To address the challenges of complex feature fusion and poor scalability in unsupervised multimodal time-series representation learning, this paper proposes a spectral-graph-theoretic framework for implicit cross-modal alignment. Methodologically, it abandons explicit multimodal feature fusion and instead adopts a single-encoder, multi-view architecture: diverse time-series views—such as frequency-domain, image-based, and symbolic representations—are generated via modality-agnostic temporal transformations, and spectral-graph-guided contrastive loss enables unsupervised cross-modal alignment. This lightweight design implicitly captures latent inter-modal dependencies, strengthening inductive bias and generalization capacity. Extensive experiments across multiple domains demonstrate state-of-the-art performance on downstream tasks—including classification, forecasting, and anomaly detection—surpassing existing unsupervised methods by an average accuracy gain of 3.2%–7.8%.

Aligning multi-modal time series features for representation learningImproving inductive bias in unsupervised time series encodingReducing feature fusion complexity to enhance scalability

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing climate data super-resolution methods, which often focus solely on single-frame spatial information while neglecting temporal dependencies and exhibiting high sensitivity to noise, thereby compromising reconstruction accuracy. To overcome these issues, we propose a temporally enhanced bidirectional alignment framework that, for the first time in climate super-resolution, incorporates a bidirectional temporal alignment mechanism. By employing paired latent-space mappings, our approach unifies spatiotemporal representations and effectively suppresses noise, enabling the exploitation of implicit temporal correlations. Departing from conventional strategies such as optical flow—ill-suited for climate data—our method leverages deep networks for end-to-end optimization, integrating forward–backward alignment with a super-resolution module. Extensive experiments on large-scale real-world climate datasets demonstrate that the proposed framework significantly improves both fine-detail recovery and spatiotemporal consistency.

climate data super-resolutionspatial resolutionstochastic noise

Existing text-to-motion generation methods are often confined to a single temporal scale, struggling to simultaneously achieve semantic alignment and temporal coherence. This work proposes a hierarchical flow-matching framework that progressively generates motion across multiple time scales: high-level layers capture semantic content and coarse structure, while low-level layers refine fine-grained temporal details. A cross-scale transfer mechanism ensures both continuity and noise consistency throughout the hierarchy. The approach integrates a text-motion diffusion Transformer, a topology-aware motion VAE, joint-aware positional encoding, and explicit skeletal topology modeling, collectively enhancing generation quality. It achieves state-of-the-art performance on the HumanML3D and KIT-ML benchmarks, with ablation studies confirming the effectiveness of the hierarchical design and individual components.

3D human motionhierarchical motionsemantic alignment

Current evaluation methods for audio-visual talking head generation rely on frame-level metrics that assume strict temporal alignment between generated and reference videos, rendering them sensitive to natural variations in speech rate, rhythm, and stylistic expression, and thereby introducing assessment bias. This work reframes evaluation as a sequence alignment problem and introduces Soft Dynamic Time Warping (Soft DTW) to align feature trajectories temporally, enhancing robustness to timing offsets while preserving sequential constraints. The proposed unified sequence-level evaluation framework subsumes frame-level metrics as a special case of rigid alignment, enabling compatibility with existing perceptual, identity, and synchronization encoders without modification. Large-scale experiments across 20 methods and 7 datasets demonstrate that the approach yields more stable evaluations with higher cross-dataset consistency, clearly disentangling trade-offs between synchronization and realism, as well as expressiveness and stability.

audio-driven talking headevaluation protocolsequence-level evaluation

Existing inference-time alignment methods operate along a single control axis, struggling to model the joint dependency between conditioning variables and latent states and exhibiting limited generalization. This work proposes PG-MAP, a framework that, without requiring any training, formulates alignment as a joint maximum a posteriori (MAP) or proximal energy optimization over both conditional variables and latent states. It introduces a forward-consistency coupling mechanism to enable cross-modal collaborative updates. PG-MAP unifies support for both diffusion and flow-matching models by integrating trajectory-level Gibbs-MAP sampling, proximal optimization, and frozen preference-reward guidance, adapting flexibly to diverse generative transport dynamics. Experiments demonstrate that PG-MAP significantly improves PickScore and Aesthetic scores on SD1.5 and SDXL, achieving 91.9% PickScore and 75.7% HPS win rate on SD3.5-medium, with human evaluations consistently outperforming strong baselines.

conditioning and latent variablesgenerative transportsinference-time alignment

Hot Scholars

YM

Yuki Mitsufuji

Distinguished Engineer, Sony
Machine LearningAudioSource SeparationMusic Technology
TW

Tai Wang

Shanghai AI Laboratory
Computer Vision3D VisionEmbodied AIDeep Learning
QW

Qilong Wang

Tianjin University
Deep LearningComputer Vision
QH

Qinghua Hu

Professor of Computer Science, Tianjin University
Machine learningData Mining
MS

Muyi Sun

School of AI, BUPT (<< NLPR CASIA << BUPT)
Multi-Modality LearningComputer VisionBiometricsMedical Image Analysis