Score
Designs algorithms and models that align representations, correspondences, or distributions across time (e.g., pre/post) and across hierarchical temporal scales, producing temporally matched features or trajectories; this includes methods for temporal domain alignment, temporal matching, and temporal distribution alignment. Builds coarse-to-fine flow decoding and multi-scale alignment mechanisms and associated regularizers that enforce trajectory consistency, spatial/temporal smoothness, and coherence across scales.
In collaborative perception, hardware heterogeneity induces feature-domain shift, while communication latency causes temporal misalignment—jointly degrading feature quality and accumulating cross-node errors. To address these challenges at the feature-level fusion stage, we propose a systematic alignment framework: (1) a consistency-preserving domain alignment module mitigates inter-device feature distribution discrepancies; (2) a progressive temporal alignment module corrects dynamic timing offsets via multi-scale motion modeling and two-stage compensation; and (3) an observability-constrained discriminator and instance-aware hierarchical aggregation strategy enhance semantic consistency. Evaluated on three benchmark datasets, our method achieves state-of-the-art performance and demonstrates significantly improved robustness under high communication latency and pose estimation errors.
This paper addresses the semantic alignment challenge in non-visual time-series-to-high-fidelity-image generation. We propose TimeArtist, a novel framework that introduces two key innovations: (1) a “warm-up–alignment” two-stage training paradigm—first performing self-supervised pretraining via dual autoencoders with a shared vector quantizer, then freezing encoders and introducing a representation-layer projection to achieve time-vision cross-modal semantic alignment; and (2) a transferable spatial prior unified architecture that maps temporal dynamics into controllable visual styles. Experiments demonstrate state-of-the-art performance on image quality metrics (e.g., FID, LPIPS) and achieve SOTA results on downstream zero-shot time-series forecasting tasks. These findings validate the effectiveness of both cross-modal semantic alignment and generative representation transfer.
Time series exhibit nonlinear temporal misalignments, hindering accurate alignment and averaging—key bottlenecks for conventional analysis methods. To address this, we propose the Diffeomorphic Temporal Alignment Network (DTAN), a framework enabling unsupervised or weakly supervised joint alignment and averaging of variable-length time series collections. Our contributions are threefold: (1) We introduce Inverse-Consistent Averaging Error (ICAE) regularization—a novel, registration-free constraint ensuring diffeomorphic consistency without explicit correspondence supervision; (2) We design Multi-Task DTAN (MT-DTAN), unifying alignment and classification in an end-to-end trainable architecture; (3) We extend Principal Component Analysis (PCA) to unaligned time series for the first time. DTAN leverages input-dependent diffeomorphic deformation prediction, warp regularization, and UCR-compatible backbone designs. Evaluated on all 128 UCR time series classification benchmarks, DTAN consistently outperforms state-of-the-art alignment and averaging methods, achieving significant gains in alignment accuracy and downstream task performance.
Diffusion models often deviate from the underlying data manifold under arbitrary guidance, degrading sample fidelity. To address this, we propose Temporal Alignment Guidance (TAG), a lightweight, plug-and-play mechanism that dynamically estimates the temporal deviation between the current sampling step and the target data manifold via an auxiliary time predictor. TAG then applies gradient-based correction to realign the sample trajectory at each step, enforcing fine-grained manifold constraints without modifying the diffusion model or requiring additional training. Empirically, TAG significantly improves generation quality across text-to-image synthesis and class-conditional generation: samples remain consistently closer to the true data manifold throughout sampling, exhibit enhanced robustness to guidance scale, achieve up to a 12.3% reduction in FID, and preserve both sampling efficiency and diversity.
To address the challenges of complex feature fusion and poor scalability in unsupervised multimodal time-series representation learning, this paper proposes a spectral-graph-theoretic framework for implicit cross-modal alignment. Methodologically, it abandons explicit multimodal feature fusion and instead adopts a single-encoder, multi-view architecture: diverse time-series views—such as frequency-domain, image-based, and symbolic representations—are generated via modality-agnostic temporal transformations, and spectral-graph-guided contrastive loss enables unsupervised cross-modal alignment. This lightweight design implicitly captures latent inter-modal dependencies, strengthening inductive bias and generalization capacity. Extensive experiments across multiple domains demonstrate state-of-the-art performance on downstream tasks—including classification, forecasting, and anomaly detection—surpassing existing unsupervised methods by an average accuracy gain of 3.2%–7.8%.
This work addresses the limitations of existing climate data super-resolution methods, which often focus solely on single-frame spatial information while neglecting temporal dependencies and exhibiting high sensitivity to noise, thereby compromising reconstruction accuracy. To overcome these issues, we propose a temporally enhanced bidirectional alignment framework that, for the first time in climate super-resolution, incorporates a bidirectional temporal alignment mechanism. By employing paired latent-space mappings, our approach unifies spatiotemporal representations and effectively suppresses noise, enabling the exploitation of implicit temporal correlations. Departing from conventional strategies such as optical flow—ill-suited for climate data—our method leverages deep networks for end-to-end optimization, integrating forward–backward alignment with a super-resolution module. Extensive experiments on large-scale real-world climate datasets demonstrate that the proposed framework significantly improves both fine-detail recovery and spatiotemporal consistency.
Existing text-to-motion generation methods are often confined to a single temporal scale, struggling to simultaneously achieve semantic alignment and temporal coherence. This work proposes a hierarchical flow-matching framework that progressively generates motion across multiple time scales: high-level layers capture semantic content and coarse structure, while low-level layers refine fine-grained temporal details. A cross-scale transfer mechanism ensures both continuity and noise consistency throughout the hierarchy. The approach integrates a text-motion diffusion Transformer, a topology-aware motion VAE, joint-aware positional encoding, and explicit skeletal topology modeling, collectively enhancing generation quality. It achieves state-of-the-art performance on the HumanML3D and KIT-ML benchmarks, with ablation studies confirming the effectiveness of the hierarchical design and individual components.
Current evaluation methods for audio-visual talking head generation rely on frame-level metrics that assume strict temporal alignment between generated and reference videos, rendering them sensitive to natural variations in speech rate, rhythm, and stylistic expression, and thereby introducing assessment bias. This work reframes evaluation as a sequence alignment problem and introduces Soft Dynamic Time Warping (Soft DTW) to align feature trajectories temporally, enhancing robustness to timing offsets while preserving sequential constraints. The proposed unified sequence-level evaluation framework subsumes frame-level metrics as a special case of rigid alignment, enabling compatibility with existing perceptual, identity, and synchronization encoders without modification. Large-scale experiments across 20 methods and 7 datasets demonstrate that the approach yields more stable evaluations with higher cross-dataset consistency, clearly disentangling trade-offs between synchronization and realism, as well as expressiveness and stability.
Existing inference-time alignment methods operate along a single control axis, struggling to model the joint dependency between conditioning variables and latent states and exhibiting limited generalization. This work proposes PG-MAP, a framework that, without requiring any training, formulates alignment as a joint maximum a posteriori (MAP) or proximal energy optimization over both conditional variables and latent states. It introduces a forward-consistency coupling mechanism to enable cross-modal collaborative updates. PG-MAP unifies support for both diffusion and flow-matching models by integrating trajectory-level Gibbs-MAP sampling, proximal optimization, and frozen preference-reward guidance, adapting flexibly to diverse generative transport dynamics. Experiments demonstrate that PG-MAP significantly improves PickScore and Aesthetic scores on SD1.5 and SDXL, achieving 91.9% PickScore and 75.7% HPS win rate on SD3.5-medium, with human evaluations consistently outperforming strong baselines.