direct timestep embedding

Designs and implements direct mappings from discrete timesteps into a model's internal embedding space—using point-wise, raw, or sinusoidal encodings—that preserve exact index-level access, magnitude/scale, and temporal trend without patching or padding. These embeddings are built to schedule or steer generation toward task-specific manifolds (including per-task timestep assignments) and can be applied in a parameter-free way to influence model behavior.

directtimestepembedding

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the high computational cost and limited temporal modeling capability of traditional implicit neural representations (INRs) when handling time-varying volumetric data, which typically rely on dense spatiotemporal coordinate sampling. The authors propose reformulating time-varying volumes as collections of time series indexed by spatial locations and replacing point-wise scalar supervision with sequence-level supervision. To better capture heterogeneous temporal dynamics, they introduce a Mixture-of-Experts (MoE) architecture that adaptively allocates model capacity across different spatial regions. This approach substantially reduces training overhead while improving reconstruction quality, demonstrates compatibility with various existing INR frameworks, and outperforms current state-of-the-art methods across multiple evaluation metrics.

coordinate-wise supervisionimplicit neural representationsspatiotemporal sampling

This work addresses the unclear nature of how neural networks internally represent the underlying geometric structure of complex dynamical systems. To resolve rotational and scale ambiguities in latent spaces, the authors propose an anchor-based, geometry-agnostic relative embedding method, establishing a reproducible framework for relative geometric analysis. Through systematic experiments across seven canonical dynamical systems using MLPs, RNNs, Transformers, and echo state networks, they find that MLPs and RNNs exhibit highly aligned internal representations, whereas Transformers and echo state networks achieve high predictive accuracy despite weaker representational alignment. This reveals that high prediction accuracy can coexist with low representational alignment, thereby uncovering a nuanced relationship between alignment and predictive performance.

dynamical systemslatent geometryneural forecasters

This study investigates how recurrent neural networks (RNNs) preserve the topological structure of invariant manifolds—such as those on tori or circles—in regular dynamical systems when trained for time-series prediction. By treating the invariant manifold of the input system as a driving signal, the work posits that the RNN’s hidden state encodes a finite history window rather than instantaneous inputs. A unified theoretical framework is developed by integrating generalized synchronization theory, differential embedding theorems, and contraction analysis. The analysis reveals that under regular dynamical driving, RNNs naturally satisfy smooth embedding conditions, circumventing the stringent requirements typical in chaotic systems. Furthermore, verifiable criteria are established that clarify the relationship between the dimensionality of the hidden state and the intrinsic dimension of the driving system, thereby elucidating the mechanism underlying topologically faithful representations.

contracting systemsrecurrent neural networksregular dynamics

This work addresses the computational redundancy in existing single-step diffusion models for multitask dense prediction, which typically rely on parameter-heavy adapters or learnable task tokens. The study is the first to reveal and exploit the fixed sinusoidal timestep embeddings inherent in diffusion models as endogenous task-conditioning signals, proposing a unified multitask learning paradigm that requires no additional parameters. Built upon pretrained diffusion models, the method leverages timestep embeddings for task guidance and incorporates manifold disentanglement to enable task-specific generation, compatible with both U-Net and DiT architectures. Experiments across ten datasets demonstrate that the approach achieves performance on par with state-of-the-art methods in monocular depth and surface normal estimation, confirming its effectiveness and broad applicability.

dense predictiondiffusion modelsmonocular vision

This work addresses the challenge of modeling irregularly and asynchronously observed time series data by proposing a continuous-time embedding method that operates without interpolation or imputation. The approach directly encodes observations as increments and constructs a continuous, injective embedding via the log-signature over intervals, enabling online computation while avoiding full path reconstruction. Built upon the Log-NCDE framework and the concept of rectangular control paths, the proposed embedding preserves the structural fidelity of the original data and maintains universality over compact subsets of the input space. Empirical evaluations demonstrate that the method achieves high accuracy, computational efficiency, and strong robustness to sparsity and asynchronicity across both synthetic dynamical systems and real-world temporal datasets.

asynchronous datacontinuous-time modelsfaithful embeddings

Latest Papers

What's happening recently
View more

This work challenges the conventional assumption that diffusion models inherently require explicit timestep embeddings, investigating their necessity in the denoising process. Through theoretical analysis and empirical validation, the study demonstrates for the first time that under certain conditions, both U-Net and Diffusion Transformer architectures can converge to a global optimum without explicit timestep conditioning, implicitly inferring the noise scale. Ablation studies and generative evaluations on CelebA and CIFAR-10 show that such timestep-agnostic models achieve competitive or superior performance compared to standard timestep-conditioned counterparts in terms of FID, precision, and recall, while preserving high structural fidelity.

denoising processdiffusion modelsnoise scales

This work addresses the challenge of achieving strict idempotence in generative models under repeated application, where output drift arises due to geometric inconsistencies between the data manifolds learned by the encoder and decoder. The study identifies this manifold misalignment as the key cause of idempotence failure—a previously unexamined issue—and introduces a novel training framework that explicitly aligns the geometric structures of both components. By enforcing the encoder’s projection and the decoder’s reconstruction to share a common underlying manifold during training, the proposed method substantially reduces idempotence error, yielding perfectly consistent outputs across repeated generations. Empirical results demonstrate significant improvements in identity preservation and information stability for image generation and editing tasks.

encoder-decoder mismatchfixed pointsgenerative models

This work addresses the challenge of analytically characterizing the dynamics of high-dimensional neural network training trajectories. To this end, it introduces—for the first time—the scalar embedding methodology from time-series analysis into the study of training dynamics, constructing low-dimensional representations that capture essential dynamical features. By defining a Lyapunov-like characteristic timescale, estimating Lyapunov exponents, performing trajectory sensitivity analyses, and examining inter-trajectory distance statistics, the study reveals the decorrelation scale and asymptotic behavior of the training process. Experimental results demonstrate that the embedded trajectories faithfully reproduce the dynamics of the original parameter space, and that the asymptotic distribution of trajectory distances consistently follows a skewed log-normal form across diverse settings, thereby validating both the efficacy and theoretical significance of the proposed approach.

loss landscapeLyapunov exponentscalar embedding

Existing time series generative models suffer from limited expressivity and poor adaptability to irregularly sampled observation grids. This work proposes G-SLiCEs, a continuous-time generative model based on Structured Linear Controlled differential equations (SLiCEs), which achieves high expressivity through continuous flow matching in path space. We establish, for the first time, that SLiCEs can approximate any continuous causal pushforward path law under the Wasserstein-∞ metric, thereby enabling universal time series generation and introducing maximal expressivity into continuous-time generative modeling. The method natively supports arbitrary observation time grids and significantly outperforms existing approaches in irregularly sampled settings, demonstrating superior performance in probabilistic forecasting and downstream tasks.

continuous-time modelsirregular gridspath-space generative modeling

Hot Scholars

LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
YM

Yue Ma

Bytedance
NLPDialogue SystemLLM
WL

Wenhan Luo

Associate Professor, HKUST
Creative AIGenerative ModelComputer VisionMachine Learning
PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
HY

Harry Yang

HKUST
computer visionmachine learning