Score
Design and implement 1D convolutional architectures that model temporal sequences by specifying causal convolutional filters, dilation patterns, and receptive-field size to control temporal context. Build and analyze variants with temporal pooling and upsampling or encoder–decoder structures, and apply model-compression and reconstruction techniques to reduce parameters and recover signals from partial observations.
Deep convolutional neural networks (CNNs) face persistent challenges in balancing representational capacity with deployment efficiency across diverse domains—including vision, language, healthcare, and speech—especially under resource constraints and data-limited regimes. Method: This work systematically surveys CNN architectural evolution from 2015 to 2025 and proposes a novel seven-dimensional unified taxonomy (covering spatial modeling, multi-path design, dimensional expansion, attention integration, etc.), alongside a synergistic optimization framework integrating sparse convolutions, depthwise separable convolutions, and attention mechanisms. It further incorporates Fourier-based preprocessing, low-precision computation, and weight compression for lightweight deployment. Contribution/Results: We comprehensively characterize the applicability boundaries of over 100 CNN variants, quantify the trade-off between computational efficiency and representation fidelity, formalize adaptation strategies for few-shot, weakly supervised, and federated learning settings, and prospectively identify emerging directions—including CNN-Transformer hybrids, vision-language joint modeling, and generative CNNs—establishing reusable design paradigms for edge deployment and cross-domain generalization.
Modeling temporal structure in infinite-dimensional dynamical systems—such as solution maps of stochastic differential equations (SDEs)—remains challenging due to the lack of expressive, causally consistent architectures operating on Banach spaces. Method: This paper introduces the *causal neural operator* framework, the first to achieve uniform approximation of Hölder-continuous or smooth trace-class causal operators on Banach spaces. It integrates a sequence architecture grounded in infinite-dimensional linear metric spaces, explicit causal constraints, functional approximation theory, and stochastic operator theory. Contributions/Results: We establish a quantitative trade-off between latent state dimension and approximation accuracy, substantially improving error bounds for RNNs in dynamical system approximation. We prove uniform approximability over arbitrary finite time horizons and compact sets, with error rates superior to classical feedforward-network-based results. By explicitly encoding temporal geometry and causality, our framework overcomes a key limitation of conventional neural operators—which neglect sequential structure—and establishes a new paradigm for causality-driven modeling of infinite-dimensional dynamics.
Causal convolutional neural networks (CNNs) lack interpretability in multimodal frequency–time series modeling. Method: We reveal that trained causal CNNs are mathematically equivalent to finite impulse response (FIR) filters and propose an analytical simplification leveraging the associativity of convolution to rigorously reduce deep causal CNNs to a single-layer equivalent FIR filter. Quasi-linear activation functions and least-squares optimization enable explicit frequency-domain feature extraction and interpretable mapping of filter parameters. Results: Evaluated on simulated beam dynamics and real bridge vibration data, our method accurately models sparse-spectrum physical system dynamics. Crucially, it establishes, for the first time, an analytical correspondence between causal CNN weights and the system’s frequency response function—thereby significantly enhancing the physical interpretability and spectral awareness of deep models in dynamic system identification.
Traditional ST-GCNs employ单一 temporal modules—either CNNs or LSTMs—leading to insufficient capture of dynamic spatiotemporal patterns. To address this, we propose a plug-and-play hybrid temporal module that, for the first time, synergistically integrates CNNs and LSTMs within a unified co-temporal block. This design jointly models local temporal features and long-range dependencies. Through theoretical analysis and cross-dataset ablation studies, we systematically characterize the intrinsic relationship between temporal module architecture and representational capacity. Evaluated on standard spatiotemporal graph benchmarks—including NTU-RGB+D and PeMSD7—our method achieves significant improvements in prediction accuracy and cross-domain generalization. It consistently outperforms pure-CNN and pure-LSTM baselines in temporal representation learning. The proposed module establishes a reusable, principled design paradigm for temporal modeling in ST-GCNs, advancing both expressiveness and architectural flexibility.
This work addresses the lack of systematic evaluation of sequence models’ ability to capture diverse temporal dependencies—such as short- and long-range, decaying, and oscillatory patterns. We propose the first synthetic benchmark framework based on controllable, parameterized memory functions. By explicitly designing memory kernel functions, our framework generates synthetic tasks with continuous-time complexity, enabling fine-grained, interpretable, and theoretically grounded analysis of model memory characteristics. We evaluate mainstream architectures—including RNNs, Transformers, and State Space Models (SSMs)—under a unified benchmark across multiple dimensions. Our experiments not only validate existing theoretical predictions but also uncover, for the first time, implicit architectural preferences for specific memory patterns and their precise failure boundaries. The results provide reproducible, interpretable, and quantitative guidance for selecting appropriate sequence modeling paradigms.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
This work addresses the ambiguity regarding whether existing single-stage video object detectors genuinely leverage temporal context, as standard evaluation metrics often fail to reveal their actual reliance on temporal information. To this end, we propose TemporalLens, a diagnostic framework that quantifies a model’s temporal dependency through controlled perturbations—including temporal shuffling, structured occlusion, and redundancy injection. Furthermore, we design YOLO-3D based on YOLOv8, explicitly preserving the temporal dimension within the backbone to enhance genuine temporal reasoning. Experiments demonstrate that TemporalLens effectively distinguishes between stacked 2D models and true temporal architectures, while YOLO-3D achieves an average mAP@50 improvement of 3.7 percentage points with 32-frame inputs, underscoring the critical role of temporal depth in performance gains.
Existing visual tokenization methods struggle to simultaneously support causal autoregressive modeling and preserve spatial structure, often leading to training instability or degraded generation quality. This work proposes CaTok, the first approach that integrates a one-dimensional causal token sequence with a MeanFlow decoder, enabling efficient and high-fidelity image generation through a temporal interval selection mechanism. Additionally, CaTok introduces REPA-A regularization to align features with those of vision foundation models. The method supports both single-step fast generation and multi-step high-quality sampling, achieving state-of-the-art performance on ImageNet reconstruction with only 0.75 FID, 22.53 PSNR, and 0.674 SSIM—comparable to current advanced autoregressive models—while requiring fewer training epochs.
Existing approaches struggle to uncover the causal influence of hidden neurons on neural network outputs, as activation patterns alone are insufficient to decipher internal computational mechanisms. This work proposes CODEC, a novel method that—unlike prior activation-based analyses—decouples network behavior into interpretable, sparse contribution modes by integrating contribution decomposition with sparse autoencoders. Applying this framework, the study reveals cross-layer causal computation pathways and demonstrates its efficacy in both image classification and retinal neural activity modeling. The approach enables precise intervention and visualization of intermediate layers, uncovers a progressive decoupling of positive and negative contributions in deeper layers, and elucidates how compositional interactions among intermediate neurons give rise to dynamic receptive fields.
This work addresses the challenge of efficiently exploiting the fine-grained, unstructured sparsity inherent in binary spikes of spiking neural networks (SNNs), particularly on SIMD-based GPUs where existing sparse computation methods fall short. The authors propose Temporal Aggregated Convolution (TAC), which pre-aggregates spikes across K time steps to reduce convolution invocations. For event-based data, they further introduce TAC-TP to preserve critical temporal information. Challenging the common assumption that “sparsity implies efficiency,” the study advocates a data-dependent temporal aggregation strategy: compressing temporal dimensions for rate-coded inputs to boost both speed and accuracy, while retaining full temporal resolution for event data. Experiments demonstrate that TAC achieves a 13.8× speedup with improved accuracy on MNIST and Fashion-MNIST, while TAC-TP reduces convolution calls by 50% and attains 95.1% accuracy on DVS128-Gesture.
This work addresses the challenge of deploying conventional recurrent spiking neural networks (SNNs) on resource-constrained edge devices, where large parameter counts and slow inference hinder the balance between efficiency and accuracy. To overcome this limitation, the authors propose a novel convolutional recurrent SNN architecture that, for the first time, integrates convolutional recurrent connections with a learnable axonal delay mechanism, accompanied by a tailored training methodology. Evaluated on audio classification tasks, the proposed model achieves a dramatic reduction in model size—cutting recurrent parameters by approximately 99%—while delivering a 52× speedup in inference latency, all without compromising classification accuracy. This advancement establishes a new paradigm for efficient edge intelligence with spiking neural networks.