Score
Designs and implements representations that parameterize the per-pixel lines of sight of pushbroom imaging sensors, encoding each pixel’s sensor-origin, direction, and scanline timing into compact numerical descriptors. Builds ray-conditioned pixel embeddings or tokens that reconcile central-projection geometry with platform orbital and scan kinematics to produce space-geometric per-pixel features for downstream analysis or models.
Existing feedforward 3D foundation models, constrained by central perspective projection, struggle to accommodate the push-broom imaging geometry of satellites, limiting their applicability in multi-view satellite 3D reconstruction. This work proposes a lightweight adaptation framework that requires no fine-tuning of the backbone network: it selects an optimal view sequence through geometric consistency constraints, parameterizes push-broom rays into geometric tokens using rational function models, and introduces a ray-direction-aware adapter to inject these tokens into a frozen Transformer backbone. This approach achieves, for the first time, an effective integration of physical imaging geometry with deep feedforward architectures, significantly enhancing accuracy and robustness in digital surface model (DSM) generation and demonstrating the critical role of explicit geometric embedding and optimized view selection.
To address the low spatial resolution of onboard hyperspectral imagery—hindering real-time downstream applications—this paper proposes a track-wise causal deep neural network architecture tailored for push-broom imaging. The method innovatively incorporates a causal memory mechanism to model spectral-spatial temporal dependencies line-by-line, achieving high reconstruction fidelity while substantially reducing memory footprint and computational cost. Integrated with lightweight network design, causal convolutions, and hardware-aware optimizations, the architecture is specifically adapted to resource-constrained, low-power onboard platforms. Experimental results demonstrate that the proposed approach matches or surpasses state-of-the-art complex models in quantitative metrics (e.g., PSNR and SSIM), enables single-line processing synchronized with imaging acquisition, and achieves, for the first time, real-time onboard super-resolution reconstruction. This work establishes a viable technical pathway for intelligent in-orbit processing of hyperspectral satellite data.
This work proposes a fully parallelizable, pixel-level distributed visual odometry and depth estimation algorithm that overcomes the inefficiencies of traditional approaches, which rely on transmitting redundant and noisy raw pixel data and struggle with on-sensor deployment. The method introduces, for the first time, an on-chip sensor architecture based on Gaussian Belief Propagation (GBP), where pixels exchange photometric observations and surface normal priors in parallel to reach consensus on camera motion. A keyframe-like anchoring mechanism is incorporated to effectively constrain inter-frame baselines and preserve geometric consistency. Experimental results demonstrate that the proposed approach achieves efficient and robust on-chip visual odometry and depth estimation on real-world datasets.
Mobile robots exhibit limited robust autonomy in complex野外 environments due to a lack of empirical design knowledge for multimodal sensor systems. Method: This paper introduces Boxi—a tightly coupled, deeply optimized sensor payload specifically designed for state estimation and mapping. Boxi integrates dual LiDARs, ten heterogeneous RGB cameras (HDR/global/shutter-rolling), an RGB-D camera, a seven-unit heterogeneous IMU array, and dual-antenna RTK GNSS. It pioneers quantitative linkage between hardware design decisions—such as nanosecond-level hardware synchronization, cross-modal joint calibration, and heterogeneous IMU configuration—and VIO/SLAM performance. Contribution/Results: We open-source a reusable sensor “recipe” design guide. Real-world field experiments demonstrate that high-precision synchronization and multi-IMU redundancy reduce state estimation error by over 40%, significantly enhancing robustness in dynamic environment mapping and long-term autonomous navigation.
This work addresses the geometric reconstruction errors introduced by existing Gaussian splatting methods that rely on perspective or affine camera approximations of Rational Polynomial Coefficient (RPC) models. We propose, for the first time, a Gaussian splatting framework natively supporting RPC cameras, directly projecting Gaussian means and covariances through the RPC model to eliminate approximation errors. Our key innovations include integrating the RPC model into the splatting pipeline, designing a Jacobian-based robust covariance projection method, and introducing ray-based depth modeling to overcome the lack of explicit depth in RPC formulations. Evaluated on the DFC2019 and IARPA2016 datasets, our approach reduces mean elevation errors by 29.6%/63.8% and 9.9%/37.9%, respectively, compared to conventional methods, significantly advancing 3D reconstruction accuracy.
Existing video generation methods struggle to robustly model camera motion and scene geometry under non-pinhole camera models such as wide-angle or fisheye lenses, limiting their performance under unified camera control. This work proposes CRePE, the first approach to explicitly encode the geometric projection paths of non-pinhole cameras into positional representations, modeling image tokens as depth-aware positional distributions along rays within a unified camera framework. By integrating monocular geometric priors via a geometric attention adapter and leveraging a Radial MixForcing mechanism, CRePE enables geometry-conditioned video generation within a frozen video DiT architecture. Experiments demonstrate that CRePE significantly outperforms existing methods—such as RayRoPE—in both geometric consistency and perceptual quality, while preserving high-fidelity video synthesis capabilities.
This work proposes a ray-based camera calibration framework tailored for 3D reconstruction, addressing the limitations of traditional reprojection error–based methods that rely on 2D calibration boards and inadequately reflect 3D geometric accuracy. Instead of reprojection error, the approach introduces reconstruction error and intersection error as more representative metrics. It employs a novel icosahedral 3D calibration target and a ring-shaped feature detector, integrated with a generalized distortion model and bootstrapping to refine both intrinsic and extrinsic parameter estimates. Experimental results on synthetic data demonstrate that the proposed method reduces average intersection error by approximately 40%, significantly enhances calibration stability, and validates that ray-level metrics provide a more faithful assessment of 3D reconstruction fidelity compared to conventional approaches.