Score
Designs, builds, and evaluates algorithms and systems that detect, localize, and classify objects in three‑dimensional sensor data and maintain their identities over time. This work includes producing 3D bounding boxes or shape estimates from point clouds, depth maps or multi‑view imagery, associating detections across frames for tracking, and analyzing detection and tracking performance with appropriate metrics.
To address perception gaps and motion blur caused by fixed frame rates in conventional LiDAR/RGB cameras under high-speed dynamic scenarios, this paper proposes the first continuous-time 3D object detection framework based on stereo event cameras. Methodologically, we design a dual-branch filtering network to jointly extract semantic and geometric features from asynchronous event streams, and introduce a spatiotemporal alignment mechanism along with a center-aligned regression optimization strategy, enabling end-to-end, purely event-driven 3D detection. Our key contributions are: (i) the first demonstration of continuous-time modeling and 3D localization using only binocular event data—bypassing frame-rate limitations entirely; and (ii) a novel dual-filtering mechanism and center-aligned regression that significantly enhance detection accuracy and robustness for fast-moving objects. Experiments on multiple high-speed dynamic sequences show substantial improvements over state-of-the-art event-based and frame-based 3D detection methods.
Existing LiDAR-based 3D detection methods are primarily designed for vehicular platforms and suffer from strong viewpoint dependency and poor cross-platform generalizability—especially on heterogeneous autonomous systems such as quadrupedal robots and UAVs. To address this, we introduce Pi3DET, the first multi-platform LiDAR 3D detection benchmark, and propose a viewpoint-invariant unified detection framework. Our approach jointly models geometric and feature alignment: geometric alignment normalizes coordinate systems and decouples scale across platforms, while feature alignment employs a cross-domain adaptive attention mechanism to bridge domain gaps. This enables effective knowledge transfer among vehicles, quadrupeds, and UAVs. Extensive experiments demonstrate that our method achieves significantly higher mean average precision (mAP) than state-of-the-art methods under cross-platform evaluation settings, validating its robustness and generalizability in complex real-world scenarios. Pi3DET establishes a new paradigm for universal 3D perception systems.
Accurate relative localization of unmanned aerial vehicles (UAVs) with respect to unmanned ground vehicles (UGVs) remains challenging in GPS-denied environments. Method: This paper proposes a real-time 3D detection and 6-DoF relative pose estimation method leveraging a UGV-mounted LiDAR and the PointPillars deep learning architecture—specifically, pillar-based voxelization followed by 2D CNN processing. To our knowledge, this is the first application of PointPillars to UAV relative localization, replacing conventional pipelines involving point cloud segmentation, Euclidean clustering, and heuristic rules. The approach performs end-to-end point cloud processing for 3D UAV detection and integrates geometric constraints to solve for full six-degree-of-freedom relative pose. Contribution/Results: Evaluated in real-world GPS-denied scenarios, the method achieves a 37.2% improvement in localization accuracy over baseline approaches, with markedly enhanced robustness and stability. Validation against ground truth confirms its effectiveness. This work establishes a scalable, lightweight deep learning paradigm for multi-agent collaborative perception and localization.
Detecting out-of-distribution (OOD) objects in open-world autonomous driving scenarios remains challenging due to the lack of 3D point cloud detectors capable of localizing and identifying unknown-category objects. Method: This paper introduces the first general-purpose 3D OOD detection framework, featuring a novel universal objectness modeling scheme that jointly learns object localization and OOD classification; it further incorporates anomaly sample augmentation and end-to-end joint optimization to eliminate reliance on predefined categories. Contribution/Results: We establish the first real-world KITTI Misc benchmark and two synthetic OOD benchmarks (nuScenes OOD and SUN-RGBD OOD). Extensive experiments across multiple baseline detectors demonstrate significant improvements in OOD recall and classification accuracy, establishing a new paradigm for open-set 3D OOD detection in unstructured environments.
Existing visual SLAM methods exhibit severe generalization deficits across diverse applications (e.g., XR, IoT, autonomous driving, UAVs, human pose tracking) and heterogeneous environments (indoor/outdoor, static/dynamic scenes, varying motion patterns), stemming from deep coupling among algorithm design, environmental characteristics, and platform motion dynamics. Method: We propose the first three-dimensional challenge taxonomy—“algorithm–environment–motion”—and systematically evaluate state-of-the-art methods (ORB-SLAM2/3, VINS-Fusion) on multi-source benchmarks (TUM, EuRoC, ARKitScenes, UAV-Human), quantifying performance via absolute trajectory error (ATE), relative pose error (RPE), and tracking loss. Contribution/Results: No method achieves robust cross-domain or intra-domain heterogeneous generalization. To address this, we introduce a principled co-optimization pathway comprising input representation disentanglement, intermediate information reuse, and output dynamic validation—establishing a reproducible benchmark and foundational design principles for universal visual localization.
This work addresses the challenge of efficient online 3D multi-object tracking and pose estimation using only multi-view monocular cameras, without relying on costly 3D annotations or computationally intensive deep models. The authors propose a fast online algorithm grounded in Bayesian optimal multi-object tracking filters, which takes as input only the outputs of a pre-trained 2D detector. By leveraging multi-camera geometric fusion and online multi-view association optimization, the method jointly infers 3D trajectories and object poses. Notably, it requires no 3D training data and remains robust under dynamic camera disconnections and reconnections. The approach achieves significantly faster runtime than existing methods while maintaining high accuracy, demonstrating its practicality and efficiency in real-world multi-camera systems.
Existing monocular animal detection methods provide only 2D bounding boxes, lacking crucial 3D structural and orientation information, and are hindered by the absence of annotated 3D animal datasets. This work proposes the first complete pipeline to generate 3D animal detection labels from monocular RGB images without requiring ground-truth 3D annotations. Leveraging the Skinned Multi-Animal Linear (SMAL) model, the method estimates 3D pose and shape, projects them into the 2D image plane via camera pose optimization, and introduces a cuboid face visibility metric to infer animal orientation. Evaluated on a newly curated Animal3D dataset, the approach achieves high-precision 3D bounding box estimation across multiple species and diverse scenes, effectively bridging the gap in unsupervised training and evaluation for 3D animal detection.
This work addresses temporal inconsistency and artifacts in existing 3D Gaussian splatting–based dynamic scene reconstruction methods under large inter-frame displacements. It introduces, for the first time, an off-the-shelf point tracking model into this framework to extract pixel trajectories and triangulate them into 3D Gaussian space, thereby guiding the initialization of Gaussian positions, rotations, and scales prior to training. This strategy effectively mitigates flickering and color shifts caused by large motions, significantly enhancing reconstruction robustness and rendering consistency. Combined with a multi-device parallelized multi-frame optimization scheme, the proposed system achieves superior reconstruction throughput on real-world datasets compared to current baselines while maintaining high-quality rendering fidelity.
This work proposes a dynamic object detection and tracking method that fuses LiDAR point clouds and fisheye camera images to address human–robot collaboration safety requirements in complex construction environments. By projecting 3D point clouds and 2D semantic detections onto a cylindrical panoramic representation for alignment, the approach enables consistent multimodal data integration. A Kalman filter is incorporated to achieve efficient and robust tracking, particularly during transitions between static and dynamic object states. The method features a concise yet effective multimodal fusion mechanism designed to handle occlusions, scale variations, and rapid motion typical of real-world construction sites. Experimental results demonstrate that the system achieves high-precision, real-time perception of dynamic objects with strong robustness and practical deployability in actual construction scenarios.