Sparse 3D Perception for Rose Harvesting Robots: A Two-Stage Approach Bridging Simulation and Real-World Applications

📅 2025-07-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address low 3D stamen-center localization accuracy, high cost of real-world annotations, and poor cross-domain generalization in rose harvesting, this paper proposes a two-stage sparse 3D localization method: first performing lightweight 2D keypoint detection from stereo images, then fusing monocular depth estimation via a neural network to achieve sub-centimeter 3D localization—bypassing traditional triangulation’s reliance on precise calibration and texture. We innovatively construct a photorealistic, dynamic synthetic dataset using Blender, enabling simulation-to-reality transfer learning without real annotations. Experiments show a 2D detection F1-score of 95.6% on synthetic data and 74.4% on real-world scenes; depth estimation error is only 3% within 2 meters. The method satisfies the real-time and accuracy requirements of resource-constrained agricultural robots, significantly enhancing scalability for automated harvesting of specialty crops.

Technology Category

Intelligent Robots: Multimodal Perception & Sensor FusionComputer Vision: 3D Computer VisionSearch and Optimization: Learning to Search

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
The global demand for medicinal plants, such as Damask roses, has surged with population growth, yet labor-intensive harvesting remains a bottleneck for scalability. To address this, we propose a novel 3D perception pipeline tailored for flower-harvesting robots, focusing on sparse 3D localization of rose centers. Our two-stage algorithm first performs 2D point-based detection on stereo images, followed by depth estimation using a lightweight deep neural network. To overcome the challenge of scarce real-world labeled data, we introduce a photorealistic synthetic dataset generated via Blender, simulating a dynamic rose farm environment with precise 3D annotations. This approach minimizes manual labeling costs while enabling robust model training. We evaluate two depth estimation paradigms: a traditional triangulation-based method and our proposed deep learning framework. Results demonstrate the superiority of our method, achieving an F1 score of 95.6% (synthetic) and 74.4% (real) in 2D detection, with a depth estimation error of 3% at a 2-meter range on synthetic data. The pipeline is optimized for computational efficiency, ensuring compatibility with resource-constrained robotic systems. By bridging the domain gap between synthetic and real-world data, this work advances agricultural automation for specialty crops, offering a scalable solution for precision harvesting.
Problem

Research questions and friction points this paper is trying to address.

Sparse 3D localization of rose centers for harvesting robots
Bridging simulation and real-world data for model training
Optimizing computational efficiency for resource-constrained robotic systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-stage 3D perception pipeline for roses
Photorealistic synthetic dataset via Blender
Lightweight deep neural network for depth
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Taha Samavati
School of Computer Engineering, Iran University of Science and Technology, Narmak, Tehran, Tehran, Iran
Mohsen Soryani
Mohsen Soryani
Associate Professor at Iran university of science and technology
Image & Video ProcessingMedical Image ProcessingComputer VisionRemote SensingPattern Recognition.
S
Sina Mansouri
George Mason University Department of Computer Science, 4400 University Drive, Virginia, United States