Reinforcement Learning for the Full Strawberry Harvesting Process: Obstacle Separation, Detachment, and Placement

πŸ“… 2026-07-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the complex contact dynamics in strawberry harvesting caused by severe occlusion and plant deformation by proposing an interaction-aware unified reinforcement learning framework that formulates obstacle separation, fruit detachment, and placement as a sequential decision-making task. The approach employs a hierarchical architecture integrating a high-level policy with low-level Cartesian impedance control, leveraging a shared interaction-aware policy to generate motion across phases and lightweight heuristic logic to coordinate task sequencing and gripper actions. Zero-shot sim-to-real transfer is achieved through feasibility-prioritized observation alignment and domain randomization. Experiments demonstrate a success rate of 89.7% in simulation and 82.0% on physical hardware, with average execution times ranging from 12.99 to 21.73 seconds across occlusion levels 1–5, confirming the system’s cross-platform efficacy and robustness.
πŸ“ Abstract
Severe occlusions and deformable plant structures introduce complex contact dynamics that challenge robotic strawberry harvesting. A policy-driven reinforcement learning (RL) framework with heuristic phase coordination was developed, in which obstacle separation, fruit detachment, and placement were formulated as a sequential decision-making task. A shared interaction-aware policy generated Cartesian motions across all task phases, while lightweight heuristic logic coordinated task progression and gripper events. A shared structured observation space was used to represent target, obstacle, end-effector, and task-context information. A hierarchical architecture combined the high-level policy with low-level Cartesian impedance control for compliant interaction. To support zero-shot sim-to-real transfer, feasibility-first observation alignment and domain randomization were adopted. The policy achieved success rates of 89.7% in simulation and 82.0% in real-world experiments. As the occlusion level increased from 1 to 5, the average execution time increased from 12.99 s to 21.73 s, reflecting greater interaction complexity. These results demonstrated effective transfer of interaction-aware harvesting behaviors to a structurally different robotic platform.
Problem

Research questions and friction points this paper is trying to address.

strawberry harvesting
occlusion
deformable structures
contact dynamics
robotic manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

reinforcement learning
sim-to-real transfer
obstacle separation
interaction-aware policy
hierarchical control
πŸ”Ž Similar Papers
No similar papers found.