HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild

📅 2026-03-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenging problem of unstructured, long-horizon strawberry harvesting in real greenhouses, where occlusions and specular reflections complicate perception and manipulation. The authors propose the first end-to-end vision–language–action (VLA) closed-loop system tailored for this setting, operating solely on RGB images from three viewpoints without requiring depth maps or geometric calibration. Leveraging less than four hours of real-world teleoperation data collected via VR, they fine-tune pi_0, pi_0.5, and WALL-OSS models using full-parameter fine-tuning combined with LoRA, and introduce an asynchronous inference–control decoupling architecture. In 50 real greenhouse trials, the fully fine-tuned pi_0.5 variant achieves a 74.0% success rate, an average cycle time of 32.6 seconds per pick, and a low fruit damage rate of 4.1%, marking the first successful deployment of VLA policies in complex agricultural environments.

Technology Category

Application Category

📝 Abstract
This work presents the first study on transferring vision-language-action (VLA) policies to real greenhouse tabletop strawberry harvesting, a long-horizon, unstructured task challenged by occlusion and specular reflections. We built an end-to-end closed-loop system on the HarvestFlex platform using three-view RGB sensing (two fixed scene views plus a wrist-mounted view) and intentionally avoided depth clouds and explicit geometric calibration. We collected 3.71 h of VR teleoperated demonstrations (227 episodes) and fine-tuned pi_0, pi_0.5, and WALL-OSS with full fine-tuning and LoRA. Under a unified 50 trials real-greenhouse protocol and metrics spanning completion, pi_0.5 with full fine-tuning achieved success rate of 74.0% with 32.6 s/pick and damage rate of 4.1%. Asynchronous inference-control decoupling further improved performance over synchronous deployment. Results showed non-trivial closed-loop picking with fewer than four hours of real data, while remaining limited by close-range observability loss and contact-dynamics mismatch. A demonstration video is available at: https://youtu.be/bN8ZowZKPMI.
Problem

Research questions and friction points this paper is trying to address.

strawberry harvesting
vision-language-action policy
occlusion
specular reflections
unstructured environment
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language-action
policy adaptation
strawberry harvesting
RGB-only perception
asynchronous inference-control
Z
Ziyang Zhao
The Intelligent Equipment Research Center, Beijing Academy of Agriculture and Forestry Sciences, Beijing 100097, China; The School of Mechanical Electrical Engineering and Automation, Shanghai University, Shanghai 20044, China
S
Shuheng Wang
The Intelligent Equipment Research Center, Beijing Academy of Agriculture and Forestry Sciences, Beijing 100097, China
Z
Zhonghua Miao
The School of Mechanical Electrical Engineering and Automation, Shanghai University, Shanghai 20044, China
Ya Xiong
Ya Xiong
Nercita, Beijing Academy of Agriculture and Forestry Sciences
Agricultural robotics