🤖 AI Summary
This study addresses the low pose estimation accuracy and frequent grasping conflicts encountered with transparent or reflective laboratory vessels under near-frontal viewpoints. To overcome these challenges, we propose a 3D-printed multi-fiducial marker benchmark that arranges four AprilTags obliquely within a single compact volume to recover perspective cues, innovatively balancing pose consistency with the contact surface requirements of parallel-jaw grippers. By integrating AprilTag detection with a Perspective-n-Point (PnP) solver, the proposed method reduces the mean orientation error from 2.18° to below 0.47° while significantly decreasing positional error. Furthermore, this work reveals the inherent trade-offs between estimation accuracy and graspability across different marker layouts, providing practical design guidelines for robotic manipulation in laboratory environments.
📝 Abstract
Robotic manipulation of labware is difficult when transparent or reflective objects must be identified and localized. Coded planar fiducials are a practical retrofit: easy to print, they leave the marked face flat and graspable. Yet a single planar tag is least reliable in near-frontal views, where perspective cues fade. Non-planar geometries restore those cues but intrude on the flat face that a parallel-jaw gripper must contact. Our idea is to tilt multiple tags within one compact footprint, so that each tag is seen at a non-frontal angle even when the marker faces the camera. We propose the quARtet marker, a 3D-printable fiducial embodying this idea: all detected corners of its four tilted AprilTags enter one Perspective-n-Point solve, and a shared configuration defines the fabricated geometry and the detector model. Because tilting consumes flat area, its three layouts trade pose-estimation consistency against graspability. In robot-referenced, same-setup fixed-camera experiments, all three layouts reduced the mean frontal orientation error from 2.18 degree for a single planar tag to 0.24-0.47 degree and the root-mean-square position error from 1.50 to 0.17-0.20 mm. A robot-mounted-camera pose-hold test confirmed this separation under closed-loop visual feedback. In swing-down trials under identical conditions, the two layouts with flat contact strips retained the object with about 2 mm of in-grasp slip, whereas the layout without flat strips slipped by roughly 100 mm. For the tested conditions, the results support a rule: the layout without flat strips when pose-estimation consistency dominates, a layout with flat strips when the marked face must remain graspable.