RoboFind: Multi-Agent Personalized Object Search for People Who Are Blind or Have Low Vision

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为视障人士设计的RoboFind系统通过智能手机与四足机器人协作,实现个性化物体搜索,解决了特定个人物品定位问题,提高了搜索成功率。
📝 Abstract
Blind and low-vision users often need to locate a specific personal object rather than an arbitrary instance of the same category. The task calls for a robot that can move through the space and reach viewpoints the user cannot, and for an accessible interface where the user says which object is meant and learns whether the right one was found. We present RoboFind, a multi-agent framework in which a smartphone teaches the target and a quadruped robot carries out the search. A Target Teaching Agent converts guided smartphone recordings into a semantic target profile and a reusable multi-view reference bank through an accessible capture flow with AR guidance, speech and haptic feedback, and screen-reader support, so later missions refer to a stored object without repeating the teaching process. At runtime, a Navigation Agent explores the environment and proposes candidate targets, a Verification Agent checks each candidate against the stored references, and a Coordination and Recovery Agent completes the mission or triggers recovery and continued search. Across 32 real-robot missions, RoboFind reaches 85.0% success against 25.0% for a reconstructed sequential first-stop baseline over 20 trials with ten targets, and reduces false success from 75.0% to 5.0%. On six shared targets it succeeds in 10/12 trials, against 5/12 for 12 independently executed GPT-6 Astra-only trials. These results show that the multi-agent design fits the demands of personalized object search, where verifying object identity before declaring completion is what makes the outcome something a user can rely on.
Problem

Research questions and friction points this paper is trying to address.

Blind and low-vision users
Personalized object search
Multi-agent framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent framework
personalized object search
target teaching agent
navigation agent
verification agent
🔎 Similar Papers
No similar papers found.
R
Ruiping Liu
Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology, Karlsruhe 76131, Germany
S
Shaofang Quan
Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology, Karlsruhe 76131, Germany
Qian Yin
Qian Yin
University of Texas at Austin
Cancer ImmunotherapyAntiviral vaccineNanoparticle therapeuticsDrug deliveryBiomaterials
J
Jingqi Zhang
Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology, Karlsruhe 76131, Germany
Junwei Zheng
Junwei Zheng
CV:HCI, KIT; CVG, ETH Zurich
Visual LocalizationScene UnderstandingAssistive Technology
Yufan Chen
Yufan Chen
Karlsruhe Institute of Technology
Document AnalysisComputer VisionRobust Deep Learning
Di Wen
Di Wen
Karlsruhe Institute of Technology
Fine-grained Action UnderstandingAnomaly DetectionRobustnessUncertainty
W
Weijia Fan
Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology, Karlsruhe 76131, Germany
Kailun Yang
Kailun Yang
Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous DrivingRobotics
M. Saquib Sarfraz
M. Saquib Sarfraz
Mercedes-Benz / Karlsruhe Institute of Technology (KIT)
Deep LearningComputer visionMachine learning
Tamim Asfour
Tamim Asfour
Karlsruhe Institute of Technology (KIT)
Humanoid RoboticsHumanoid Robots
Kunyu Peng
Kunyu Peng
Karlsruhe Institute of Technology
video understandingopen set recognitiongeneralizable deep learning
Rainer Stiefelhagen
Rainer Stiefelhagen
Karlsruhe Institute of Technology, Karlsruhe, Germany
Computer visionMultimodal interactionAccessibility