🤖 AI Summary
This study addresses the limitation of passive, instruction-dependent systems in human-robot collaboration by proposing an active assistance method driven by interaction outcomes. An event-driven framework monitors state changes and extracts pre- and post-interaction snapshots, which are processed by a frozen vision-language model (VLM) for reasoning and action sequence generation. Requiring no training or fine-tuning, the approach combines constrained action primitives with an integer ID referencing mechanism to achieve cross-task zero-shot active assistance relying solely on semantic priors. Evaluated across three real-world tabletop tasks, the proposed method demonstrates performance comparable to instruction-driven baselines, thereby validating both its executability and effectiveness.
📝 Abstract
Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an action rather than waiting for instructions. Motivated by this, we investigate an event-driven formulation of proactive assistance, where human--object interaction outcomes initiate assistive reasoning without user-provided task specifications at inference time. To this end, we propose an event-driven framework that monitors workspace state changes with an event monitor and, upon event completion, extracts stabilized pre/post snapshots that characterize the resulting state transition. A frozen pretrained Vision-Language Model (VLM) then uses its semantic priors to infer the task context, decide whether assistance is appropriate, and, when needed, generate a sequence of assistive actions from the observed transition. To make outputs executable and verifiable, we restrict actions to a set of action primitives and reference objects via integer IDs.We evaluate the same framework across three distinct real world tabletop collaboration tasks without task-specific training or fine-tuning. The event-driven framework achieves performance comparable to variants given user instructions.