Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决物理任务视觉指导不适应用户环境的问题,提出了一种生成式教程框架,通过增强现实技术生成符合用户环境的目标图像和演示视频。
📝 Abstract
Visual instructions for physical tasks are typically authored in one context and followed in another, requiring users to translate demonstrated tools, materials, and spatial relationships into their own environment. We introduce Generative Tutorial, a conceptual framework for live visual instruction that depicts intended outcomes and actions within the user's environment and task flow. A formative evaluation of state-of-the-art image and video generation identifies failures and potential benefits across 15 physical tasks. Drawing on these findings, we build an augmented-reality prototype system that proactively generates goal images and demonstration videos using observed workspace context and predicted visual outcomes of preceding actions. A 24-participant lab study found higher task performance quality, greater perceived workspace correspondence, and shorter step-confirmation intervals with the system than with pre-authored guidance. Qualitative findings highlighted how contextual resemblance shapes trust, how generation errors affect interpretation, and how guidance delivery should adapt to users' needs, informing future designs.
Problem

Research questions and friction points this paper is trying to address.

visual instructions
physical tasks
contextualized
environment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Tutorial
Live Visual Instruction
Augmented Reality
Contextualized Guidance
Proactive Generation
🔎 Similar Papers
No similar papers found.