Live Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究通过构建liveassistant框架,利用自回归策略处理直播中的视听内容和观众互动,以决定是否、何时及向谁提供帮助。
📝 Abstract
Livestreams are long-lasting interactive environments where audiovisual content, viewer activity, host behavior, and platform signals evolve together, creating assistance needs that emerge from the stream itself. We introduce \liveassistant, a framework for mixed-initiative, role-conditioned assistance that formulates livestream interaction as four coupled decisions: \textbf{whether to act, when to act, whom to address, and what to communicate}. At each 10-second interval, one autoregressive policy consumes native audio and video with synchronized comments, gifts, viewer dynamics, and room metadata, then selects \textsc{OBS}, \textsc{MEM}, or \textsc{ANS}. \textsc{OBS} remains silent, \textsc{MEM} records a private semantic update, and \textsc{ANS} specifies a recipient, task, and grounded message. To support this task, we build a trajectory engine that reconstructs real livestream sessions into structured causal supervision, yielding over 320 hours of optimization trajectories and a human-reviewed benchmark of 275 clips and 13,812 decision intervals. We train the policy with Marker-Aware Multiturn Supervised Fine-Tuning (MA-MSFT), which strengthens sparse structured decisions, followed by Streaming Multiturn GSPO (SM-GSPO), which optimizes self-generated trajectories with turn- and trajectory-level credit. On the held-out benchmark, \liveassistant reaches 71.14 state accuracy, 72.67 recipient accuracy, and 58.41 task accuracy, with consistent gains over representative streaming and general multimodal baselines. Together, the formulation, benchmark, and training framework establish livestream assistance as selective participation in a shared social stream.
Problem

Research questions and friction points this paper is trying to address.

Livestreams
Assistance
Interactive Environments
Viewer Activity
Host Behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

mixed-initiative
role-conditioned assistance
autoregressive policy
trajectory engine
Marker-Aware Multiturn Supervised Fine-Tuning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shujian Gao
Institute of Trustworthy Embodied AI, Fudan University
J
Jiamei Yan
ByteDance TikTok
Y
Yuchen Yang
ByteDance TikTok
P
Penghao Zhou
ByteDance TikTok
Q
Qinglei Wang
ByteDance TikTok
Tiehan Fan
Tiehan Fan
Nanjing University
AIGCMultiModal Learning
Yuan Wang
Yuan Wang
Zhejiang University
Medical MLLM
Zuxuan Wu
Zuxuan Wu
Fudan University
Y
Yu-gang Jiang
Institute of Trustworthy Embodied AI, Fudan University