Omni2Web: Benchmarking Audiovisual Website Development

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建Omni2Web基准,利用音频和视觉信息解决网页编辑请求中指示代词指代不明的问题,评估多种模型在直接编辑、意图恢复及指令实用性上的表现。
📝 Abstract
Screen-recorded web editing requests contain weak deictic expressions such as ``this'' and ``there,'' whose referents depend on speech, cursor trajectories, page state, and edit history. Such requests require intent recovery beyond the explicit specifications assumed by many existing web-editing benchmarks. We introduce Omni2Web, a bilingual benchmark of 918 instances spanning 13,907 edit steps. It defines three complementary tracks: Direct Editing evaluates webpage editing from recordings, Instruction Recovery measures explicit intent recovery, and Instruction Utility tests whether recovered instructions can drive a fixed code executor. We evaluate 17 open- and closed-source models. The best models attain 51.17 on the Edit Fidelity Score (EFS) for Direct Editing and 49.14 on the Instruction Recovery Score (IRS); under the fixed executor, the strongest recovered instructions reach 51.08 EFS, still far below the 89.69 EFS obtained with oracle instructions. Step-level analyses show that correct grounding does not guarantee successful edits, while some Omni models recover instructions that the fixed coding model executes substantially better than their direct edits. Controlled ablations further demonstrate the value of temporally aligned audiovisual evidence, while alternative judges preserve the leader and broad ordering. Together, these findings reveal substantial headroom in multimodal intent recovery and code execution and highlight the promise of pairing Omni rewriters with coding models.
Problem

Research questions and friction points this paper is trying to address.

weak deictic expressions
intent recovery
web editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Intent Recovery
Audiovisual Evidence
Bilingual Benchmark
💼 Related Jobs
No related jobs found.
M
Minghao Han
FDU; Alibaba Token Hub, Alibaba Group
Zhenghao Xing
Zhenghao Xing
The Chinese University of Hong Kong
Multimodal LearningComputer Vision
X
Xize Cheng
Alibaba Token Hub, Alibaba Group
Y
Yuxuan Wang
Alibaba Token Hub, Alibaba Group
J
Junming Lin
Alibaba Token Hub, Alibaba Group; THU
L
Ling Wang
Alibaba Token Hub, Alibaba Group; PolyU
Y
Yinsong Yan
Alibaba Token Hub, Alibaba Group; PolyU
Yunfei Chu
Yunfei Chu
Alibaba Group
machine learning
Qize Yang
Qize Yang
Tongyi Lab, Alibaba Group
Computer VisionDeep Learning
Jin Xu
Jin Xu
Qwen Team, Alibaba Group
Multimodal InteractionLarge Language ModelSpeech SynthesisVideo/Audio Processing