VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the oversight of scene text rendering accuracy in existing video generation evaluations, which often yields visually realistic videos containing erroneous text. To bridge this gap, this work proposes the first systematic benchmark for video text rendering, encompassing diverse scenarios such as advertisements. It introduces an automated evaluation pipeline with a chained query mechanism and develops a novel keyframe-guided multi-agent collaborative iterative optimization framework that leverages multimodal techniques to enhance text rendering quality. Experiments reveal significant deficiencies in mainstream models, with even the best-performing model exhibiting a character error rate of 0.250. These findings highlight critical limitations in current approaches and provide clear directions for advancing precise text generation in videos.
📝 Abstract
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.
Problem

Research questions and friction points this paper is trying to address.

Video Generation
Visual Text Rendering
Evaluation Benchmark
Text Fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Text Rendering
Video Generation Benchmark
Automated Evaluation Pipeline
Keyframe-Guided Agentic Framework
Text Fidelity
🔎 Similar Papers
No similar papers found.