Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancement

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of temporal correlation and compression quality in screen content videos caused by abrupt motion transitions and high-frequency details. To tackle these challenges, this work proposes STM-Net, an enhancement framework that incorporates a prior-guided spatiotemporal scheduler and a parallel stream routing mechanism. These components adaptively handle abrupt transitions without explicit detection while preventing feature contamination. Furthermore, by integrating bidirectional temporal feature extraction with a cascaded multi-scale distillation module, the proposed method effectively preserves critical high-frequency information. Experimental results demonstrate that STM-Net outperforms existing state-of-the-art approaches in both objective metrics and subjective visual quality, offering a robust solution for mitigating compression artifacts in screen content videos.
📝 Abstract
Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text and graphics. Conventional video enhancement methods, which rely heavily on temporal continuity, often suffer from performance degradation when processing SCVs due to the disruption of temporal correlations. To address these challenges, we propose the Spatial-Temporal Multi-scale Network (STM-Net), a novel framework specifically tailored for compressed SCV enhancement. Our approach integrates three complementary components: a Prior-Guided Spatio-Temporal Dispatcher (PG-STD) that routes input into three parallel streams to avoid feature contamination, a Bidirectional Temporal Feature Extraction (BTFE) module that adaptively handles abrupt transitions without explicit detection, and a Cascaded Multi-scale Feature Distillation (CMFD) module that preserves critical high-frequency details. Experimental results demonstrate that STM-Net outperforms state-of-the-art methods in both objective metrics and subjective visual quality, providing a robust solution for screen content artifacts. Code is available at https://github.com/HUANGZiyin1/STM-Net.
Problem

Research questions and friction points this paper is trying to address.

Screen Content Video
Video Quality Enhancement
Compressed Video
Temporal Discontinuity
High-frequency Details
Innovation

Methods, ideas, or system contributions that make the work stand out.

Screen Content Video
Spatial-Temporal Multi-scale Network
Feature Distillation
Video Quality Enhancement
Bidirectional Temporal Feature Extraction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziyin Huang
School of Artificial Intelligence, Shenzhen Polytechnic University, Shenzhen, China; and Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong, China
S
Sik-Ho Tsang
Department of Computer Science, Hong Kong Chu Hai College, Hong Kong, China
X
Xinyuan Qin
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong, China
Y
Yui-Lam Chan
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong, China
X
Xueling Zhou
College of Elite Engineers, Dongguan University of Technology, Dongguan, China
F
Feiyu Chen
Department of Computer Science, City University of Hong Kong, Hong Kong, China