🤖 AI Summary
This study addresses the limitation of conventional video super-resolution evaluation, which relies on distortion metrics that fail to reflect actual recognition performance in surveillance scenarios. To this end, we propose a semantics-oriented evaluation framework. By constructing a stochastic composite degradation benchmark to replace standard bicubic downsampling, we introduce two task-specific metrics, FaceRecBox and TextRecBox, which directly quantify downstream utility through face and license plate recognition pipelines. Validation using RCDM and its residual-gated variant demonstrates that the proposed approach improves face identification accuracy to 86.93% and significantly enhances license plate text recognition scores. These results confirm that our framework effectively strengthens visual information recovery under complex surveillance conditions, offering a more practical assessment paradigm for real-world deployment.
📝 Abstract
Video super-resolution (VSR) is normally judged by PSNR and SSIM on clips that were downsampled bicubically, although in surveillance its purpose is to make faces and licence plates \emph{recognisable}. We present FANVIDv2, a benchmark that scores VSR by what a recognition pipeline can do with its output. FANVIDv2 provides $320\times180$ low-resolution (LR) clips with high-resolution (HR) references for 48 public figures (with one HR gallery image each) and 375 licence-plate clips covering 360 distinct plate strings. LR clips are generated with a randomised compound degradation (blur, resize jitter, sensor noise, JPEG compression, final downsampling) rather than bicubic downsampling alone. Two metrics score recognition \emph{inside} detections: FaceRecBox rewards a face only if it is localised and correctly identified, and TextRecBox scores plate transcriptions by normalised edit distance weighted by localisation quality. With a 2.3\,M-parameter VSR baseline (RCDM), FaceRecBox rises from 0.6864 to 0.7222, identity accuracy on matched faces from 84.35\% to 86.93\%, and TextRecBox from 0.3088 to 0.3667; a residual-map gated variant (RCDM-RMGF) reaches 0.3801 on plates. We describe the degradation model, the baseline architectures and the scorers in detail, and release annotations, metadata, download and degradation scripts and evaluation code.