FANVID: A Benchmark for Face and License Plate Recognition in Low-Resolution Videos

📅 2025-06-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Low-resolution (180×320) surveillance videos hinder reliable per-frame identification of faces and license plates. To address this, we introduce FANVID—the first cross-modal recognition benchmark for low-resolution video—comprising 1,463 LR video clips (20–60 FPS), 63 face identities, and 49 license plate identities, explicitly requiring temporal modeling for recognizing targets indiscernible in individual frames. We propose two novel tasks: video-level face matching against high-resolution reference photos, and dictionary-free license plate text recognition. To enhance realism, we incorporate distractors and design a joint evaluation protocol combining identity-centric mAP@0.5 and character-level accuracy. Our end-to-end baseline integrates pretrained video super-resolution, temporal-aware detection, and recognition modules, built upon 31,096 manually refined bounding boxes. It achieves 0.58 and 0.42 on the respective tasks. We publicly release the dataset, annotation guidelines, evaluation code, and models to advance research on temporal recognition in low-resolution video.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Data and user privacy-enhancing technologies for the WebSecurity and Privacy: Large-scale security measurements
📝 Abstract
Real-world surveillance often renders faces and license plates unrecognizable in individual low-resolution (LR) frames, hindering reliable identification. To advance temporal recognition models, we present FANVID, a novel video-based benchmark comprising nearly 1,463 LR clips (180 x 320, 20--60 FPS) featuring 63 identities and 49 license plates from three English-speaking countries. Each video includes distractor faces and plates, increasing task difficulty and realism. The dataset contains 31,096 manually verified bounding boxes and labels. FANVID defines two tasks: (1) face matching -- detecting LR faces and matching them to high-resolution mugshots, and (2) license plate recognition -- extracting text from LR plates without a predefined database. Videos are downsampled from high-resolution sources to ensure that faces and text are indecipherable in single frames, requiring models to exploit temporal information. We introduce evaluation metrics adapted from mean Average Precision at IoU>0.5, prioritizing identity correctness for faces and character-level accuracy for text. A baseline method with pre-trained video super-resolution, detection, and recognition achieved performance scores of 0.58 (face matching) and 0.42 (plate recognition), highlighting both the feasibility and challenge of the tasks. FANVID's selection of faces and plates balances diversity with recognition challenge. We release the software for data access, evaluation, baseline, and annotation to support reproducibility and extension. FANVID aims to catalyze innovation in temporal modeling for LR recognition, with applications in surveillance, forensics, and autonomous vehicles.
Problem

Research questions and friction points this paper is trying to address.

Recognizing faces and license plates in low-resolution surveillance videos
Matching low-resolution faces to high-resolution mugshots without errors
Extracting accurate text from blurry license plates without reference databases
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video-based benchmark for LR recognition
Temporal modeling for face and plate matching
Adapted evaluation metrics for identity accuracy
🔎 Similar Papers
No similar papers found.
Kavitha Viswanathan
Kavitha Viswanathan
Research Scholar
Deep LearningMachine Learning
V
Vrinda Goel
Department of Electrical Engineering, Indian Institute of Technology Bombay, India
S
Shlesh Gholap
Department of Electrical Engineering, Indian Institute of Technology Bombay, India
D
Devayan Ghosh
Department of Electrical Engineering, Indian Institute of Technology Bombay, India
M
Madhav Gupta
Department of Electrical Engineering, Indian Institute of Technology Bombay, India
D
Dhruvi Ganatra
Department of Electrical Engineering, Indian Institute of Technology Bombay, India
S
Sanket Potdar
Department of Electrical Engineering, Indian Institute of Technology Bombay, India
Amit Sethi
Amit Sethi
Indian Institute of Technology Bombay, Indian Institute of Technology Guwahati, University of
Image processingcomputer visionmachine learningmedical image processing