CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of reconstructing structural changes in historical street views caused by uneven historical coverage. We construct VIGOR-his, the first decade-spanning cross-view dataset, and propose CrossTimeEdit, which reformulates historical street view generation as an editing task guided by discrepancies between recent street-level and satellite imagery. Methodologically, building upon FLUX.2 as the backbone, we design a flow-matching reinforcement learning framework steered by three-dimensional rewards encompassing instruction alignment, background preservation, and physical plausibility. This framework is optimized through supervised fine-tuning, Flow-GRPO, and a group reward decoupled normalization strategy. Experimental results demonstrate that our model improves three editing metrics by 17.12% over baselines, surpassing existing cross-view generation models in scene consistency, visual realism, and perceptual quality.
📝 Abstract
Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12\% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at https://luhanwen67.github.io/CrossTimeEdit-release/.
Problem

Research questions and friction points this paper is trying to address.

historical street-view generation
cross-view dataset
urban evolution
scene editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Time Editing
Flow-GRPO
Reward-Guided Reinforcement Learning
Cross-View Generation
Historical Street-View
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hanwen Lu
Sun Yat-sen University
Jun He
Jun He
Sun Yat-sen University
Computer VisionDeep LearningRemote SensingUrban Perception
Mingjia Yang
Mingjia Yang
University of Michigan
H
Hao Wei
Sun Yat-sen University
J
Jinhao Huang
Sun Yat-sen University
Yi Lin
Yi Lin
Sun Yat-sen University
Deep LearningUrban PerceptionCross-View DataRemote Sensing
X
Xiang Zhang
Sun Yat-sen University