🤖 AI Summary
This study addresses the challenge of reconstructing structural changes in historical street views caused by uneven historical coverage. We construct VIGOR-his, the first decade-spanning cross-view dataset, and propose CrossTimeEdit, which reformulates historical street view generation as an editing task guided by discrepancies between recent street-level and satellite imagery. Methodologically, building upon FLUX.2 as the backbone, we design a flow-matching reinforcement learning framework steered by three-dimensional rewards encompassing instruction alignment, background preservation, and physical plausibility. This framework is optimized through supervised fine-tuning, Flow-GRPO, and a group reward decoupled normalization strategy. Experimental results demonstrate that our model improves three editing metrics by 17.12% over baselines, surpassing existing cross-view generation models in scene consistency, visual realism, and perceptual quality.
📝 Abstract
Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12\% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality. The implementation code, dataset, and model weights are available at https://luhanwen67.github.io/CrossTimeEdit-release/.