🤖 AI Summary
This work addresses the challenges of large-scale image semantic editing—namely semantic distortion, misalignment, and visible seam artifacts—which often compromise generation quality and content coherence. The authors propose a training-free, model-agnostic black-box pipeline that, for the first time, enables high-quality editing of large images by leveraging closed-source vision-language models (VLMs) without fine-tuning. The method integrates overlapped tiling, black-box VLM inpainting, geometric and color consistency correction, seam-risk-aware multi-candidate ranking, and dynamic-programming-based curved seam blending. This framework substantially reduces seam visibility, supports arbitrary-region semantic modifications, and is compatible with diverse off-the-shelf VLMs, achieving natural and seamless editing results without requiring model adaptation.
📝 Abstract
Semantic region editing for large images must satisfy two requirements at the same time: high generative quality and natural integration with surrounding content. Some related methods rely on white-box models and leave the strong generation capability of closed-source models underexplored. Directly applying closed-source models to tiled editing, however, introduces several failure modes: semantic deformation, canvas-level alignment drift, and visible seam artifacts. This paper presents SeamEdit, a training-free and model-agnostic pipeline that treats any VLM with inpainting capability as a black-box oracle. SeamEdit mitigates these issues through a five-stage post-hoc pipeline: overlay-based tile decomposition, black-box VLM inpainting, geometric and color-consistency correction, seam-risk-based multi-candidate ranking, and dynamic-programming curved seam fusion. The pipeline reduces seam visibility and supports semantic modification of arbitrary tile regions.