SeamEdit: A Black-Box VLM-Agnostic Pipeline for Large-Image Semantic Editing

📅 2026-06-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of large-scale image semantic editing—namely semantic distortion, misalignment, and visible seam artifacts—which often compromise generation quality and content coherence. The authors propose a training-free, model-agnostic black-box pipeline that, for the first time, enables high-quality editing of large images by leveraging closed-source vision-language models (VLMs) without fine-tuning. The method integrates overlapped tiling, black-box VLM inpainting, geometric and color consistency correction, seam-risk-aware multi-candidate ranking, and dynamic-programming-based curved seam blending. This framework substantially reduces seam visibility, supports arbitrary-region semantic modifications, and is compatible with diverse off-the-shelf VLMs, achieving natural and seamless editing results without requiring model adaptation.
📝 Abstract
Semantic region editing for large images must satisfy two requirements at the same time: high generative quality and natural integration with surrounding content. Some related methods rely on white-box models and leave the strong generation capability of closed-source models underexplored. Directly applying closed-source models to tiled editing, however, introduces several failure modes: semantic deformation, canvas-level alignment drift, and visible seam artifacts. This paper presents SeamEdit, a training-free and model-agnostic pipeline that treats any VLM with inpainting capability as a black-box oracle. SeamEdit mitigates these issues through a five-stage post-hoc pipeline: overlay-based tile decomposition, black-box VLM inpainting, geometric and color-consistency correction, seam-risk-based multi-candidate ranking, and dynamic-programming curved seam fusion. The pipeline reduces seam visibility and supports semantic modification of arbitrary tile regions.
Problem

Research questions and friction points this paper is trying to address.

semantic editing
large-image editing
seam artifacts
black-box VLM
tile-based inpainting
Innovation

Methods, ideas, or system contributions that make the work stand out.

black-box editing
seamless image fusion
large-image inpainting
model-agnostic pipeline
semantic region editing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiangyu Lyu
Technische Universität Darmstadt, Darmstadt, Germany
D
Dan Lei
Fine-Arts Educator, Yuncheng Middle School, Yuncheng, Shanxi, China