π€ AI Summary
This study addresses the difficulty of quantitatively analyzing and scenario-testing manually annotated large-scale road markings by proposing a controllable synthesis framework based on a Bird's-Eye-View (BEV) pipeline. Methodologically, it employs a conditional Rectified-Flow DiT model guided by text and mask inputs to generate missing center-region markings. The approach introduces a topology-aware auxiliary loss, a fuzzy target training strategy, and a training-free Structured Gaussian Rendering (SGR) post-processing technique, combined with BΓ©zier curve fitting to optimize geometric precision. Experiments demonstrate that the proposed method achieves Buffered F1 scores of 80.8 and 88.0 and clDice scores of 50.2 and 66.2 on the Argoverse2 and Waymo datasets, respectively. These results significantly outperform baselines, effectively enhancing the connectivity and clarity of fine, sparse road markings.
π Abstract
Lane and road markings provide critical guidance for vehicle navigation and multi-agent coordination, yet authoring them at scale remains a manual workflow that limits quantitative analysis and scenario testing. We introduce Controllable Road Marking Generation, which synthesizes a missing center-region marking layout from a drivable-area mask, optional outer-ring markings, and a textual description. Our benchmark uses deterministic, metadata-derived prompts and three output channels: lane dividers, road dividers, and pedestrian crossings. We develop a conditional bird's-eye-view (BEV) pipeline that combines (i) a text-conditioned latent rectified-flow DiT trained with a topology-aware auxiliary loss, (ii) Gaussian-blurred training targets that stabilize learning of thin, sparse markings, and (iii) Structured Gaussian Render (SGR), a training-free post-process that recovers crisp divider geometry by extracting polylines, fitting cubic B\'ezier curves, and re-rendering them as anisotropic super-Gaussian primitives. On 4,597 Argoverse~2 test tiles, our system achieves Buffered F1 of 80.8 and clDice of 50.2, compared with 38.8 and 24.6 for an adapted state-of-the-art mask-refinement baseline. On Waymo dataset, it yields 88.0 Buffered F1 and 66.2 clDice. Component ablations show complementary connectivity gains from topology-aware supervision and SGR. Text-editing experiments reveal that stronger guidance improves edit success but also increases changes to non-target structures. We see this framework as a step toward simulation-ready road-marking variation, automated map completion, and early-stage infrastructure design exploration.