🤖 AI Summary
This study addresses the challenge of auditing the quality of re-described image-text corpora, where length-based proxy metrics often fail and downstream benchmarks introduce confounding factors. To overcome this, we propose a five-axis auditing framework under a fixed budget, introducing Controllable Basic Units (CBUs) as a unified claim unit to enable standardized cross-corpus evaluation. By integrating vision-language models, a dual-judge mechanism based on Qwen and Gemma, and a text budget control algorithm, we construct a reusable, budget-matched auditing pipeline. Experiments demonstrate that our approach significantly increases the number of supported CBUs on corpora such as CC12M. Furthermore, we release a large-scale multi-source re-described corpus comprising approximately 490 million samples, along with the complete suite of auditing artifacts.
📝 Abstract
Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($π$), captioner ($V_c$), and source corpus ($C$). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions $D_{π,V_c,C}$: at a fixed text budget of $B = 64$ it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the common claim unit. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by $+3.39$ to $+6.36$ under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus ($\approx$ 490M) together with the audit-artifact bundle.