IMPLICIT-Bench: Measuring Implicit Bias in Text-to-Image Models under Neutral Prompts

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对文本到图像模型在中性提示下的隐式偏见问题,提出IMPLICIT-Bench基准测试方法来测量这种偏见,并评估去偏见方法的效果。
📝 Abstract
Text-to-image (T2I) models are typically evaluated for bias using slot-based templates such as ``a photo of a [profession]''. Such templates probe only \emph{explicit} demographic attributes (e.g., gender, skin tone) in isolation. They overlook a broader \emph{implicit} bias that arises in natural prompts: when stereotype-relevant attributes are left unspecified, models still default to stereotypical outputs. We introduce IMPLICIT-Bench, a benchmark for measuring implicit bias in T2I models under such prompts. The key design is a structured-knowledge-graph (KG) construction of controlled prompt triplets: neutral, stereotype, and anti-stereotype variants that differ only along a single bias dimension while preserving scene semantics. This enables precise attribution of bias effects that template benchmarks cannot achieve. IMPLICIT-Bench comprises 5,493 prompts across 11 bias categories, validated through multi-model agreement, CLIP-based verification, and human evaluation. Using this benchmark, we show that state-of-the-art T2I models exhibit systematic bias under neutral prompts, a failure mode largely invisible to existing evaluations. We then use IMPLICIT-Bench to evaluate debiasing methods, uncovering a fundamental trade-off between bias reduction and semantic fidelity.
Problem

Research questions and friction points this paper is trying to address.

Implicit Bias
Text-to-Image Models
Neutral Prompts
Innovation

Methods, ideas, or system contributions that make the work stand out.

IMPLICIT-Bench
structured-knowledge-graph
implicit bias
text-to-image models
neutral prompts