GroundedGEO: Auditing the Evidence Gap in Generative Search Rankings

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过建立证据配对基准和GroundedGEO重排序模型,旨在解决生成式搜索系统中难以区分真实与伪造内容的问题。
📝 Abstract
Generative search systems rank products and services for consequential decisions, and publishers can cheaply make candidate text look relevant. Yet evidence status is not a text property but a claim-evidence relation: text-only rankers and defenses cannot separate honest detailed content from fabricated detail, creating an identifiability gap. We audit this gap with an evidence-paired benchmark (50 e-commerce queries, 1,950 cases) and a claim-level reranker, GroundedGEO, that penalizes query-relevant claims lacking support in a supplied packet. Matched rich variants control format and volume; packet twins add attestations at fixed text, while thinned packets withdraw them. On the frozen listwise ranker Qwen2.5-7B, unsupported-rich variants show significant normalized rank gain over clean candidates (+0.065 to +0.092 across claim profiles, Holm-corrected), while supported and neutral controls do not; the effect is model-dependent (marginal on MiMo-v2.5, absent on GLM-5.3-Flash). On a frozen pointwise scorer, oracle evidence labels cut the unsupported-rich top-3 rate from 0.65 to 0.43 (laundering from 0.61 to 0.39) at lambda=40 with zero false suppression; packet twins restore the original rates without changing text. Against a 370-claim human gold, all tested automatic judges fail the preregistered reliability gate, although the best local judge retains 79-100% of oracle suppression with zero measured false suppression on protected arms. Separately, stripping attestation coverage increases false suppression by 0.307. These diagnostic effects identify two limits on the evidence channel: label quality and packet coverage. They do not validate an automatic defense, and interpretation of the adverse human-gold arm remains pending adjudication.
Problem

Research questions and friction points this paper is trying to address.

Generative Search Systems
Evidence Gap
Text-only Rankers
Claim-Evidence Relation
Identifiability Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Search
Evidence Gap
Claim-Evidence Relation
GroundedGEO
Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yihan Xia
Shenzhen University
H
Huiling Fan
Shenzhen University
K
Kangrong Zhong
Shenzhen University
Taotao Wang
Taotao Wang
Shenzhen University
Blockchain and Blockchain NetworksWireless Communications and Networking