Sample-Efficient Multiple Testing with Adaptive Data Collection

๐Ÿ“… 2026-09-22
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡็ ”็ฉถไบ†้€š่ฟ‡่‡ช้€‚ๅบ”ๆ•ฐๆฎๆ”ถ้›†่งฃๅ†ณๅคš้‡ๆต‹่ฏ•้—ฎ้ข˜๏ผŒๆๅ‡บๅŸบไบŽeๅ€ผ็š„ๅŽ้ชŒๆŠฝๆ ทๆ–นๆณ•(e-PS)๏ผŒๆœ‰ๆ•ˆๆŽงๅˆถ้”™่ฏฏๅ‘็Žฐ็އๅนถๅ‡ๅฐ‘ๆ ทๆœฌ้œ€ๆฑ‚ใ€‚
๐Ÿ“ Abstract
This paper studies adaptive experimental design for multiple testing, where an experimenter sequentially chooses which hypothesis to sample. We propose the e-value-based posterior sampling (e-PS) procedure, which uses the empirical average of log e-value increments to guide randomized sampling and applies e-BH to construct rejection sets. Under conditionally valid e-value increments, the procedure controls the false discovery rate at arbitrary stopping times and produces nested rejection sets. We establish high-probability bounds on the number of samples needed to discover all nonnull hypotheses in terms of the growth and concentration of the underlying e-processes. We specialize these bounds to simple-versus-simple, composite-versus-simple, and simple-versus-composite testing. Simulations and experiments using joke ratings and watermarked text illustrate the procedure's power under limited sampling budgets.
Problem

Research questions and friction points this paper is trying to address.

multiple testing
adaptive data collection
false discovery rate
nonnull hypotheses
Innovation

Methods, ideas, or system contributions that make the work stand out.

e-value-based posterior sampling
adaptive experimental design
false discovery rate control