🤖 AI Summary
This study addresses the high cost and lack of systematic design in manual review of candidate record pairs, which hinders the simultaneous optimization of accuracy, representativeness, and uncertainty coverage under limited budgets. The authors model review as finite-population sampling based on fine-grained stratification and introduce a novel multidimensional stratification framework that integrates match-score intervals, comparison patterns, record-level ambiguity, and demographic groups. Ambiguity is quantified using match probability bands—derived from deciles of model scores—and conditional candidate perplexity. A tunable allocation mechanism is achieved through within-band error tolerances and global budget scaling. Experiments demonstrate that by reviewing only 7% of samples (compared to 23% in the baseline), the framework preserves high-segment matching accuracy and ambiguity distribution, confirming its effectiveness and flexibility under resource constraints.
📝 Abstract
Clerical review of candidate record pairs remains the de facto gold standard for evaluating record linkage, but it is resource-intensive and often designed informally. We propose a design-based framework that treats clerical review as finite-population sampling over fine-grained strata defined by match weight, comparison pattern, record-level ambiguity, and demographic group. Match-probability bands are constructed from model-based score deciles. Within bands, strata combine comparison patterns with an ambiguity factor derived from matchability and conditional candidate perplexity. A band-specific margin-of-error profile encodes substantive priorities, such as tighter precision in high-score bands, while a single scaling parameter enforces the overall clerical budget. We evaluate the framework using a labelled dataset deduplicated in Splink, comprising 50,000 records and approximately 478,000 candidate pairs. We compare a baseline design reviewing about 23% of pairs with a budget-constrained design reviewing about 7%. The baseline accurately estimates global and band-specific match rates, while sampled distributions of comparison patterns, gender, and ambiguity broadly track the population. Under the budget design, global error approximately doubles, with the largest band-level errors in middle-score bands where matches, non-matches, and ambiguous cases are intermixed. Accuracy in the highest-score bands and the band-wise ambiguity profile are largely preserved, although representativeness by comparison pattern and gender declines. The framework generalises to other clerical-review objectives and can incorporate gold-standard data as prior information for design and calibration. It makes explicit the negotiable trade-offs between workload, precision, representativeness, and coverage of linkage uncertainty.