🤖 AI Summary
This study addresses the challenge of fine-grained food safety risk prediction under real-world constraints of limited inspection resources and sparse regional sampling. The authors propose a Transformer-based multi-source fusion framework that integrates 11 million inspection records with heterogeneous socioeconomic and environmental data. Their approach innovatively combines Wilson interval-based confidence modeling with semi-supervised label refinement and employs a three-stage pretraining strategy to effectively mitigate data sparsity. The method reveals, for the first time, the threshold-driven heuristic decision-making mechanism inherent in regulatory practices, offering a novel pathway for AI-assisted policy decisions. Experimental results demonstrate that the model significantly outperforms baseline methods on nationwide 2022 data, and its deployment in Zhejiang province led to an increased detection rate of violations and improved efficiency in inspection resource allocation.
📝 Abstract
Ensuring food safety represents a critical public health challenge, particularly when inspection resources are limited and regional sampling data are sparse. This study proposes a Transformer-based framework capable of forecasting fine-grained, city-level food safety risks by unifying over 11 million inspection records with supplemental demographic, economic, and environmental indicators extracted from the Statistical Yearbook. A three-stage pretraining design leverages partial supervision from the Wilson interval (capturing both safety and risk rankings), together with semi-supervised label refinement, to effectively utilize historical records even when local sample sizes are insufficient. Experimental evaluations on data from 2022 show that the proposed approach outperforms baselines significantly. A subsequent field experiment in collaboration with the Zhejiang Provincial Administration for Market Regulation further demonstrates improved detection rates and more efficient allocation of inspection resources compared to a manually developed plan. Observations of regulatory decision-making reveal a threshold-based heuristic employed by inspectors, hinting that additional training or decision-support interfaces could further enhance the impact of AI-generated risk scores. Overall, these findings underscore that a rigorous integration of large-scale public inspection data, Wilson interval-based confidence modeling, and advanced deep learning can facilitate earlier and more granular identification of food safety threats. By reducing reliance on reactive measures alone, the proposed framework has the potential to advance proactive, data-driven oversight of the global food supply.