🤖 AI Summary
This study addresses the bottleneck in sustainable protein discovery caused by reliance on costly sensory testing and slow iteration cycles. To this end, it establishes a multimodal sensory prediction benchmark encompassing molecular classification and food ranking, alongside a privacy-preserving evaluation platform. Methodologically, this work is the first to quantify the reliability ceiling of sensory data, constructing a cross-scale evaluation framework to delineate the performance boundaries of machine learning. It further integrates multimodal learning with concordance analyses, such as Krippendorff's Alpha, to mine large-scale evaluation datasets. The optimal model achieves a pairwise accuracy of 0.661, matching the median performance of human panelists. These findings provide a reliable computational proxy standard for alternative protein screening, significantly accelerating the discovery pipeline while reducing dependence on traditional sensory panels.
📝 Abstract
Sustainable protein discovery lacks the fast computational proxies, analogous to molecular docking or density functional theory, that accelerate drug and materials discovery. Evaluating whether a novel food tastes like its animal-based target requires expensive human sensory panels, bottlenecking the design-build-test loop. We introduce TasteBench, a multimodal benchmark and privacy-preserving competition for sensory prediction, spanning two tasks: a food-level ranking task built on 21K+ human evaluations across 215 plant-based foods in 24 product categories, yielding 935 within-category ranking pairs, and a supporting molecular-level taste classification task over 15K flavor molecules. To enable rigorous interpretation of model performance, we characterize the ground truth: inter-rater agreement among panelists is low (Krippendorff's $α= .077$), and the split-half reliability ceiling of panel-aggregated rankings is .825, establishing the range within which ML systems on this benchmark should be assessed. We evaluate baselines across four input modalities; on the same pairs panelists rated, the best model achieves .661 pairwise accuracy, competitive with the median individual panelist (.650), and .683 across all within-category pairs. TasteBench provides the evaluation infrastructure and baselines for measuring progress on computational screening for sustainable protein discovery.