🤖 AI Summary
This study addresses the lack of reliable evaluation standards in style transfer and the inability of existing metrics to reflect human preferences by proposing the ASTRA benchmark alongside a learning-based evaluation framework. Methodologically, the authors design a two-stage pairwise comparison protocol for user studies to construct ranking ground truth, and train a deep evaluation network on image triplets. Experimental results demonstrate that the proposed ASTRA-Score achieves significantly higher correlation with human perceptual preferences than conventional metrics. This work effectively bridges the gap in standardized evaluation within the field, establishing a robust and perception-aligned automated assessment mechanism for style transfer.
📝 Abstract
Style transfer lacks a reliable evaluation standard: ground truth is inherently ill-defined, and existing automatic metrics often fail to reflect human preference. This paper introduces ASTRA (Assessment of Style TRansfer Algorithms), an approach for automatic evaluation of style transfer algorithms; it contains two components, ASTRA-Data and ASTRA-Score. ASTRA-Data consists of a benchmark image set of content and style references, a collection of style transfer results generated on the benchmark set, and user study data capturing human judgements through a two-stage pairwise comparison protocol. From these annotations, we derive ranking-based ground truth for content preservation, style fidelity, and overall preference. Based on ASTRA-Data, we construct ASTRA-Score, a learnt evaluator that predicts preference-aligned scores from content-style-stylization image triplets, enabling automatic and scalable evaluation of new models applied to the benchmark set. Experimental results demonstrate that ASTRA-Score achieves substantially higher correlation with human rankings compared to prior metrics. Overall, ASTRA establishes a robust mechanism for standardised evaluation of style transfer methods.