🤖 AI Summary
This study addresses the inefficiency of manual methods in large-scale typographic comparison of 17th-century Spanish playbills by proposing a character prototype–based statistical framework. The approach automatically extracts, clusters, and aligns character images to compute inter-book font distances and introduces, for the first time, an a contrario significance test to assess the reliability of observed typographic differences, enabling robust automated comparison of both roman and italic typefaces. By overcoming the scalability limitations of traditional analyses, the method—validated by domain experts—successfully uncovers new printer attributions and revises existing conclusions, thereby advancing digital bibliographical scholarship toward large-scale, automated investigation.
📝 Abstract
We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. This provides a tool that accelerates philological analysis. Using character prototypes derived from clustering and aligning automatically extracted character images, the method defines a typeface distance between any two books. To produce actionable outputs, we develop an a contrario statistical framework to interpret the significance of the computed typeface distances. We apply the method to the philological study of 17 th -century Spanish printed theatre chapbooks in a quantity that exceeds the capabilities of systematic visual inspection by human experts. Our method enables the automatic comparison of Roman and Italic types extracted from different books. After validation by human experts, our method has led to new printer attributions being discovered, and former printer attributions being revised. This success strongly suggests that our method has the potential to enable digital bibliography on a larger scale than was previously possible.