🤖 AI Summary
This study addresses the interpretability limitations of traditional classification prototypes by proposing an explainable document classification method. The core idea involves rendering documents as HSV images, replacing abstract vectors with visual prototypes and performing template matching via deformable row alignment. Methodologically, this work introduces image-value prototypes and named linguistic factor channels to support end-to-end color-space learning and generative decoding. Training is further optimized through a Skip-Gram objective, a four-dimensional bottleneck, and dynamic time warping-style alignment. Experimental results demonstrate that the proposed approach outperforms vector-based prototypes while effectively validating its layout-preservation capabilities. Furthermore, the findings reveal inherent limitations of distance-based matching, thereby establishing a novel paradigm for interpretable document classification.
📝 Abstract
Prototypes in classification models are almost always vectors, and a vector has no readable form. This paper asks what happens when a prototype is an image. Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for. ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors. A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping. Because the match is a spatial pattern comparison rather than a linear readout, the model reports where an input departs from its archetype and along which channel, and a generative head decodes each archetype back into text. The image representation works: it beats an otherwise identical model with vector prototypes in all three paired seeds, by between 4.3 and 11.8 points on a ten-class task. The distance-based matching does not. A diagnostic that keeps the representation fixed and swaps only the classifier recovers the sequence baselines, which locates a 20.6-point shortfall in the matching rather than in the color compression, and a benchmark built so that a pair of documents shares a bag of words and differs only in arrangement confirms the layout-preservation it was designed for. We report both directions, because for a representation whose whole purpose is inspect ability, the failure modes are as informative as the gains.