ProtoSemImage: Image-Valued Prototypes with Deformable Row Alignment for Interpretable Document Classification

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the interpretability limitations of traditional classification prototypes by proposing an explainable document classification method. The core idea involves rendering documents as HSV images, replacing abstract vectors with visual prototypes and performing template matching via deformable row alignment. Methodologically, this work introduces image-value prototypes and named linguistic factor channels to support end-to-end color-space learning and generative decoding. Training is further optimized through a Skip-Gram objective, a four-dimensional bottleneck, and dynamic time warping-style alignment. Experimental results demonstrate that the proposed approach outperforms vector-based prototypes while effectively validating its layout-preservation capabilities. Furthermore, the findings reveal inherent limitations of distance-based matching, thereby establishing a novel paradigm for interpretable document classification.
📝 Abstract
Prototypes in classification models are almost always vectors, and a vector has no readable form. This paper asks what happens when a prototype is an image. Documents give the question a natural form, because a document can be rendered as a multi-channel image in which every token becomes a pixel, so a class representative can take the same shape and the same channel semantics as the inputs it stands for. ProtoSemImage represents each class by one or more visual archetypes: prototype images in a four-channel HSV space whose channels carry named linguistic factors. A Skip-Gram objective learns that color space end to end through a four-dimensional bottleneck, discourse boundary rows become differentiable typed difference rows, and classification reduces to 2D visual template matching: a deformable row alignment between a document image and the archetype bank, in the spirit of dynamic time warping. Because the match is a spatial pattern comparison rather than a linear readout, the model reports where an input departs from its archetype and along which channel, and a generative head decodes each archetype back into text. The image representation works: it beats an otherwise identical model with vector prototypes in all three paired seeds, by between 4.3 and 11.8 points on a ten-class task. The distance-based matching does not. A diagnostic that keeps the representation fixed and swaps only the classifier recovers the sequence baselines, which locates a 20.6-point shortfall in the matching rather than in the color compression, and a benchmark built so that a pair of documents shares a bag of words and differs only in arrangement confirms the layout-preservation it was designed for. We report both directions, because for a representation whose whole purpose is inspect ability, the failure modes are as informative as the gains.
Problem

Research questions and friction points this paper is trying to address.

interpretable classification
image-valued prototypes
document classification
prototype-based learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Image-valued prototypes
Deformable row alignment
Interpretable document classification
Multi-channel HSV space
2D visual template matching
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mohammad Zare
Artificial Intelligence Lab at AriooBarzan Engineering Team, Shiraz, Iran
P
Pirooz Shamsinejadbabaki
Department of Computer Engineering and Information Technology, Shiraz University of Technology, Shiraz, Iran