🤖 AI Summary
This study systematically investigates the necessity and effectiveness of temporal modeling in sketch representation learning: Is it essential to treat sketches as sequences, and which types of sequential information—such as coordinate encoding (absolute vs. relative), decoding paradigm (autoregressive vs. non-autoregressive), or positional encoding—most critically impact representation quality? Through controlled ablation experiments, we evaluate these factors across multiple tasks—including classification, retrieval, and generation—on standard multi-task benchmarks. Results show that: (1) absolute coordinates yield significantly better performance than relative coordinates; (2) non-autoregressive decoders consistently outperform autoregressive ones across all evaluated tasks; and (3) the benefit of temporal modeling is highly task- and sequence-dependent, not universally advantageous. To our knowledge, this is the first work to empirically demonstrate the *conditional validity* of temporal modeling for sketches, providing reproducible, lightweight design principles for efficient sketch representation learning.
📝 Abstract
Sketches are simple human hand-drawn abstractions of complex scenes and real-world objects. Although the field of sketch representation learning has advanced significantly, there is still a gap in understanding the true relevance of the temporal aspect to the quality of these representations. This work investigates whether it is indeed justifiable to treat sketches as sequences, as well as which internal orders play a more relevant role. The results indicate that, although the use of traditional positional encodings is valid for modeling sketches as sequences, absolute coordinates consistently outperform relative ones. Furthermore, non-autoregressive decoders outperform their autoregressive counterparts. Finally, the importance of temporality was shown to depend on both the order considered and the task evaluated.