🤖 AI Summary
This paper addresses key challenges in logic library cell vector representation learning—namely, weak semantic modeling, heavy reliance on manual annotations, and difficulty capturing multi-attribute and electrical characteristics. To this end, we propose Lib2Vec, a self-supervised framework that automatically learns joint embeddings of cells and arcs directly from standard Liberty files, supporting variable-length pin modeling and attribute-specific representations. Our contributions include: (i) a novel quantitative evaluation mechanism based on regularity testing; (ii) the first generalizable architecture for multi-attribute joint embedding; and (iii) discovery and validation of interpretable logical analogies in the embedding space (e.g., BUF − INV + NAND ≈ AND). Experiments demonstrate that Lib2Vec significantly improves modeling of both functional and electrical similarity among cells, and effectively enhances downstream circuit learning tasks—especially under annotation-scarce conditions.
📝 Abstract
We propose Lib2Vec, a novel self-supervised framework to efficiently learn meaningful vector representations of library cells, enabling ML models to capture essential cell semantics. The framework comprises three key components: (1) an automated method for generating regularity tests to quantitatively evaluate how well cell representations reflect inter-cell relationships; (2) a self-supervised learning scheme that systematically extracts training data from Liberty files, removing the need for costly labeling; and (3) an attention-based model architecture that accommodates various pin counts and enables the creation of property-specific cell and arc embeddings. Experimental results demonstrate that Lib2Vec effectively captures functional and electrical similarities. Moreover, linear algebraic operations on cell vectors reveal meaningful relationships, such as vector(BUF) - vector(INV) + vector(NAND) ~ vector(AND), showcasing the framework's nuanced representation capabilities. Lib2Vec also enhances downstream circuit learning applications, especially when labeled data is scarce.