semantic id design

Design identifier formats, encoding and decoding procedures, and grounding rules that embed recoverable semantic information into identifiers so they can be mapped to catalog items, preserved across transformations, and interpreted or operated on by language models and downstream systems.

semanticiddesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Closed-class words (e.g., prepositions, conjunctions, articles) in source code identifiers—grammatically essential in natural language yet systematically understudied in programming language research—lack empirical characterization and theoretical grounding. Method: We construct CCID, the first manually annotated dataset of 1,275 closed-class identifiers, and integrate extended syntactic pattern modeling, grounded theory coding, and statistical analysis to uncover how such words encode control flow, data transformation, temporal logic, and behavioral roles via part-of-speech sequences. Contribution/Results: We propose a syntax-pattern–based framework for identifier semantic analysis and empirically demonstrate strong correlations between high-frequency closed-class patterns and program behavior. This work fills a critical gap in programming linguistics by providing the first large-scale empirical study of closed-class words in identifiers, with implications for identifier naming assistance, code comprehension, and programming pedagogy.

Analyzes linguistic structure of identifier names with closed syntactic categoriesExplores relationship between closed-category grammar patterns and program behaviorInvestigates how developers encode behavior in source code via naming

The widespread deployment of algorithms—particularly large language models—in high-stakes domains such as healthcare, criminal justice, and finance has intensified challenges surrounding accountability, transparency, and traceability. This work proposes a novel framework that systematically integrates Digital Object Identifiers (DOIs) into algorithmic governance by establishing a unique identity system for algorithms enriched with metadata parameters. Complemented by a dedicated cryptographic authentication protocol and secure API mechanisms, the framework enables end-to-end auditable tracking across the algorithm’s entire lifecycle. It thereby facilitates reliable provenance tracing, bias mitigation, and scientific reproducibility, while laying a verifiable and audit-ready governance foundation for AI agents and multimodal large language models.

AI ethicsalgorithm accountabilityalgorithm traceability

Permanent Data Encoding (PDE): A Visual Language for Semantic Compression and Knowledge Preservation in 3-Character Units

Jul 27, 2025
YT
Yoshiharu Tsuyuki
🏛️ Yokohama City University | Matsumoto Dental University

To address challenges in long-term knowledge preservation—namely, overreliance on digital systems, offline inaccessibility, and intergenerational unintelligibility—this paper proposes a non-electric, human-readable visual language framework. The method employs 2–3-character glyphs as atomic semantic units, integrated with a public dictionary protocol and rule-based semantic expansion, enabling high-density semantic compression and transparent, self-contained visual parsing. Its core contribution lies in unifying lightweight encoding with logically derivable syntax, thereby supporting persistent, maintenance-free storage, manual decoding, and logical reconstruction without power. Experimental evaluation demonstrates robustness and interpretability in disaster recovery and human-AI collaborative scenarios. The framework establishes a deployable, zero-maintenance semantic substrate for intergenerational knowledge infrastructure.

Develops a visual language for long-term knowledge preservationEnables human-readable data without digital systemsEncodes semantic content into compact alphanumeric units

This work addresses the limitations of conventional text encoders in generative recommendation systems, where fragmented tokenization disrupts semantic coherence in item descriptions and misaligns textual embeddings with the geometric structure of visual embeddings, thereby degrading multimodal fusion. To overcome this, the authors propose rendering item text as images and encoding them using a vision-based OCR model to construct semantic IDs grounded in visual signals. This approach represents the first systematic exploration of treating text as a visual modality for semantic representation, yielding more consistent and stable embeddings in both unimodal and multimodal generative recommendation settings. Experiments across four datasets and two backbone architectures demonstrate that OCR-based text representations match or surpass standard text encoders, maintaining robustness even under extreme resolution compression and significantly enhancing cross-modal alignment stability and deployment efficiency.

Generative RecommendationMultimodal FusionSemantic ID

Latest Papers

What's happening recently
View more

This work proposes a deterministic, evidence-driven approach to semantic recovery in production data warehouses lacking documentation and semantic annotations. By integrating a language model with a validation framework, the method leverages structured evidence—including value fingerprints, a library of 26 semantic patterns, and verification rules—to generate column-level semantic descriptions accompanied by confidence scores and provenance traces. A novel “capability detector” mechanism enables calibrated abstention when evidence is insufficient. Evaluated on 680 hidden columns, the approach achieves an accuracy of 0.475, substantially outperforming the baseline of 0.223. In blind tests on clinical data, it recovers 95.5% of ICD-9 codes, fully abstains on columns with no supporting evidence, and maintains 86% execution accuracy with 59% coverage under completely opaque conditions.

column semanticsdata profilingmetadata reconstruction

Existing semantic identifier–based approaches for multimodal CTR prediction struggle to simultaneously preserve embedding semantic coherence, fine-grained continuous signals, and scalable hierarchical identifiers. To address this, this work proposes PaletteID, inspired by color palettes, which constructs a set of carefully selected prototype items as semantic anchors to bridge pretrained multimodal representations with recommendation models. These prototypes are chosen via a Semantic Quality-aware Determinantal Point Process (SQ-DPP) to balance local density and global diversity. PaletteID then employs a retrieval-aggregation mechanism to generate identifier representations that are interpretable, robust, and scalable. Experiments on two public datasets demonstrate that PaletteID significantly improves CTR prediction performance—particularly for long-tail items—while achieving more stable identifier assignments and enhanced semantic interpretability.

codebook assignmentembedding discretizationmultimodal CTR prediction

Hot Scholars

TS

Teng Shi

Renmin University of China
Recommender SystemInformation Retrieval
GP

Gustavo Penha

Spotify Research
Recommender SystemsInformation RetrievalNatural Language ProcessingMachine Learning
XZ

Xiangyu Zhao

Associate Professor, City University of Hong Kong
RecommendationsLarge Language Models (LLMs)TrustworthyAISearch Engine