Score
Design identifier formats, encoding and decoding procedures, and grounding rules that embed recoverable semantic information into identifiers so they can be mapped to catalog items, preserved across transformations, and interpreted or operated on by language models and downstream systems.
Closed-class words (e.g., prepositions, conjunctions, articles) in source code identifiers—grammatically essential in natural language yet systematically understudied in programming language research—lack empirical characterization and theoretical grounding. Method: We construct CCID, the first manually annotated dataset of 1,275 closed-class identifiers, and integrate extended syntactic pattern modeling, grounded theory coding, and statistical analysis to uncover how such words encode control flow, data transformation, temporal logic, and behavioral roles via part-of-speech sequences. Contribution/Results: We propose a syntax-pattern–based framework for identifier semantic analysis and empirically demonstrate strong correlations between high-frequency closed-class patterns and program behavior. This work fills a critical gap in programming linguistics by providing the first large-scale empirical study of closed-class words in identifiers, with implications for identifier naming assistance, code comprehension, and programming pedagogy.
The widespread deployment of algorithms—particularly large language models—in high-stakes domains such as healthcare, criminal justice, and finance has intensified challenges surrounding accountability, transparency, and traceability. This work proposes a novel framework that systematically integrates Digital Object Identifiers (DOIs) into algorithmic governance by establishing a unique identity system for algorithms enriched with metadata parameters. Complemented by a dedicated cryptographic authentication protocol and secure API mechanisms, the framework enables end-to-end auditable tracking across the algorithm’s entire lifecycle. It thereby facilitates reliable provenance tracing, bias mitigation, and scientific reproducibility, while laying a verifiable and audit-ready governance foundation for AI agents and multimodal large language models.
To address challenges in long-term knowledge preservation—namely, overreliance on digital systems, offline inaccessibility, and intergenerational unintelligibility—this paper proposes a non-electric, human-readable visual language framework. The method employs 2–3-character glyphs as atomic semantic units, integrated with a public dictionary protocol and rule-based semantic expansion, enabling high-density semantic compression and transparent, self-contained visual parsing. Its core contribution lies in unifying lightweight encoding with logically derivable syntax, thereby supporting persistent, maintenance-free storage, manual decoding, and logical reconstruction without power. Experimental evaluation demonstrates robustness and interpretability in disaster recovery and human-AI collaborative scenarios. The framework establishes a deployable, zero-maintenance semantic substrate for intergenerational knowledge infrastructure.
This work addresses the limitations of conventional text encoders in generative recommendation systems, where fragmented tokenization disrupts semantic coherence in item descriptions and misaligns textual embeddings with the geometric structure of visual embeddings, thereby degrading multimodal fusion. To overcome this, the authors propose rendering item text as images and encoding them using a vision-based OCR model to construct semantic IDs grounded in visual signals. This approach represents the first systematic exploration of treating text as a visual modality for semantic representation, yielding more consistent and stable embeddings in both unimodal and multimodal generative recommendation settings. Experiments across four datasets and two backbone architectures demonstrate that OCR-based text representations match or surpass standard text encoders, maintaining robustness even under extreme resolution compression and significantly enhancing cross-modal alignment stability and deployment efficiency.
This work proposes a deterministic, evidence-driven approach to semantic recovery in production data warehouses lacking documentation and semantic annotations. By integrating a language model with a validation framework, the method leverages structured evidence—including value fingerprints, a library of 26 semantic patterns, and verification rules—to generate column-level semantic descriptions accompanied by confidence scores and provenance traces. A novel “capability detector” mechanism enables calibrated abstention when evidence is insufficient. Evaluated on 680 hidden columns, the approach achieves an accuracy of 0.475, substantially outperforming the baseline of 0.223. In blind tests on clinical data, it recovers 95.5% of ICD-9 codes, fully abstains on columns with no supporting evidence, and maintains 86% execution accuracy with 59% coverage under completely opaque conditions.
Existing semantic identifier–based approaches for multimodal CTR prediction struggle to simultaneously preserve embedding semantic coherence, fine-grained continuous signals, and scalable hierarchical identifiers. To address this, this work proposes PaletteID, inspired by color palettes, which constructs a set of carefully selected prototype items as semantic anchors to bridge pretrained multimodal representations with recommendation models. These prototypes are chosen via a Semantic Quality-aware Determinantal Point Process (SQ-DPP) to balance local density and global diversity. PaletteID then employs a retrieval-aggregation mechanism to generate identifier representations that are interpretable, robust, and scalable. Experiments on two public datasets demonstrate that PaletteID significantly improves CTR prediction performance—particularly for long-tail items—while achieving more stable identifier assignments and enhanced semantic interpretability.