Score
Designs, trains, and evaluates models that map discrete or structured inputs to continuous vector representations (embeddings); and builds or analyzes the resulting embeddings for tasks such as similarity search, nearest‑neighbor retrieval, clustering, dimensionality reduction, or as feature inputs to downstream learning systems.
To address the challenge of predicting analytical operator outcomes on unseen, massive heterogeneous datasets, this paper introduces NumTabData2Vec—the first end-to-end dataset-level vectorization model. It maps raw tabular data to low-dimensional semantic embeddings, enabling cross-dataset semantic similarity search and analytical result inference. By jointly modeling data structure, statistical features, and operator semantics, NumTabData2Vec achieves high-accuracy outcome prediction on previously unseen, real-world multi-source datasets. Evaluated on multiple real-world benchmarks, it significantly outperforms state-of-the-art methods in prediction accuracy, reduces execution latency by over an order of magnitude, and effectively discriminates among diverse practical scenarios. This work establishes a scalable meta-learning paradigm for large-scale data analysis, offering robust generalization across heterogeneous datasets without requiring task-specific fine-tuning.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.
This work addresses the challenge of detecting generative AI–produced content. We propose an unsupervised, interpretable embedding-space analysis method: semantic embeddings of text or images are extracted using pre-trained large language or multimodal models; subsequently, dimensionality reduction (e.g., PCA) uncovers an intrinsic, low-dimensional distributional shift between AI-generated and human-created samples—rendering them highly separable without supervision. This phenomenon is systematically validated for the first time and endowed with human-interpretable semantic meaning (e.g., topic coherence, syntactic redundancy). Experiments across diverse generative models—including ChatGPT, Gemini, and Stable Diffusion—demonstrate that high-accuracy separation is achieved solely from raw embeddings and unsupervised projection, without fine-tuning, labeled data, or model-specific detectors. Our approach thus significantly enhances both generalizability and interpretability of AI-content detection.
This work investigates the learnability of high-dimensional embedding vectors from discrete data, focusing on how sample size, token frequency, and embedding–correlation strength jointly govern estimation accuracy. We propose a low-rank approximate Approximate Message Passing (AMP) algorithm grounded in a correlation–similarity coupled probabilistic model. This marks the first systematic integration of the AMP framework into the theoretical analysis of embedding estimation, enabling rigorous characterization of the phase transition boundary for estimation performance. Leveraging tools from high-dimensional statistical inference and random matrix theory, we derive precise quantitative relationships between embedding estimation error and key problem parameters. Extensive experiments on synthetic data and real-world text tasks validate our theoretical predictions, demonstrating substantial improvements in statistical efficiency and robustness—particularly in high-dimensional, sparse regimes.
High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.
This study addresses the challenge of natural language–driven simulation model discovery by systematically investigating the impact of data representation, Transformer-based embedding models, and reranking strategies on retrieval performance. By constructing multimodal model metadata and leveraging standard information retrieval metrics, the work presents the first quantitative evaluation of open-source embedding models for this task. Experimental results demonstrate that the proposed approach achieves strong performance in recall@5 and nDCG@5, with reranking substantially enhancing effectiveness on complex queries. These contributions establish the first benchmark framework for AI-enabled model reusability, composability, and interoperability in simulation model retrieval.
This study addresses the lack of effective code embedding methods for visual programming languages such as Scratch by systematically evaluating four large language models combined with five embedding strategies on textual representations of Scratch programs. The authors construct datasets for token prediction and program functionality classification tasks, demonstrating that embedding models trained on large-scale Scratch data effectively integrate structural and semantic information. Notably, these embeddings accurately predict the functional correctness of student programs without requiring task-specific fine-tuning. The findings offer a transferable and efficient embedding framework for classroom-level learning analytics, thereby filling a critical gap in representation learning for visual programming education.
This study investigates the relationship between the performance of embedding models and the structural properties of their embedding spaces, with the aim of predicting downstream task effectiveness. Leveraging the MTEB benchmark, the authors evaluate 25 prominent embedding models across five tasks in both English and multilingual settings. They characterize the local and linear structures of embedding spaces using nearest-neighbor overlap and independent component analysis (ICA). The work reveals, for the first time, a remarkably high correlation (up to 0.97) between the degree of local structure preservation in embedding spaces and model performance on downstream tasks. Furthermore, it demonstrates that different tasks exhibit distinct dependencies on local versus linear structural information. These findings indicate that structural characteristics of embedding spaces can effectively predict model performance across diverse tasks, including retrieval, bilingual text mining, pair classification, and summarization.