Score
Design and train models that encode individual items as embeddings conditioned on the set or partial selection they appear in, and that construct embeddings representing whole sets. These representations capture combinatorial synergies among set members (e.g., card interactions in a deck) and are produced to serve as features for prediction, ranking, or other downstream analyses.
This paper presents a systematic survey of set function learning—a machine learning paradigm for modeling permutation-invariant mappings over unordered, variable-length sets (e.g., point clouds, multi-label instances). Addressing the theoretical foundations and modeling paradigms, it unifies deep approaches—including DeepSets, Set Transformers, and graph neural network variants—with non-deep methods such as symmetric function approximation, kernel-based techniques, and combinatorial optimization. It introduces, for the first time, a comprehensive taxonomy for invariance-aware modeling that explicitly covers representation, aggregation, and equivariance properties. The survey rigorously delineates the applicability boundaries of each method across canonical tasks like point cloud classification and multi-label prediction, empirically validating scalability and generalization performance. By clarifying the fundamental capabilities and limitations of prevailing models, this work consolidates set modeling as a theoretically grounded and practically impactful subfield of machine learning.
High-dimensional data analysis faces fundamental challenges including empirically unjustified embedding method selection, poorly characterized performance bounds, and fragmented theoretical debates. This study systematically reviews mainstream dimensionality reduction techniques—including t-SNE, UMAP, PCA, and autoencoders—synthesizing scattered literature and key controversies to propose the first practice-oriented three-dimensional framework for low-dimensional embeddings: “generation–evaluation–application.” We conduct a comprehensive empirical evaluation across diverse real-world datasets and downstream tasks, rigorously characterizing each algorithm’s trade-offs in preserving local versus global structure, robustness to noise and hyperparameter variation, and interpretability. Our analysis establishes clear applicability boundaries and inherent limitations for each method. The resulting best-practice guidelines integrate theoretical rigor with engineering feasibility, providing the field with standardized evaluation protocols and principled criteria for method selection. (149 words)
Existing machine learning datasets inadequately support AI-assisted open-ended research in professional mathematics—particularly in algebraic combinatorics—due to insufficient scale, structural richness, and formal verifiability. Method: We introduce the Algebraic Combinatorics Dataset Repository (ACD Repo), the first benchmark suite designed for cutting-edge mathematical research, covering nine unsolved problems with over one million structured, formally verifiable examples per problem. Our approach integrates supervised narrow-model training, model interpretability analysis, large language model–driven program synthesis, and symbolic encoding of combinatorial structures—establishing a novel “interpretable modeling + program synthesis” paradigm for conjecture generation. Contribution/Results: We release nine high-quality, reproducible datasets; substantially lower the barrier to AI-augmented original mathematical conjecturing; and empirically characterize the limits of neural models in abstract pattern induction—demonstrating both their capacity for nontrivial structural generalization and their systematic failures in higher-order combinatorial reasoning.
Multilayer networks pose significant challenges for embedding learning and link prediction due to their heterogeneous connection types and structural complexity. This work presents a systematic review of existing approaches, introduces a refined taxonomy for modeling methods, and establishes a fair and reproducible evaluation framework. Notably, it designs a novel testing protocol tailored specifically for directed multilayer networks. By doing so, this study establishes the first standardized evaluation paradigm for multilayer network embedding learning, substantially enhancing both link prediction performance and cross-method comparability. The proposed framework advances the field toward more efficient and rigorous research practices.
This work addresses the interpretability and safety of deep neural network representations. We propose a novel geometric modeling paradigm: treating neural representations as coordinate systems embedded in the data distribution, abstracting generic data distributions as random lattice structures, and—crucially—introducing percolation theory for the first time to analyze their structural properties. Building on this foundation, we establish a unified classification framework for three representation types—contextual, compositional, and surface features—qualitatively integrating multiple mechanistic interpretability findings. By unifying geometric representation modeling, random lattice theory, and percolation analysis, we derive mathematically verifiable links between data distribution structure and neural representation type. This yields the first theoretically rigorous and empirically testable mathematical framework for representation disentanglement, and charts a principled path toward safe, reliable AI representation learning. (149 words)
This work addresses the challenge of detecting generative AI–produced content. We propose an unsupervised, interpretable embedding-space analysis method: semantic embeddings of text or images are extracted using pre-trained large language or multimodal models; subsequently, dimensionality reduction (e.g., PCA) uncovers an intrinsic, low-dimensional distributional shift between AI-generated and human-created samples—rendering them highly separable without supervision. This phenomenon is systematically validated for the first time and endowed with human-interpretable semantic meaning (e.g., topic coherence, syntactic redundancy). Experiments across diverse generative models—including ChatGPT, Gemini, and Stable Diffusion—demonstrate that high-accuracy separation is achieved solely from raw embeddings and unsupervised projection, without fine-tuning, labeled data, or model-specific detectors. Our approach thus significantly enhances both generalizability and interpretability of AI-content detection.
This work proposes a formal framework for characterizing the machine learnability of sets of binary strings through bounded-complexity Boolean autoencoders. The framework unifies recognizability, generatability, and learnability from examples as structural properties of discrete sets, realized via Boolean threshold function networks. By integrating an iterative optimization algorithm, the approach effectively approximates and converges to strictly learnable states. Empirical evaluations demonstrate that diverse discrete structures—including Rorschach patterns and complex “wild” sets—exhibit machine learnability under this formalism, underscoring both its expressive power and practical utility.
本文通过传播随机特征,无需复杂模型设计和训练,利用随机游走和匿名游走诱导的隐式层次结构生成图嵌入,有效捕捉节点邻近性和结构角色。
This work investigates the intrinsic structural properties preserved by winning subnetworks in the lottery ticket hypothesis within feature space. By constructing an interpretable compositional toy model and integrating feature-space distance metrics, lightweight probing techniques, structured SGD training, and sparse retraining, the study reveals that winning tickets do not rely on specific weight values or neuron identities. Instead, they correspond to a family of compatible encoding positions that are proximate to their final representations at initialization and exhibit low mutual interference. The proposed lightweight probe significantly outperforms conventional weight-based lottery identification methods in both accuracy and encoding recovery fidelity, thereby demonstrating that the geometric structure of feature space plays a dominant role in the lottery ticket mechanism.
This work addresses the challenge of general-purpose algorithm selection without requiring domain-specific knowledge. It proposes ZeroFolio, a novel framework that treats problem instances as raw text and leverages pretrained text embeddings to generate representations, thereby eliminating the need for handcrafted features or task-specific training. Algorithm selection is performed via a weighted k-nearest neighbors approach based on Manhattan distance. ZeroFolio constitutes the first fully feature-free and tuning-free algorithm selection method, applicable across diverse problem domains represented in textual form. Evaluated on 11 ASlib scenarios, ZeroFolio outperforms random forest models using handcrafted features in 9 cases and surpasses scenario-tuned random forests in 8, achieving performance comparable to AutoFolio—all without any configuration or hyperparameter tuning.
This study addresses the challenge of predicting deck strength in Magic: The Gathering draft scenarios, where decisions must be made under incomplete information and complex card synergies. To tackle this problem, we propose a sequential encoder model that introduces, for the first time, a set-based, context-aware card embedding method to effectively capture combinatorial effects within pick sequences. Leveraging a large-scale dataset of real draft logs, our approach establishes the first learnable benchmark for deck strength prediction, significantly outperforming conventional linear models. The proposed framework provides a data-driven solution that enhances draft decision-making by accurately modeling the nuanced interactions among selected cards throughout the drafting process.