Score
Designs and implements methods to represent, infer, and manipulate relations among neighboring entities, covering explicit spatial adjacency, attribute-based neighbor similarity, and dual or composite relations between samples. Builds algorithms that mine pairwise attribute relations, identify attribute-similar neighbors, assign or adjust neighbor weights (e.g., downweight visually similar but irrelevant samples), and generate relational supervision signals for downstream models or analyses.
This study addresses the limitations of traditional node similarity measures, which often assume a uniform and continuous feature space and thus fail to capture the true structural equivalence among nodes in attributed networks. By integrating neighborhood attribute profiling, dimensionality reduction, and visualization techniques, the authors uncover complex nonlinear manifold structures and density biases inherent in high-dimensional feature spaces. Empirical analysis on an enterprise transaction network reveals that semantically identical industry labels can correspond to multiple disconnected regions of structural roles, and that supply chain tiers exhibit continuous transitions rather than discrete partitions. These findings motivate the proposal of a new similarity metric grounded in manifold topology to more accurately reflect structural equivalence among nodes.
In modern data analytics, entities and their attribute relationships often span multiple granularities, complicating critical attribute derivation and target entity retrieval. Existing OLAP operators, window functions, and aggregation constructs suffer from limitations in composability, formal expressiveness, and runtime performance. To address this, we propose Multi-Relational Algebra (MRA), a novel algebraic framework grounded in the “slice”—a semantic unit comprising a region of tuples and an associated feature table—thereby transcending conventional single-table or single-column constraints. MRA supports dynamic heterogeneous schema modeling and cross-schema composition, and introduces a formal algebraic system, a slice-based computational model, a unified logical execution engine, and a query optimization framework tailored for data insight discovery. The system has been deployed in production, supporting millions of daily operations and effectively handling complex analytical tasks that resist modeling under traditional relational paradigms.
Traditional regional data modeling often assumes constant spatial dependence strength between adjacent regions, leading to distorted covariance structures; while treating each edge weight as an independent parameter alleviates this issue, it introduces high-dimensional estimation challenges. This paper proposes a low-dimensional basis-function expansion for parameterizing the edge-weight matrix—marking the first application of dimensionality reduction to graph edge-weight estimation—enabling flexible characterization of heterogeneous spatial dependence via a small set of basis coefficients. Integrating graph neural covariance modeling with spatial statistical inference, the method achieves significant improvements in both covariance estimation accuracy and computational efficiency in simulations and empirical studies. It enables robust, scalable modeling of large-scale regional data.
Existing graph mining methods primarily focus on topological subgraph discovery and lack a unified mechanism for jointly modeling syntactic and semantic aspects of association rules over attributed graphs. This paper proposes the MINE GRAPH RULE operator, the first to enable integrated syntactic–semantic expression of graph association rules in an attributed graph database—implemented as an extension to Neo4j. Syntactically, conditions are specified via Cypher-like queries; semantically, rule quality is evaluated using support and confidence metrics. The operator tightly couples graph structure with attribute semantics, leverages Neo4j’s native query optimization, and incorporates relational association rule pruning strategies to ensure efficiency and portability. Experiments demonstrate strong scalability across multidimensional parameters. An open-source plugin implementing the operator significantly enhances both the expressiveness and practical utility of graph association rule mining.
This work addresses program synthesis by proposing a general relational decomposition framework: input-output examples are encoded as sets of logical facts, and the mapping between them is explicitly modeled as a logical relation. Methodologically, it formalizes program synthesis as a relational subtask decomposition problem—marking the first such formulation—and leverages inductive logic programming (ILP) for interpretable, model-agnostic relation learning and inference. Crucially, no domain-specific architecture or task customization is required, enabling cross-task generalization. Evaluated on four challenging benchmarks, the approach significantly outperforms standard sequence- and tree-based neural models. Moreover, by interfacing with off-the-shelf ILP solvers, it surpasses state-of-the-art domain-specific synthesizers across multiple benchmarks. The contribution is a novel, interpretable, modular, and model-independent paradigm for program synthesis grounded in relational logic and ILP.
This work addresses the challenge of efficiently supporting both vector similarity search and arbitrary attribute filtering in high-dimensional approximate nearest neighbor retrieval. The authors propose a lightweight graph-based indexing algorithm that seamlessly integrates attribute filtering into the graph traversal process, overcoming the efficiency bottlenecks of existing methods when handling unseen query vectors combined with complex attribute constraints. Experimental results on multiple real-world datasets demonstrate that the proposed approach significantly outperforms state-of-the-art techniques, achieving substantially faster query latency while maintaining high recall. The method thus offers a compelling balance among flexibility, efficiency, and scalability for hybrid vector-and-attribute search scenarios.
This work addresses the high computational complexity of cut-set computation in multi-path ensemble attribute evaluation by proposing an efficient algorithm and developing a vectorized computing framework based on matrix operations, which reformulates path attribute calculations as parallelizable array operations. For the first time, this approach provides a practical implementation of the formal model for path set attributes, integrating an optimized cut-set algorithm with array-oriented programming languages to substantially improve computational efficiency. Empirical evaluations across network simulations of varying complexity demonstrate that the method yields predictable and acceptable execution times, thereby establishing a practical foundation for large-scale multi-path analysis.
This work addresses a key limitation of existing representative possible world approaches, which preserve only node-level features and thus struggle to support graph mining tasks—such as link prediction—that rely on common-neighbor structures. To overcome this, the paper introduces the Common-neighbor-aware Representative Possible World (CRPW) problem, extending representative possible world modeling from node-level statistics to structural relationships between node pairs for the first time. The authors propose a two-stage optimization algorithm that integrates integer-count acceleration with a Beta-adaptive termination mechanism to efficiently approximate this NP-hard problem. Experimental results on real-world uncertain graphs demonstrate that the proposed method significantly outperforms current state-of-the-art techniques, achieving superior performance particularly on tasks dependent on common-neighbor structures.