Score
Designs, implements, and evaluates methods and pipelines to detect, cluster, match, and reconcile duplicate or near-duplicate predicates and records across schemas and datasets. This includes building embedding- and similarity-based representations, ranking candidate predicate pairs, performing predicate/schema matching and standardization, merging or preserving canonical representatives, adding curation-time duplicate checks and automated deduplication for modalities such as documents and images, and measuring/removing test–train overlap.
Weak schema constraints in knowledge graphs often lead to predicate redundancy, resulting in semantic duplication, hindered reuse, and degraded data quality. This work is the first to frame predicate redundancy as a core data quality issue and proposes a closed-loop governance framework encompassing detection, resolution, and prevention. The approach integrates automated techniques—such as embedding-based clustering—with human-in-the-loop validation and embeds this synergy into a crowdsourced knowledge graph evolution pipeline. By extending the SciKGDash platform with interactive review capabilities and support for predicate merging or deletion, the system enables semi-automated curation. Evaluation on ORKG reveals that up to 30% of predicates are redundant, primarily due to user behavior and interface design flaws, thereby demonstrating the effectiveness of the proposed human–machine collaborative strategy.
This paper identifies a systemic bias arising from the misuse of BigCloneBench as a ground-truth benchmark for semantic clone detection: 93% of its Weak Type-3/Type-4 clone pairs (N=406, manually verified) are functionally dissimilar, indicating severe label inaccuracy. Method: We conduct bibliometric analysis across 179 papers and apply a truth-quality assessment framework to evaluate ground-truth reliability. Contribution/Results: Our analysis casts doubt on the validity of conclusions in 139 semantic clone studies; high F1 scores often reflect overfitting to spurious dataset patterns rather than genuine semantic understanding. This work is the first to systematically demonstrate—via large-scale manual adjudication, statistical sampling, and meta-analysis—the substantive threat posed by dataset misuse to research validity in this domain. We rigorously establish that BigCloneBench is suitable only for syntactic or textual clone evaluation; its application beyond this scope risks generating misleading scientific conclusions.
To address progressive entity resolution under time-sensitive and resource-constrained settings, this paper proposes the first systematic progressive entity matching framework, decomposing the process into four stages: filtering, weighting, scheduling, and matching. It introduces a unified design space encompassing mainstream approaches and a novel confidence-prioritized dynamic scheduling mechanism, significantly improving early recall and response efficiency. The framework integrates hybrid (rule- and learning-based) filtering, multi-feature weighting, and configurable matching functions, enabling end-to-end, on-demand retrieval of high-confidence results. Extensive experiments across 10 linkage and 8 deduplication datasets demonstrate that the optimal configuration achieves 90% recall of true duplicate pairs 42% earlier on average, while simultaneously improving both F1 score and time efficiency.
To address low accuracy and poor scalability in complex multi-table schema matching for enterprise data integration, this paper proposes LLMatch, a unified framework. Methodologically, LLMatch introduces (1) a novel two-stage semantic optimization strategy: the Rollup module aggregates semantically related columns to enhance generalization, while the Drilldown module enables fine-grained mapping reconstruction; (2) an LLM-driven column-level alignment method integrating schema preprocessing, candidate table filtering, and hierarchical semantic reduction; and (3) SchemaNet—the first benchmark dataset designed for realistic, complex schema-matching scenarios. Experimental results demonstrate that LLMatch significantly improves matching accuracy on multi-table tasks and substantially enhances engineers’ efficiency and debuggability in practical data integration workflows.
This work addresses the incompleteness of query results in decentralized knowledge graph querying caused by lexical heterogeneity. It proposes a method that dynamically discovers and applies local schema alignment rules during link-traversal query processing (LTQP), without requiring prior centralized alignment or altering the original traversal behavior. This approach enables, for the first time, runtime online schema alignment through scoped, on-demand semantic mappings, significantly improving query completeness. Implemented within the Comunica framework, the system integrates a web interface, command-line tools, and a reusable library. Its effectiveness is demonstrated in a decentralized social media scenario, where it recovers complete results with low overhead, establishing a practical foundation for Web-scale distributed LTQP.
This work addresses the poor generalizability of existing self-service entity resolution methods on unseen datasets and their difficulty in balancing precision and recall, which often leads to error propagation and cascading erroneous merges. To overcome these limitations, the authors propose a robust self-service entity resolution system that integrates three key innovations: automated selection among multiple algorithms, decoupled optimization of precision and recall—enforcing precision through rule-based veto mechanisms while enhancing recall via diverse candidate generation—and non-transitive cross-cluster merge validation to prevent error propagation. Extensive experiments on six real-world benchmark datasets, ranging from 864 to 5 million records, demonstrate that the proposed approach significantly reduces incorrect merges, enhances practical utility, and substantially lowers the trial-and-error burden for practitioners.
This work addresses the joint schema-value matching problem between pivot tables and relational tables in data lakes, which demands semantic consistency, value compatibility, and generalization under anonymized data. To this end, we propose PiLLar, a novel framework that, for the first time, formulates the task as a large language model (LLM)-guided Monte Carlo Tree Search (MCTS), enabling unsupervised, training-free cross-domain adaptation. We provide a dynamic theoretical error analysis that guarantees asymptotic convergence. Furthermore, we construct PTbench, the first real-world benchmark for this problem. Experimental results show that PiLLar achieves an average matching accuracy of 87.94% on PTbench, significantly outperforming existing methods and demonstrating its effectiveness and strong generalization capability.
This work addresses the pervasive issue of sub-document-level redundancy in large-scale pretraining corpora, which existing methods struggle to identify efficiently across distributed shards while flexibly preserving redundant copies. The authors propose a scalable sub-document deduplication framework that decouples duplicate detection from copy retention: it leverages natural boundary segmentation, normalized exact hashing, and distributed aggregation to identify duplicate groups, and introduces—for the first time—a frequency- and length-aware adaptive copy retention strategy that overcomes the limitations of fixed heuristic rules. Experiments on FineWeb-Edu and web corpora containing code demonstrate that the proposed approach significantly enhances model training performance, underscoring the importance of explicit control over copy retention during deduplication.