Score
Designs, implements, and operates systems that detect and remove duplicate or near‑duplicate records and events across datasets or real‑time streams, including end‑to‑end deduplication pipelines, algorithms, and operational strategies (去重算法 / 去重策略). Evaluates and optimizes matching and indexing techniques (e.g., hashing, fingerprinting, clustering), state and storage management, and tradeoffs between precision, recall, latency, and resource use.
This work addresses the pervasive issue of sub-document-level redundancy in large-scale pretraining corpora, which existing methods struggle to identify efficiently across distributed shards while flexibly preserving redundant copies. The authors propose a scalable sub-document deduplication framework that decouples duplicate detection from copy retention: it leverages natural boundary segmentation, normalized exact hashing, and distributed aggregation to identify duplicate groups, and introduces—for the first time—a frequency- and length-aware adaptive copy retention strategy that overcomes the limitations of fixed heuristic rules. Experiments on FineWeb-Edu and web corpora containing code demonstrate that the proposed approach significantly enhances model training performance, underscoring the importance of explicit control over copy retention during deduplication.
This work addresses the redundant computation caused by duplicate texts in retrieval-augmented generation (RAG) by proposing a byte-level exact block deduplication method that significantly compresses context while preserving generation quality. For the first time, the compression efficacy of this approach is quantified across three real-world RAG scenarios—academic, enterprise, and conversational—achieving compression ratios of 0.16%, 24.03%, and 80.34%, respectively. Rigorous evaluation via multi-vendor large language model APIs, a five-category human-in-the-loop noise filtering protocol, and statistical validation using Wilson confidence intervals consistently demonstrates that all tested cases remain within a <5% quality degradation threshold. These results establish a deterministic optimization pathway that guarantees zero quality regression.
Semantic-level duplicate detection in large-scale data remains challenging, and manual annotation incurs prohibitively high costs. Method: This paper proposes the first end-to-end deduplication model integrating active learning with pre-trained Transformers, reformulating deduplication as a sequence-to-classification task. It innovatively introduces active learning into semantic deduplication for the first time and designs an R-Drop–based enhancement strategy to improve the generalization capability of each annotation round. The approach unifies Transformer pre-training, active sampling, R-Drop regularization, and sequence classification fine-tuning. Results: On benchmark datasets, the method achieves a 28% improvement in Recall over existing state-of-the-art approaches, significantly reducing annotation effort while enhancing model robustness and generalization.
In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.
This work addresses the combinatorial explosion in memory and time that plagues scalable flow- and context-sensitive pointer analysis, where existing optimizations often compromise precision. To overcome this challenge, the authors propose a Multi-level Deduplication Engine (MDE) that recursively identifies structured redundancies, assigns unique identifiers to equivalent computation states, and integrates memoization of operations to enable efficient reuse—thereby surpassing the limitations of traditional non-recursive deduplication techniques. Implemented in C++ and integrated into a pointer analysis framework, MDE demonstrates substantial performance gains on the SPEC benchmark suite, achieving up to an 18.1× reduction in peak memory usage and an 8.15× speedup in runtime. Notably, the optimization benefits intensify with increasing program scale, highlighting MDE’s effectiveness for large real-world applications.
Weak schema constraints in knowledge graphs often lead to predicate redundancy, resulting in semantic duplication, hindered reuse, and degraded data quality. This work is the first to frame predicate redundancy as a core data quality issue and proposes a closed-loop governance framework encompassing detection, resolution, and prevention. The approach integrates automated techniques—such as embedding-based clustering—with human-in-the-loop validation and embeds this synergy into a crowdsourced knowledge graph evolution pipeline. By extending the SciKGDash platform with interactive review capabilities and support for predicate merging or deletion, the system enables semi-automated curation. Evaluation on ORKG reveals that up to 30% of predicates are redundant, primarily due to user behavior and interface design flaws, thereby demonstrating the effectiveness of the proposed human–machine collaborative strategy.
This work addresses the high analytical latency and severe resource contention in HTAP systems caused by traditional ETL pipelines, which entail frequent data movement. To overcome these limitations, the authors propose offloading data transformation logic to an intelligent storage layer that leverages near-data or in-storage computing capabilities to perform format conversion and preprocessing directly at the storage tier. This approach eliminates the overhead of data migration and significantly reduces interference with foreground transactional workloads, thereby enhancing both the performance and real-time responsiveness of analytical queries. Experimental results demonstrate that the proposed architecture achieves a highly reusable, low-latency data processing paradigm under mixed workloads across multiple execution engines, offering an efficient and scalable storage-compute co-design for HTAP systems.