build deduplication pipelines

Designs, implements, and operates systems that detect and remove duplicate or near‑duplicate records and events across datasets or real‑time streams, including end‑to‑end deduplication pipelines, algorithms, and operational strategies (去重算法 / 去重策略). Evaluates and optimizes matching and indexing techniques (e.g., hashing, fingerprinting, clustering), state and storage management, and tradeoffs between precision, recall, latency, and resource use.

builddeduplicationpipelines

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$213K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the pervasive issue of sub-document-level redundancy in large-scale pretraining corpora, which existing methods struggle to identify efficiently across distributed shards while flexibly preserving redundant copies. The authors propose a scalable sub-document deduplication framework that decouples duplicate detection from copy retention: it leverages natural boundary segmentation, normalized exact hashing, and distributed aggregation to identify duplicate groups, and introduces—for the first time—a frequency- and length-aware adaptive copy retention strategy that overcomes the limitations of fixed heuristic rules. Experiments on FineWeb-Edu and web corpora containing code demonstrate that the proposed approach significantly enhances model training performance, underscoring the importance of explicit control over copy retention during deduplication.

duplicate retentionfrequency-awarelarge language model pretraining

This work addresses the redundant computation caused by duplicate texts in retrieval-augmented generation (RAG) by proposing a byte-level exact block deduplication method that significantly compresses context while preserving generation quality. For the first time, the compression efficacy of this approach is quantified across three real-world RAG scenarios—academic, enterprise, and conversational—achieving compression ratios of 0.16%, 24.03%, and 80.34%, respectively. Rigorous evaluation via multi-vendor large language model APIs, a five-category human-in-the-loop noise filtering protocol, and statistical validation using Wilson confidence intervals consistently demonstrates that all tested cases remain within a <5% quality degradation threshold. These results establish a deterministic optimization pathway that guarantees zero quality regression.

byte-exact deduplicationinference compute savingsquality preservation

A Pre-trained Data Deduplication Model based on Active Learning

Jul 31, 2023
XL
Xin-Yang Liu
🏛️ Southwest Jiaotong University

Semantic-level duplicate detection in large-scale data remains challenging, and manual annotation incurs prohibitively high costs. Method: This paper proposes the first end-to-end deduplication model integrating active learning with pre-trained Transformers, reformulating deduplication as a sequence-to-classification task. It innovatively introduces active learning into semantic deduplication for the first time and designs an R-Drop–based enhancement strategy to improve the generalization capability of each annotation round. The approach unifies Transformer pre-training, active sampling, R-Drop regularization, and sequence classification fine-tuning. Results: On benchmark datasets, the method achieves a 28% improvement in Recall over existing state-of-the-art approaches, significantly reducing annotation effort while enhancing model robustness and generalization.

Big DataDuplicate DataSemantic Deduplication

In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.

Enhances auditability and traceability in regulated data pipelinesIntegrates data quality control into continuous DataOps managementUnifies rule-based, statistical, and AI methods for anomaly detection

Latest Papers

What's happening recently
View more

This work addresses the combinatorial explosion in memory and time that plagues scalable flow- and context-sensitive pointer analysis, where existing optimizations often compromise precision. To overcome this challenge, the authors propose a Multi-level Deduplication Engine (MDE) that recursively identifies structured redundancies, assigns unique identifiers to equivalent computation states, and integrates memoization of operations to enable efficient reuse—thereby surpassing the limitations of traditional non-recursive deduplication techniques. Implemented in C++ and integrated into a pointer analysis framework, MDE demonstrates substantial performance gains on the SPEC benchmark suite, achieving up to an 18.1× reduction in peak memory usage and an 8.15× speedup in runtime. Notably, the optimization benefits intensify with increasing program scale, highlighting MDE’s effectiveness for large real-world applications.

memory usagepointer analysisredundancy

Weak schema constraints in knowledge graphs often lead to predicate redundancy, resulting in semantic duplication, hindered reuse, and degraded data quality. This work is the first to frame predicate redundancy as a core data quality issue and proposes a closed-loop governance framework encompassing detection, resolution, and prevention. The approach integrates automated techniques—such as embedding-based clustering—with human-in-the-loop validation and embeds this synergy into a crowdsourced knowledge graph evolution pipeline. By extending the SciKGDash platform with interactive review capabilities and support for predicate merging or deletion, the system enables semi-automated curation. Evaluation on ORKG reveals that up to 30% of predicates are redundant, primarily due to user behavior and interface design flaws, thereby demonstrating the effectiveness of the proposed human–machine collaborative strategy.

data qualityduplicate predicatespredicate redundancy

This work addresses the high analytical latency and severe resource contention in HTAP systems caused by traditional ETL pipelines, which entail frequent data movement. To overcome these limitations, the authors propose offloading data transformation logic to an intelligent storage layer that leverages near-data or in-storage computing capabilities to perform format conversion and preprocessing directly at the storage tier. This approach eliminates the overhead of data migration and significantly reduces interference with foreground transactional workloads, thereby enhancing both the performance and real-time responsiveness of analytical queries. Experimental results demonstrate that the proposed architecture achieves a highly reusable, low-latency data processing paradigm under mixed workloads across multiple execution engines, offering an efficient and scalable storage-compute co-design for HTAP systems.

data movementdata transformationHTAP

Hot Scholars

HM

Haitao Mi

Principal Researcher, Tencent US
Large Language Models
MW

Martin Weyssow

Research Scientist, Singapore Management University
Deep Learning for CodeLarge Language ModelsAI4SE
PO

Pedro Ortiz Suarez

Principal Research Scientist, Common Crawl Foundation
Language modelingCorpus linguisticsNamed Entity RecognitionComputational Linguistics
YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
QY

Qifan Yu

Zhejiang University
MLLMmultimodal learningimage generation & editing