curate dataset edge cases

Design and build data-collection and curation processes that deliberately identify, gather, annotate, and integrate rare, ambiguous, adversarial, or boundary-condition examples across multiple modalities (e.g., text, images, audio, video, sensors). Create sampling strategies, annotation guidelines, quality-control checks, metadata and test splits or benchmarks that document and enable evaluation of model behavior specifically on these dataset edge cases.

curatedatasetedgecases

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Effectively obtaining acoustic, visual and textual data from videos

Sep 06, 2025
JE
Jorge E. León
🏛️ Adolfo Ibáñez University (UAI) | Diego Portales University (UDP)

A critical shortage exists of high-quality, large-scale, semantically aligned acoustic–visual–textual multimodal datasets. Method: This paper proposes an end-to-end framework for constructing video-based multimodal data, integrating three key components: (i) video content filtering, (ii) cross-modal synchronization triplet extraction (audio–frame–subtitle), and (iii) fine-grained description synthesis leveraging image-to-text generation models—ensuring temporal and semantic alignment across all three modalities. Contribution/Results: The resulting publicly released dataset spans diverse real-world scenarios and substantially advances performance on cross-modal retrieval and joint embedding learning tasks, achieving state-of-the-art results across multiple benchmarks. By providing a scalable, high-fidelity resource, this work establishes a new foundation for training and evaluating foundational multimodal models.

Creating high-quality audio-image-text datasetsEnsuring semantic connections between modalitiesExtracting multimodal data from videos

Machine learning (ML) suffers from weak data curation practices and insufficient documentation of ethical, environmental, and data management information. Method: We systematically evaluated 60 datasets from the NeurIPS Datasets and Benchmarks Track (2021–2023), introducing bibliometric data cataloging theory from library and information science to ML for the first time. We developed a literature-driven, four-dimensional evaluation framework—assessing documentation completeness, ethical impact, environmental footprint, and data management—and designed an actionable, structured rubric alongside an open-source assessment toolkit. Contribution/Results: We released the first exemplar metadata repository showcasing best practices. Our analysis revealed widespread deficiencies across all four dimensions. Based on these findings, we formulated actionable guidelines for conference reviewers and community adoption. All artifacts—including framework, rubric, toolkit, and metadata—are openly shared to advance ML datasets toward higher quality, reusability, and standardization.

Data CurationEthical ConsiderationsMachine Learning

To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.

Automate video annotation for computer vision modelsDevelop efficient video tracking and segmentation toolReduce time and resources for labeled data generation

Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation

Feb 12, 2025
JX
Jinda Xu
🏛️ Shanghai Jiao Tong University | HAOMO.AI Technology Co., Ltd.

Web-scraped datasets commonly suffer from low quality, high redundancy, and class imbalance, rendering existing heuristic filtering methods inadequate for modeling complex multimodal features—often introducing bias or erroneously discarding relevant samples. To address this, we propose EcoDatum, the first quality-centric multimodal collaborative filtering framework. Its core innovations are: (1) a quality-guided multimodal deduplication mechanism that jointly leverages visual, linguistic, and cross-modal embeddings for fine-grained similarity assessment; and (2) a weakly supervised ensemble optimization framework integrating automated hyperparameter search with multi-operator collaborative scoring. Evaluated on the DataComp benchmark, EcoDatum achieves a mean score of 0.182—outperforming prior baselines by 28%—and ranks first overall. Empirical results demonstrate substantial improvements in downstream model training efficiency and generalization performance.

Addressing unstructured dataset challengesEnhancing data curation efficiencyImproving model performance quality

Latest Papers

What's happening recently
View more

This study addresses the lack of domain priors in general-purpose data and the high cost of manual annotation by exploring synthetic data curation strategies for task-specific visual perception. Shifting focus from generating more data to determining what data to generate, this work proposes a task-oriented data curation paradigm. Methodologically, it systematically integrates three complementary paradigms—procedural rendering, physics-based simulation, and generative AI—leveraging limited real seed samples to learn sensor appearance characteristics and construct customized synthetic pipelines. Experiments demonstrate that this hybrid strategy achieves reliable Sim-to-Real transfer in surface defect detection, old photo restoration, and 6DoF pose estimation. These results validate that combining controllable supervision with appearance learning constitutes an effective pathway for enhancing model generalization capabilities.

Data CurationSim-to-Real TransferSynthetic Data

This study addresses a critical gap in machine learning education: the overreliance on pre-labeled datasets, which often obscures the subjectivity and ambiguity inherent in data annotation, leading students to place undue trust in model outputs. To counter this, the authors introduce an innovative pedagogical intervention that transforms manual annotation into an active learning tool. Students annotated hair coverage in skin lesion images using a three-point scale, followed by structured reflections via questionnaires. A cross-institutional experiment involving 43 participants from Fontys University of Applied Sciences (Netherlands) and the IT University of Copenhagen (Denmark) demonstrated that this approach significantly enhanced learners’ awareness of annotation ambiguity, dataset biases, and model limitations. Most participants acknowledged the influence of personal interpretation on labeling decisions and reported higher engagement compared to traditional instruction. This work provides the first empirical evidence supporting subjective annotation as an effective strategy for cultivating critical thinking about AI systems.

biasdata annotationinterpretive diversity

Oversampling and Downsampling with Core-Boundary Awareness: A Data Quality-Driven Approach

Sep 24, 2025
SB
Samir Brahim Belhaouari
🏛️ Hamad Bin Khalifa University | Maastricht University | KTH Royal Institute of Technology

To address the challenge in imbalanced classification where models struggle to distinguish boundary-critical samples from core-redundant ones, this paper proposes a core-boundary-aware data resampling framework. Methodologically, it is the first to systematically model data distribution geometry to adaptively identify decision-boundary-sensitive regions and high-density core regions, applying boundary-enhancing oversampling and core-compressing undersampling respectively. The framework is generalizable to text, multimodal, and self-supervised learning settings. Empirically, it achieves up to a 10% improvement in F1-score across 96% of benchmark datasets, while attaining a 90% data compression ratio without accuracy loss and accelerating training tenfold—significantly reducing computational overhead. Its core innovation lies in shifting sample importance assessment from class-frequency-driven heuristics to geometry-aware, boundary-sensitivity-driven evaluation.

Addresses imbalanced classification by differentiating boundary and core data instancesAims to improve computational efficiency in training like LLMs through quality-driven samplingProposes oversampling boundary data and reducing core data to enhance model performance

This work addresses the challenge that critical evidence in scientific literature is scattered across lengthy texts, tables, and figures, hindering existing agents from efficiently performing cross-modal structured extraction and reasoning. To overcome this limitation, the authors propose Beaver, a novel framework that integrates task scaffolding, multimodal evidence tools, and provenance tracking into an agent workflow, enabling auditable, phased autonomous research with iterative diagnose-and-correct cycles. Evaluated on the Gold-Referenced Attribute Score (GRAS), Beaver achieves 81.0—surpassing the current state-of-the-art agent by 23 percentage points—with particularly pronounced gains on high-value attributes requiring cross-modal reasoning.

cross-modal reasoningevidence integrationmultimodal sources

This study systematically investigates the key factors in data curation for multimodal reasoning under fixed model architectures and training protocols. Framed within the NeurIPS 2025 DCVLR Challenge, the work proposes a difficulty-aware sampling strategy grounded in alignment with foundational datasets and conducts ablation studies to assess the impact of data scale, diversity, and synthetic augmentation. The findings reveal that sample difficulty is the dominant driver of performance gains; merely increasing data volume reduces variance without necessarily improving accuracy, while data diversity and synthetic augmentation offer limited benefits. The proposed approach secured first place in the challenge, underscoring the critical role of alignment and difficulty-aware sampling in data-efficient multimodal reasoning.

data curationdata efficiencydataset selection

Hot Scholars

QZ

Qiquan Zhang

UNSW, Australia | NUS, Singapore | HIT, China
speech processingspeech enhancementaudio-visual learningNLP
SG

Stephan Getzmann

IfADo Leibniz Research Centre for Working Environments and Human Factors
QW

Qingbo Wu

University of Electronic Science and Technology of China
video codingimage and video quality assessment