Score
Design and build data-collection and curation processes that deliberately identify, gather, annotate, and integrate rare, ambiguous, adversarial, or boundary-condition examples across multiple modalities (e.g., text, images, audio, video, sensors). Create sampling strategies, annotation guidelines, quality-control checks, metadata and test splits or benchmarks that document and enable evaluation of model behavior specifically on these dataset edge cases.
Existing dataset characterization methods—statistical, structural, and model-driven—lack sufficient interpretability and deep structural insight. To address this, we propose a novel tensor-based representation paradigm that transcends conventional two-dimensional assumptions. Our approach leverages high-order tensor decomposition, multilinear modeling, and cross-modal joint representation to explicitly capture high-dimensional, nonlinear, and multi-source relational structures inherent in complex data. Extensive experiments demonstrate that the proposed method significantly outperforms baseline approaches in three key aspects: (i) disentangling intricate data structures, (ii) enhancing feature interpretability, and (iii) enabling traceable downstream task reasoning. This work establishes a unified tensor modeling framework for dataset representation and pioneers a data-driven discovery pathway tailored for explainable AI. It contributes both theoretical advances—through formalizing multilinear structure learning—and practical utility—by providing an interpretable, computationally grounded toolkit for transparent data analysis.
A critical shortage exists of high-quality, large-scale, semantically aligned acoustic–visual–textual multimodal datasets. Method: This paper proposes an end-to-end framework for constructing video-based multimodal data, integrating three key components: (i) video content filtering, (ii) cross-modal synchronization triplet extraction (audio–frame–subtitle), and (iii) fine-grained description synthesis leveraging image-to-text generation models—ensuring temporal and semantic alignment across all three modalities. Contribution/Results: The resulting publicly released dataset spans diverse real-world scenarios and substantially advances performance on cross-modal retrieval and joint embedding learning tasks, achieving state-of-the-art results across multiple benchmarks. By providing a scalable, high-fidelity resource, this work establishes a new foundation for training and evaluating foundational multimodal models.
Machine learning (ML) suffers from weak data curation practices and insufficient documentation of ethical, environmental, and data management information. Method: We systematically evaluated 60 datasets from the NeurIPS Datasets and Benchmarks Track (2021–2023), introducing bibliometric data cataloging theory from library and information science to ML for the first time. We developed a literature-driven, four-dimensional evaluation framework—assessing documentation completeness, ethical impact, environmental footprint, and data management—and designed an actionable, structured rubric alongside an open-source assessment toolkit. Contribution/Results: We released the first exemplar metadata repository showcasing best practices. Our analysis revealed widespread deficiencies across all four dimensions. Based on these findings, we formulated actionable guidelines for conference reviewers and community adoption. All artifacts—including framework, rubric, toolkit, and metadata—are openly shared to advance ML datasets toward higher quality, reusability, and standardization.
本文介绍NeMo数据设计器,一种用于生成多模态合成数据的开源框架,通过声明式配置和插件系统提高数据集多样性和工作流可重复性。
To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.
Web-scraped datasets commonly suffer from low quality, high redundancy, and class imbalance, rendering existing heuristic filtering methods inadequate for modeling complex multimodal features—often introducing bias or erroneously discarding relevant samples. To address this, we propose EcoDatum, the first quality-centric multimodal collaborative filtering framework. Its core innovations are: (1) a quality-guided multimodal deduplication mechanism that jointly leverages visual, linguistic, and cross-modal embeddings for fine-grained similarity assessment; and (2) a weakly supervised ensemble optimization framework integrating automated hyperparameter search with multi-operator collaborative scoring. Evaluated on the DataComp benchmark, EcoDatum achieves a mean score of 0.182—outperforming prior baselines by 28%—and ranks first overall. Empirical results demonstrate substantial improvements in downstream model training efficiency and generalization performance.
This study addresses the lack of domain priors in general-purpose data and the high cost of manual annotation by exploring synthetic data curation strategies for task-specific visual perception. Shifting focus from generating more data to determining what data to generate, this work proposes a task-oriented data curation paradigm. Methodologically, it systematically integrates three complementary paradigms—procedural rendering, physics-based simulation, and generative AI—leveraging limited real seed samples to learn sensor appearance characteristics and construct customized synthetic pipelines. Experiments demonstrate that this hybrid strategy achieves reliable Sim-to-Real transfer in surface defect detection, old photo restoration, and 6DoF pose estimation. These results validate that combining controllable supervision with appearance learning constitutes an effective pathway for enhancing model generalization capabilities.
This study addresses a critical gap in machine learning education: the overreliance on pre-labeled datasets, which often obscures the subjectivity and ambiguity inherent in data annotation, leading students to place undue trust in model outputs. To counter this, the authors introduce an innovative pedagogical intervention that transforms manual annotation into an active learning tool. Students annotated hair coverage in skin lesion images using a three-point scale, followed by structured reflections via questionnaires. A cross-institutional experiment involving 43 participants from Fontys University of Applied Sciences (Netherlands) and the IT University of Copenhagen (Denmark) demonstrated that this approach significantly enhanced learners’ awareness of annotation ambiguity, dataset biases, and model limitations. Most participants acknowledged the influence of personal interpretation on labeling decisions and reported higher engagement compared to traditional instruction. This work provides the first empirical evidence supporting subjective annotation as an effective strategy for cultivating critical thinking about AI systems.
To address the challenge in imbalanced classification where models struggle to distinguish boundary-critical samples from core-redundant ones, this paper proposes a core-boundary-aware data resampling framework. Methodologically, it is the first to systematically model data distribution geometry to adaptively identify decision-boundary-sensitive regions and high-density core regions, applying boundary-enhancing oversampling and core-compressing undersampling respectively. The framework is generalizable to text, multimodal, and self-supervised learning settings. Empirically, it achieves up to a 10% improvement in F1-score across 96% of benchmark datasets, while attaining a 90% data compression ratio without accuracy loss and accelerating training tenfold—significantly reducing computational overhead. Its core innovation lies in shifting sample importance assessment from class-frequency-driven heuristics to geometry-aware, boundary-sensitivity-driven evaluation.
This work addresses the challenge that critical evidence in scientific literature is scattered across lengthy texts, tables, and figures, hindering existing agents from efficiently performing cross-modal structured extraction and reasoning. To overcome this limitation, the authors propose Beaver, a novel framework that integrates task scaffolding, multimodal evidence tools, and provenance tracking into an agent workflow, enabling auditable, phased autonomous research with iterative diagnose-and-correct cycles. Evaluated on the Gold-Referenced Attribute Score (GRAS), Beaver achieves 81.0—surpassing the current state-of-the-art agent by 23 percentage points—with particularly pronounced gains on high-value attributes requiring cross-modal reasoning.
This study systematically investigates the key factors in data curation for multimodal reasoning under fixed model architectures and training protocols. Framed within the NeurIPS 2025 DCVLR Challenge, the work proposes a difficulty-aware sampling strategy grounded in alignment with foundational datasets and conducts ablation studies to assess the impact of data scale, diversity, and synthetic augmentation. The findings reveal that sample difficulty is the dominant driver of performance gains; merely increasing data volume reduces variance without necessarily improving accuracy, while data diversity and synthetic augmentation offer limited benefits. The proposed approach secured first place in the challenge, underscoring the critical role of alignment and difficulty-aware sampling in data-efficient multimodal reasoning.