automated data labeling

Designs, builds, and configures automated data-labeling systems and pipelines that produce ground-truth annotations using model-assisted or LLM-assisted workflows, including auto-labeling tools, labeling pipeline design, and labeling strategy development. Develops and evaluates components for multi-label and sequence labeling, noise-robust labeling methods, labeling protocols and quality-control procedures, and the tooling and integration required to automate and manage end-to-end labeling operations.

automateddatalabeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$216K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation

May 02, 2024
MP
Maja Pavlovic
🏛️ Queen Mary University of London | University of Utrecht

This work systematically evaluates the efficacy and limitations of large language models (LLMs) as annotators for subjective tasks. We survey 12 prior studies and empirically compare opinion distribution alignment between GPT-series models and human annotators across four subjective datasets—introducing, for the first time, a novel evaluation paradigm centered on *perspective diversity alignment*. Results reveal substantial distributional biases in LLMs, including underrepresentation of minority viewpoints, prompt sensitivity, English-language bias, and embedded societal prejudices; most existing annotation methods overlook such distributional discrepancies, and only a few strategies effectively capture opinion diversity. Our findings expose critical reliability risks in deploying LLMs for subjective annotation and establish a reproducible statistical framework—comprising quantitative metrics and methodological guidelines—for assessing annotation quality in subjective NLP tasks.

Addressing limitations like bias and prompt sensitivityComparing human and GPT-generated opinion distributionsEvaluating LLMs' effectiveness in data annotation tasks

Must-Read Papers

Most classic and influential ideas
View more

AutoDCWorkflow: LLM-based Data Cleaning Workflow Auto-Generation and Benchmark

Dec 09, 2024
LL
Lan Li
🏛️ University of Illinois, Urbana-Champaign

Data cleaning remains highly manual, inefficient, and error-prone. This paper proposes the first goal-driven LLM-based framework for automatic workflow generation: given a dirty table and a target query, it end-to-end generates a minimal viable clean table along with executable cleaning steps—including deduplication, missing-value imputation, and format standardization. Our contributions are threefold: (1) We introduce the first benchmark dataset comprising annotated quadruples of (goal, dirty table, cleaning workflow, cleaned answer); (2) We design a zero-shot, multi-stage prompting framework—requiring no fine-tuning—that decomposes the task into goal column identification, data quality diagnosis, and operation-parameter generation; (3) We empirically validate that off-the-shelf LLMs possess inherent reasoning capabilities sufficient to generate high-quality, executable cleaning workflows across three major LLM families, significantly reducing human intervention.

Addressing format inconsistencies, type errors, duplicates in datasetsAutomating data cleaning workflow generation using LLMsEvaluating workflow quality against human-curated benchmarks

Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance

Oct 24, 2024
ON
Omer Nahum
🏛️ Technion - Institute of Technology | Google Research

Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.

Comparing annotation quality from experts, crowdsourcing, and LLMsDetecting label errors in NLP benchmark datasetsMitigating mislabeled data effects on model performance

AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks

Feb 23, 2024
ZY
Zekang Yang
🏛️ SenseTime | Tetras.AI | The University of Hong Kong

Current computer vision research lacks an end-to-end system that automatically generates deployable models directly from natural language requirements. Method: This paper proposes a novel “request-to-model” paradigm, introducing AutoMMLab—the first open-source platform for end-to-end vision modeling—and the LAMP benchmark. It further designs HPO-LLaMA, an LLM-driven hyperparameter optimization algorithm that integrates natural language understanding, automated training and deployment pipelines, and a multi-stage evaluation framework. Contributions/Results: (1) First realization of a fully automated, language-instruction-driven vision modeling pipeline; (2) HPO-LLaMA achieves over 40% improvement in hyperparameter search efficiency across multiple CV tasks; (3) Comprehensive open-sourcing of datasets, code, and benchmarks to advance accessible, reproducible vision model development.

Automatic Model GenerationComputer VisionNatural Language Processing

The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators

Jun 25, 2024
TH
Tzu-Heng Huang
🏛️ University of Wisconsin-Madison

Conventional pretraining data annotation relying on black-box API calls faces critical bottlenecks—including prohibitively high invocation costs, non-editable outputs, and poor auditability. Method: This paper proposes a novel paradigm—“large language models (LLMs) generating executable annotation programs”—where LLMs synthesize Python-based annotation code that executes locally, enabling lightweight, iterative validation and refinement. Contribution/Results: The approach achieves annotation quality on par with or up to 12.9% higher than baseline methods across multiple tasks, while reducing total annotation cost by approximately 500×. It ensures high fidelity, reproducibility, transparency, and reusability—overcoming the limitations of static datasets and costly external API dependencies. To our knowledge, this is the first work to integrate program synthesis into the data annotation pipeline, establishing a new paradigm for efficient, controllable, and sustainable data engineering.

Cost-efficiencyData annotationPre-trained models

AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML

Oct 03, 2024
PT
Patara Trirat
🏛️ KAIST | DeepAuto.ai

Existing AutoML systems rely heavily on expert configuration, resulting in low usability; although LLM-assisted approaches have emerged, they typically target isolated pipeline stages and fail to harness LLMs’ end-to-end reasoning capabilities. This paper introduces the first multi-agent large language model framework for full-stack AutoML—spanning data acquisition, preprocessing, model search, hyperparameter tuning, and deployment. Our approach innovatively integrates retrieval-augmented multi-stage planning, task-parallel decomposition, multi-stage program verification, and domain-adaptive prompt engineering to enable natural-language-driven fully automated machine learning. Evaluated across 14 diverse datasets and 7 downstream task categories, our framework achieves significantly higher end-to-end automation success rates. Generated models maintain high cross-domain performance, while human intervention is reduced by over 70%.

Automates full-pipeline AutoML for non-expertsEnhances exploration with retrieval-augmented planningImproves efficiency via multi-agent parallel task execution

Latest Papers

What's happening recently
View more

LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls

Nov 12, 2025
KZ
Kangning Zhang
🏛️ Shanghai Jiao Tong University | Xiaohongshu Inc.

Existing LLM tool-use methods rely on static data pipelines, decoupling data generation from model training—hindering adaptive focus on model weaknesses and effective removal of noisy labels, thus impairing training efficiency. This paper introduces the first open-source, model-aware data evolution framework, establishing a closed-loop training paradigm comprising three tightly integrated modules: *capability diagnosis*, *label verification*, and *error-driven expansion*. It jointly optimizes data and model through iterative refinement: greedy capability probing identifies model deficiencies; discriminator-guided label verification purifies training data; and error feedback steers targeted data augmentation. The resulting 8B model achieves state-of-the-art performance on BFCL-v3 and ACEBench—surpassing same-scale SOTA models and even outperforming its 32B data generator—marking the first demonstration of data–model co-evolution within an open-source ecosystem.

Addressing static synthetic data pipelines with non-interactive processesClosing the data-training loop for robust LLM tool callsCorrecting noisy labels and focusing on model weaknesses adaptively

Existing LLM-driven feature engineering methods are not designed for multi-label learning, thus failing to model label dependencies and lacking task-specificity. To address this, we propose FEAML—a novel framework that pioneers the integration of LLM-based code generation into multi-label settings. FEAML automatically constructs highly discriminative features by jointly leveraging metadata and label co-occurrence matrices. It introduces label-dependency-aware prompt engineering and a Pearson correlation-based redundancy detection mechanism, coupled with closed-loop optimization guided by classification accuracy. This yields an interpretable, low-redundancy, and self-optimizing feature generation paradigm. Extensive experiments on multiple standard multi-label benchmark datasets demonstrate that FEAML significantly outperforms conventional feature engineering approaches, achieving substantial average improvements in classification accuracy—thereby validating its effectiveness and generalizability.

FEAML automates feature engineering for multi-label classification tasks.It models label dependencies using metadata and co-occurrence matrices.The method integrates feedback to optimize LLM-generated features iteratively.

Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code

Nov 26, 2025
AK
Apu Kumar Chakroborti
🏛️ Georgia State University

Domain scientists often lack sufficient programming expertise to conduct data analysis efficiently. This paper addresses the low reliability and poor trustworthiness of large language models (LLMs) in scientific code generation by introducing the first benchmark suite for Python-based data analysis and visualization grounded in real-world research tasks. We propose three synergistic strategies: data-aware prompt disambiguation, retrieval-augmented prompt optimization, and iterative error repair—integrated with retrieval-augmented generation (RAG) and automated execution validation. Experiments demonstrate substantial improvements in code executability and functional correctness. However, domain-context understanding remains a critical bottleneck. This work contributes both a reusable, realistic evaluation benchmark and a systematic technical framework for developing trustworthy AI-powered scientific tools.

LLMs generate code for scientific data analysis and visualizationStrategies improve code reliability but need further refinementTrustworthiness of LLM-generated code is limited without human intervention

This study addresses the lack of systematic understanding regarding the practical usage patterns, reliability mechanisms, and autonomy levels of large language model (LLM) agents in low-code/no-code platforms. Drawing on over 6,000 publicly available n8n workflows, the authors employ large-scale data mining, structured log analysis, and qualitative coding to empirically characterize how LLM agents are deployed in real-world automation scenarios—specifically examining task distribution, workflow structure, tool invocation, and degrees of autonomy. The findings reveal that while LLMs are commonly embedded within complex workflows featuring control logic and human review steps, such workflows generally lack structured fault tolerance, repair loops, and approval mechanisms. Based on these insights, the study articulates ten empirical observations and five design implications to inform the development of more reliable and governable low-code platforms.

Agentic WorkflowsHuman-AI CollaborationLarge Language Models

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
JS

Jing Shao

Research Scientist, Shanghai AI Laboratory/Shanghai Jiao Tong University
Computer VisionMulti-Modal Large Language Model
CX

Chaowei Xiao

University of Wisconsin - Madison/NVIDIA
Trustworthy Machine LearningAdversarial Machine LearningAI SafetyRobust AI
WL

Weihua Luo

Alibaba
natural language processingmachine learningartificial intelligence
LW

Longyue Wang

Alibaba International
Large Language ModelMachine TranslationNatural Language ProcessingLanguange Agent