Score
Designs, builds, and configures automated data-labeling systems and pipelines that produce ground-truth annotations using model-assisted or LLM-assisted workflows, including auto-labeling tools, labeling pipeline design, and labeling strategy development. Develops and evaluates components for multi-label and sequence labeling, noise-robust labeling methods, labeling protocols and quality-control procedures, and the tooling and integration required to automate and manage end-to-end labeling operations.
High annotation costs and prolonged turnaround times plague NLP development, necessitating efficient and reliable data labeling paradigms. This paper proposes an LLM-powered Human-in-the-Loop (HITL) hybrid annotation framework that systematically integrates synthetic data generation, active learning, and human-AI collaboration, augmented with built-in mechanisms for annotation quality assessment, annotator management, and cost-benefit analysis. Unlike prior work—largely theoretical or narrowly scoped—this study introduces the first deployable, plug-and-play industrial-grade annotation methodology, bridging the critical gap between methodological research and real-world engineering practice. Empirical validation across multiple production NLP projects demonstrates that the framework consistently reduces annotation costs and cycle time by 30–50%, while maintaining label quality within required thresholds.
This work systematically evaluates the efficacy and limitations of large language models (LLMs) as annotators for subjective tasks. We survey 12 prior studies and empirically compare opinion distribution alignment between GPT-series models and human annotators across four subjective datasets—introducing, for the first time, a novel evaluation paradigm centered on *perspective diversity alignment*. Results reveal substantial distributional biases in LLMs, including underrepresentation of minority viewpoints, prompt sensitivity, English-language bias, and embedded societal prejudices; most existing annotation methods overlook such distributional discrepancies, and only a few strategies effectively capture opinion diversity. Our findings expose critical reliability risks in deploying LLMs for subjective annotation and establish a reproducible statistical framework—comprising quantitative metrics and methodological guidelines—for assessing annotation quality in subjective NLP tasks.
Data cleaning remains highly manual, inefficient, and error-prone. This paper proposes the first goal-driven LLM-based framework for automatic workflow generation: given a dirty table and a target query, it end-to-end generates a minimal viable clean table along with executable cleaning steps—including deduplication, missing-value imputation, and format standardization. Our contributions are threefold: (1) We introduce the first benchmark dataset comprising annotated quadruples of (goal, dirty table, cleaning workflow, cleaned answer); (2) We design a zero-shot, multi-stage prompting framework—requiring no fine-tuning—that decomposes the task into goal column identification, data quality diagnosis, and operation-parameter generation; (3) We empirically validate that off-the-shelf LLMs possess inherent reasoning capabilities sufficient to generate high-quality, executable cleaning workflows across three major LLM families, significantly reducing human intervention.
Widespread label noise (10–25%) in NLP benchmark datasets leads to systematic underestimation of model performance, with many purported “LLM failures” attributable to annotation errors rather than model limitations. Method: We propose LLM-as-a-judge—a framework leveraging ensemble judgments from GPT-4, Claude, and Llama, combined with consistency voting and error-sensitivity analysis to automatically detect mislabeled instances; we further apply label smoothing and confident learning for robust label recalibration. Contribution/Results: Comprehensive evaluation across the TRUE benchmark suite reveals substantial disparities in quality and efficiency among expert, crowdsourced, and LLM-generated annotations. After correction, state-of-the-art models achieve average accuracy gains of 3.2–7.8 percentage points. This work provides the first empirical evidence of systematic label-noise interference in LLM evaluation and introduces a scalable, collaborative adjudication paradigm that reframes data correction as model performance recalibration.
Current computer vision research lacks an end-to-end system that automatically generates deployable models directly from natural language requirements. Method: This paper proposes a novel “request-to-model” paradigm, introducing AutoMMLab—the first open-source platform for end-to-end vision modeling—and the LAMP benchmark. It further designs HPO-LLaMA, an LLM-driven hyperparameter optimization algorithm that integrates natural language understanding, automated training and deployment pipelines, and a multi-stage evaluation framework. Contributions/Results: (1) First realization of a fully automated, language-instruction-driven vision modeling pipeline; (2) HPO-LLaMA achieves over 40% improvement in hyperparameter search efficiency across multiple CV tasks; (3) Comprehensive open-sourcing of datasets, code, and benchmarks to advance accessible, reproducible vision model development.
Conventional pretraining data annotation relying on black-box API calls faces critical bottlenecks—including prohibitively high invocation costs, non-editable outputs, and poor auditability. Method: This paper proposes a novel paradigm—“large language models (LLMs) generating executable annotation programs”—where LLMs synthesize Python-based annotation code that executes locally, enabling lightweight, iterative validation and refinement. Contribution/Results: The approach achieves annotation quality on par with or up to 12.9% higher than baseline methods across multiple tasks, while reducing total annotation cost by approximately 500×. It ensures high fidelity, reproducibility, transparency, and reusability—overcoming the limitations of static datasets and costly external API dependencies. To our knowledge, this is the first work to integrate program synthesis into the data annotation pipeline, establishing a new paradigm for efficient, controllable, and sustainable data engineering.
Existing AutoML systems rely heavily on expert configuration, resulting in low usability; although LLM-assisted approaches have emerged, they typically target isolated pipeline stages and fail to harness LLMs’ end-to-end reasoning capabilities. This paper introduces the first multi-agent large language model framework for full-stack AutoML—spanning data acquisition, preprocessing, model search, hyperparameter tuning, and deployment. Our approach innovatively integrates retrieval-augmented multi-stage planning, task-parallel decomposition, multi-stage program verification, and domain-adaptive prompt engineering to enable natural-language-driven fully automated machine learning. Evaluated across 14 diverse datasets and 7 downstream task categories, our framework achieves significantly higher end-to-end automation success rates. Generated models maintain high cross-domain performance, while human intervention is reduced by over 70%.
Existing LLM tool-use methods rely on static data pipelines, decoupling data generation from model training—hindering adaptive focus on model weaknesses and effective removal of noisy labels, thus impairing training efficiency. This paper introduces the first open-source, model-aware data evolution framework, establishing a closed-loop training paradigm comprising three tightly integrated modules: *capability diagnosis*, *label verification*, and *error-driven expansion*. It jointly optimizes data and model through iterative refinement: greedy capability probing identifies model deficiencies; discriminator-guided label verification purifies training data; and error feedback steers targeted data augmentation. The resulting 8B model achieves state-of-the-art performance on BFCL-v3 and ACEBench—surpassing same-scale SOTA models and even outperforming its 32B data generator—marking the first demonstration of data–model co-evolution within an open-source ecosystem.
Existing LLM-driven feature engineering methods are not designed for multi-label learning, thus failing to model label dependencies and lacking task-specificity. To address this, we propose FEAML—a novel framework that pioneers the integration of LLM-based code generation into multi-label settings. FEAML automatically constructs highly discriminative features by jointly leveraging metadata and label co-occurrence matrices. It introduces label-dependency-aware prompt engineering and a Pearson correlation-based redundancy detection mechanism, coupled with closed-loop optimization guided by classification accuracy. This yields an interpretable, low-redundancy, and self-optimizing feature generation paradigm. Extensive experiments on multiple standard multi-label benchmark datasets demonstrate that FEAML significantly outperforms conventional feature engineering approaches, achieving substantial average improvements in classification accuracy—thereby validating its effectiveness and generalizability.
Domain scientists often lack sufficient programming expertise to conduct data analysis efficiently. This paper addresses the low reliability and poor trustworthiness of large language models (LLMs) in scientific code generation by introducing the first benchmark suite for Python-based data analysis and visualization grounded in real-world research tasks. We propose three synergistic strategies: data-aware prompt disambiguation, retrieval-augmented prompt optimization, and iterative error repair—integrated with retrieval-augmented generation (RAG) and automated execution validation. Experiments demonstrate substantial improvements in code executability and functional correctness. However, domain-context understanding remains a critical bottleneck. This work contributes both a reusable, realistic evaluation benchmark and a systematic technical framework for developing trustworthy AI-powered scientific tools.
This study addresses the lack of systematic understanding regarding the practical usage patterns, reliability mechanisms, and autonomy levels of large language model (LLM) agents in low-code/no-code platforms. Drawing on over 6,000 publicly available n8n workflows, the authors employ large-scale data mining, structured log analysis, and qualitative coding to empirically characterize how LLM agents are deployed in real-world automation scenarios—specifically examining task distribution, workflow structure, tool invocation, and degrees of autonomy. The findings reveal that while LLMs are commonly embedded within complex workflows featuring control logic and human review steps, such workflows generally lack structured fault tolerance, repair loops, and approval mechanisms. Based on these insights, the study articulates ten empirical observations and five design implications to inform the development of more reliable and governable low-code platforms.