Score
Designs and implements annotation schemes, labeling tools, and data pipelines for assigning and managing hierarchical (two-level) category labels on visual content. Builds protocols and interfaces for collecting and reconciling multimodal annotations (e.g., aligning image and non-image signals), and analyzes label quality, inter-annotator agreement, and consistency across the two category levels and modalities.
This study addresses the inconsistency in human annotation caused by ambiguous category definitions in traditional content moderation. To resolve this, the authors propose an AI-driven constitutional annotation framework: large language models first assist humans in formulating structured, interpretable category “constitutions,” which then guide automated dual-axis labeling of intent and content safety. This approach shifts human effort from case-by-case judgments to high-level semantic definition. Evaluated on harassment, hate speech, and non-violent criminal conduct tasks, the method reduces cross-model annotation inconsistency by up to 57-fold compared to conventional paragraph-based rules and effectively exposes latent gaps in existing policy formulations.
Scientific image annotation projects face cross-domain managerial challenges—including scarce data acquisition, inefficient resource allocation, inadequate annotator training, and pronounced human bias. To address these issues, this paper proposes the first general-purpose framework for preparing scientific image annotation projects. The framework systematically integrates objective definition, data availability assessment, multi-role team configuration, bias mitigation strategies, and an iterative annotator training mechanism, complemented by a recommended toolchain supporting integrated project management, quality control, and collaborative annotation. A novel closed-loop workflow—comprising bias detection, feedback integration, and retraining—is introduced to significantly enhance annotation consistency and efficiency. Empirical evaluation across multiple disciplines demonstrates that the framework reduces annotation costs by over 20%, improves project success rates, and strengthens knowledge base construction quality—thereby filling a critical research gap in standardized preparation guidelines for complex scientific image annotation.
High annotation costs and prolonged turnaround times plague NLP development, necessitating efficient and reliable data labeling paradigms. This paper proposes an LLM-powered Human-in-the-Loop (HITL) hybrid annotation framework that systematically integrates synthetic data generation, active learning, and human-AI collaboration, augmented with built-in mechanisms for annotation quality assessment, annotator management, and cost-benefit analysis. Unlike prior work—largely theoretical or narrowly scoped—this study introduces the first deployable, plug-and-play industrial-grade annotation methodology, bridging the critical gap between methodological research and real-world engineering practice. Empirical validation across multiple production NLP projects demonstrates that the framework consistently reduces annotation costs and cycle time by 30–50%, while maintaining label quality within required thresholds.
NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.
Existing element-attribute grid representations for graphic design completion tasks struggle to model variable-length, type-heterogeneous, and multimodal (text-image) structures. Method: We propose a unified interleaved multimodal tokenized document model that jointly encodes syntactic and semantic structures of markup languages (e.g., SVG/HTML) alongside variable-size, alpha-channel-aware local image generation. We introduce a specialized image quantizer for efficient transparent-image tokenization and integrate an enhanced code-large language model with an interleaved multimodal sequence architecture. Contribution/Results: Our model achieves significant improvements over baselines on three design completion tasks—missing template attributes, image synthesis, and text generation—demonstrating its effectiveness in jointly modeling structural logic and visual semantics in design documents.
This study addresses the common yet often overlooked issue of subjective disagreement in multi-label sentiment annotation, which traditional approaches typically treat as noise and discard along with its underlying structural information. To better capture annotator uncertainty, the work proposes replacing hard majority voting with soft labels derived from vote proportions and intensity-weighted confidence, and introduces Soft Bernoulli Cross-Entropy (SoftBCE) for soft-supervised model training. Additionally, it incorporates a probabilistic alignment metric for evaluation and a data-driven diagnostic framework to analyze annotation discrepancies. Experimental results show that while hard labels yield marginally higher F1 scores, soft labels more faithfully represent the inherent uncertainty among annotators. This research establishes a novel paradigm and offers practical guidance for label aggregation, model training, and evaluation in multi-label sentiment analysis.
This work addresses the challenges of ambiguous many-to-many mappings and subjective annotations arising from the semantic gap between visual content and textual descriptions in object recognition datasets. To mitigate these issues, the authors propose an interactive crowdsourcing framework that integrates knowledge representation, natural language processing, and computer vision. The framework dynamically generates guided questions based on a predefined category hierarchy and visual attribute constraints, iteratively refining annotations through worker feedback. Experimental results demonstrate that the proposed approach significantly improves annotation consistency and quality, effectively narrows the semantic gap, and enhances the overall design of the crowdsourcing workflow.
This work addresses the challenge of chart annotation generation, which requires integrated understanding of chart semantics, inference of communicative intent, and generation of appropriate textual or graphical elements. Despite the growing capabilities of multimodal large language models (MLLMs), their performance on this task has lacked systematic evaluation. To bridge this gap, we introduce ChartAnno, the first benchmark comprising 1,200 real-world charts paired with their source code and three-tiered instructions, supporting three input modalities: code-only, code-plus-image, and image-only. Through comprehensive automatic evaluation and ablation studies, we find that current MLLMs still face significant bottlenecks in abstract intent reasoning. While proprietary models generally outperform open-source counterparts, the performance gap is narrowing. Moreover, concrete instructions substantially improve annotation quality, and image inputs provide modest gains on design-oriented metrics.
Current visual language model (VLM) training is hindered by the scarcity of high-quality annotated data that jointly captures unified spatial coordinates, open-vocabulary semantics, structural attributes, and topological relationships. Moreover, conventional annotation tools suffer from limited expressiveness, a disconnect between annotation and training pipelines, and poor reusability. To address these challenges, this work proposes ScreenAnnotator, which introduces a unified atomic annotation schema and integrates an online policy-based annotation loop with an embedded Bayesian verifier alongside a template-driven multitask data synthesis mechanism. This framework enables efficient and reusable construction of visual reasoning datasets. Evaluated on flowchart and GUI screenshot annotations, the system achieves acceptance rates of 99.7% and 77%, respectively, with consistently decreasing per-image annotation time. Fine-tuning VLMs on the generated data yields a 76.1% accuracy on flowchart understanding tasks, representing an absolute improvement of 35.1 percentage points.