create absa dataset

Designs and builds annotated corpora for aspect-based sentiment analysis by collecting relevant texts, defining and applying annotation guidelines, marking aspect terms/targets, labeling sentiment polarity (and any aspect-level attributes) per aspect, and organizing the dataset with quality checks and class balance for model training and evaluation.

createabsadataset

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the scarcity of high-quality annotated data for German aspect-based sentiment analysis (ABSA) and the unclear impact of annotation sources on model performance. It presents the first systematic comparison of annotation quality among experts, students, crowdworkers, and large language models (LLMs) in the German ABSA context. The authors construct a gold-standard dataset through expert re-annotation and evaluate the effectiveness of each annotation type on two core tasks: aspect category sentiment analysis (ACSA) and aspect term and sentiment detection (TASD). Leveraging state-of-the-art models—including BERT, T5, and LLaMA—with both fine-tuning and instruction-based prompting, the experiments demonstrate that expert annotations yield significantly higher consistency and downstream task performance. The study also quantifies the trade-offs of using LLM-generated and non-expert annotations under resource-constrained conditions, highlighting their practical feasibility alongside inherent limitations.

Annotation QualityAspect-Based Sentiment AnalysisGerman NLP

Implicit aspect extraction in Aspect-Based Sentiment Analysis (ABSA) remains challenging due to the scarcity of real-world annotated data, particularly for emerging domains. Method: This study investigates the capability boundaries of large language models (LLMs) on implicit aspect extraction in a novel domain—sports—and introduces the first synthetic, generation-oriented dataset tailored for generative LLMs. It reformulates ABSA as a structured generation task and proposes a novel aspect–polarity pair evaluation metric. Systematic zero-shot and few-shot evaluations are conducted on open-source LLMs (e.g., LLaMA, Phi). Contribution/Results: Results demonstrate that LLMs possess preliminary capacity for implicit aspect identification, yet exhibit limited cross-domain generalization. The proposed metric significantly enhances comparability and interpretability of generated outputs. This work establishes a new paradigm for data construction and evaluation in low-resource ABSA, bridging the gap between generative modeling and fine-grained sentiment analysis.

Assessing LLMs' performance with synthetic sports dataEvaluating LLMs for aspect extraction in ABSAProposing metrics for generative model evaluation

This work addresses the limitations of existing aspect-based sentiment analysis (ABSA) annotation tools, which lack support for the full task spectrum, offer limited customizability, and fail to provide context-aware intelligent assistance. To overcome these challenges, we present an open-source web-based annotation platform that, for the first time, enables flexible configuration across the entire ABSA task spectrum. The system innovatively integrates large language models (LLMs) with retrieval-augmented generation (RAG) to deliver dynamic, context-aware suggestions during human-in-the-loop annotation by leveraging semantically similar previously annotated examples. Through few-shot prompting and real-time retrieval, the platform continuously refines suggestion quality as annotation progresses. Released under the MIT license, this tool significantly enhances annotation efficiency and consistency, making it suitable for both academic research and industrial applications.

Annotation ToolAspect-Based Sentiment AnalysisHuman-in-the-Loop

Czech Dataset for Complex Aspect-Based Sentiment Analysis Tasks

Aug 11, 2025
JŠ
Jakub Šmíd
🏛️ University of West Bohemia | University of West Bohemia

Existing Czech ABSA datasets support only aspect term extraction or polarity classification, lacking unified annotations for joint target–aspect–category detection. To address this gap, we introduce the first Czech fine-grained ABSA dataset conforming to the SemEval-2016 format, comprising 3.1K manually double-annotated restaurant reviews with 90% inter-annotator agreement—enabling standardized modeling of complex Czech ABSA tasks for the first time. We further release 24 million unlabeled Czech reviews to facilitate unsupervised learning. Extensive experiments with Transformer-based models establish strong baselines across all subtasks. All resources—including code, annotated data, and detailed annotation guidelines—are publicly available under a non-commercial academic license. This work significantly advances cross-lingual ABSA benchmarking and low-resource sentiment analysis, particularly for morphologically rich languages like Czech.

Enables cross-lingual comparisons with unified annotation formatIntroduces Czech dataset for complex aspect-based sentiment analysisProvides annotated and unannotated reviews for diverse learning approaches

This study addresses the lack of systematic approaches for constructing, storing, and sharing high-quality annotated corpora. It proposes a generalizable and reusable end-to-end methodology encompassing annotation guideline development, corpus annotation, data storage, sharing mechanisms, and value realization, with an emphasis on full lifecycle management and cross-domain applicability. Integrating linguistic annotation theory, data management standards, and collaborative research practices, the approach is articulated through a structured framework and illustrative examples to yield a clear and actionable guide. The resulting methodology provides standardized support for diverse research domains, significantly enhancing the efficiency and quality with which researchers can build and utilize annotated textual data.

annotated corpusannotation guidelinescorpus creation

Latest Papers

What's happening recently
View more

This study addresses the scarcity of high-quality annotated data for fine-grained aspect-based sentiment analysis in Korean e-commerce reviews by introducing EVAD, a novel Korean fashion-domain review dataset. The authors propose an extended ABSA framework that supports unary, binary, and multi-value aspect classification, innovatively modeling aspect values according to their value types. They achieve efficient and precise fine-grained annotation by integrating semi-supervised symbol propagation (SSP) with linguistic resources based on finite-state transducers (FSTs). Experimental results demonstrate the quality of the dataset and effectiveness of the approach: KoBERT and KcBERT models trained on EVAD attain F1 scores of 0.88 and 0.90, respectively, on the aspect–value pair extraction task.

Aspect-Based Sentiment Analysisaspect-value pairse-commerce reviews

This work addresses the challenge of acquiring high-quality annotated data for fine-grained opinion analysis tasks—such as Aspect Sentiment Triplet Extraction (ASTE) and Aspect-Category Opinion-Sentiment (ACOS)—which are hindered by high annotation costs and substantial human effort, particularly in multi-domain settings. To mitigate these limitations, the authors propose an automated labeling and arbitration framework grounded in large language models (LLMs), integrated with a declarative annotation pipeline. This approach significantly reduces inconsistencies arising from manual prompt engineering while achieving high inter-annotator agreement on both ASTE and ACOS tasks. By minimizing reliance on human annotators and lowering data construction costs, the method enhances the reliability and scalability of cross-model annotations, thereby facilitating broader practical deployment across diverse domains.

annotation costdomain-specific datasetsfine-grained opinion analysis

This study addresses the scarcity of aspect-based sentiment analysis (ABSA) resources for low-resource languages like Czech, particularly the lack of datasets annotated with opinion terms and effective cross-lingual transfer methods. To bridge this gap, the authors construct the first Czech ABSA benchmark dataset in the restaurant domain, featuring opinion-term annotations and supporting three levels of task complexity. They systematically evaluate a range of Transformer and large language models (LLMs) under monolingual, cross-lingual, and multilingual settings, and propose an LLM-driven translation–label alignment strategy to enhance cross-lingual transfer. Experimental results demonstrate that the proposed approach significantly outperforms baseline methods on Czech ABSA, while also revealing limitations of current models in capturing fine-grained opinion expressions, thereby establishing a new benchmark for sentiment analysis in low-resource languages.

aspect-based sentiment analysiscross-lingual challengesCzech

This work addresses the high cost of manual annotation in complex aspect-based sentiment analysis (ABSA) tasks such as aspect sentiment quadruple prediction (ASQP). To mitigate this, the authors propose LA-ABSA, a framework that systematically leverages large language models (LLMs) as annotators, guided by only a few human-provided examples through in-context learning to generate high-quality labeled data for fine-tuning lightweight downstream models. This approach substantially reduces reliance on both extensive human annotation and expensive LLM inference. Evaluated on five benchmark datasets—including SemEval Rest16—the method achieves strong performance, attaining an F1 score of 49.85 on ASQP, which closely approaches the in-context learning performance of Gemma-3-27B (51.10) while significantly lowering computational overhead.

Aspect-Based Sentiment Analysisdata annotation costlow-resource scenarios

This study addresses the lack of efficient, structured collaborative annotation tools in existing Aspect-Based Sentiment Analysis (ABSA) research, which often necessitates cumbersome manual processing for data integration, relation reconstruction, and inter-annotator agreement (IAA) computation. To overcome these limitations, this work proposes the first web-based collaborative annotation platform supporting four ABSA subtasks, featuring an innovative integration of multi-task collaborative annotation and an automated ETL pipeline. The system automatically aligns annotations upon export, preserves character-level positions and dual-span offsets, and computes IAA metrics in real time. Validation on 1,002 restaurant reviews demonstrates a median annotation time of 31.58 seconds per instance, with raw IAA scores ranging from 0.78 to 0.86 across tasks, enabling the direct generation of high-quality, ready-to-use training data and substantially improving both annotation efficiency and structural integrity.

Annotation ToolAspect-Based Sentiment AnalysisCollaborative Annotation