Institution profile

Universidade da Coruña

Academic institutioneurope · es
Official website
Research library90linked papers
Opportunities0open roles
Selected work

Representative Papers

A Syntax-Injected Approach for Faster and More Accurate Sentiment Analysis

Jun 21, 2024arXiv.org

In sentiment analysis (SA), conventional dependency parsing enhances accuracy and interpretability but incurs prohibitive computational overhead, hindering practical deployment. This paper proposes the Sequence Labeling Syntactic Parser (SELSP), the first approach to formulate dependency parsing as a lightweight sequence labeling task, enabling efficient syntactic integration. SELSP incorporates a polarity-aware sentiment lexicon, employs ternary and quinary classification schemes, and is rigorously evaluated via multi-model ablation studies. Compared to Stanza and VADER, SELSP achieves significant gains in both accuracy and inference speed; against Transformer-based baselines, it accelerates inference by multiple orders of magnitude while retaining competitive performance on ternary sentiment classification. Key contributions are: (i) pioneering a sequence labeling paradigm for dependency parsing; (ii) empirically validating that sentiment lexicons grounded in polarity discrimination differences yield superior performance; and (iii) achieving a balanced optimization of inference speed, predictive accuracy, and model interpretability.

1 citationsRead paper

Structured Sentiment Analysis Using Sequence Labeling as Dependency Graph Parsing

Oct 08, 2026

This study addresses the challenge of fine-grained sentiment analysis, which requires precisely identifying complex relationships among holders, targets, and sentiment expressions. Traditional structured parsing methods rely on intricate graph models with high computational costs. Departing from conventional paradigms, this work reformulates structured sentiment analysis as a dependency graph parsing task. It introduces a linearized graph encoding strategy that solves graph structure prediction directly through sequence labeling models, thereby circumventing complex decoding procedures. The proposed lightweight architecture enables efficient inference while achieving strong performance across five languages and seven benchmark datasets. Its results are comparable to, or even surpass, those of existing complex single-task models, offering a concise and effective unified solution for cross-lingual fine-grained sentiment analysis.

0 citationsRead paper

I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers

Oct 07, 2026

This study investigates the proliferation of contrastive constructions, such as "rather than," in NLP papers generated by large language models (LLMs), which introduces content redundancy and elicits reviewer dissatisfaction. Through text mining, human annotation, preference data analysis, and reward model evaluation, this work systematically examines the causes underlying this surge in LLM-assisted academic writing. Our findings reveal that post-training mechanisms based on human preference data inadvertently induce the overuse of contrastive structures as a side effect, and we quantify their detrimental impact on reader experience. Notably, the frequency of such constructions in 2026 is seven times that of 2019, with the majority judged as low-quality expressions that degrade academic rigor.

0 citationsRead paper

Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming

Oct 07, 2026

This study addresses the reliance on expert knowledge in production scheduling modeling and the high code hallucination rates of general-purpose large language models (LLMs), which hinder practical deployment. To overcome these challenges, this work proposes an automated modeling framework integrating a multi-agent architecture with retrieval-augmented generation. Leveraging fine-tuning-free LLM agents, the method dynamically retrieves solver documentation via the Model Context Protocol to automatically translate natural language descriptions into executable constraint programming code, effectively bridging modeling and AI-driven decision-making. Experimental results demonstrate that the single-run success rate of generated scripts improves from 14.8% to 59.3%, reaching 80.6% for medium-complexity problems. These findings significantly mitigate hallucinations and validate the feasibility of LLM-assisted modeling.

0 citationsRead paper

Argument Structure Prediction in Online Conversations: A Comparative Study of Modeling Paradigms and Task Architectures

Sep 30, 2026

This study addresses the unclear trade-offs among modeling paradigms and task architectures for online conversational argument structure prediction. We present the first multi-paradigm comparative analysis under strict constraints, developing a dedicated data processing pipeline that transforms Informational Argumentation Trees (IAT) into bipolar argument structures. Furthermore, this work systematically evaluates the end-to-end performance and efficiency of supervised fine-tuning and large language models within both single-step and multi-step architectures. Our findings reveal that relation identification constitutes the primary bottleneck and delineate the respective advantages and limitations of each paradigm. To facilitate future research, we release the complete framework as open source.

0 citationsRead paper
Recent publications

Latest Papers

Structured Sentiment Analysis Using Sequence Labeling as Dependency Graph Parsing

Oct 08, 2026

This study addresses the challenge of fine-grained sentiment analysis, which requires precisely identifying complex relationships among holders, targets, and sentiment expressions. Traditional structured parsing methods rely on intricate graph models with high computational costs. Departing from conventional paradigms, this work reformulates structured sentiment analysis as a dependency graph parsing task. It introduces a linearized graph encoding strategy that solves graph structure prediction directly through sequence labeling models, thereby circumventing complex decoding procedures. The proposed lightweight architecture enables efficient inference while achieving strong performance across five languages and seven benchmark datasets. Its results are comparable to, or even surpass, those of existing complex single-task models, offering a concise and effective unified solution for cross-lingual fine-grained sentiment analysis.

0 citationsRead paper

I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers

Oct 07, 2026

This study investigates the proliferation of contrastive constructions, such as "rather than," in NLP papers generated by large language models (LLMs), which introduces content redundancy and elicits reviewer dissatisfaction. Through text mining, human annotation, preference data analysis, and reward model evaluation, this work systematically examines the causes underlying this surge in LLM-assisted academic writing. Our findings reveal that post-training mechanisms based on human preference data inadvertently induce the overuse of contrastive structures as a side effect, and we quantify their detrimental impact on reader experience. Notably, the frequency of such constructions in 2026 is seven times that of 2019, with the majority judged as low-quality expressions that degrade academic rigor.

0 citationsRead paper

Agentic AI-Assisted Modeling for Production Scheduling: Assessment in Constraint Programming

Oct 07, 2026

This study addresses the reliance on expert knowledge in production scheduling modeling and the high code hallucination rates of general-purpose large language models (LLMs), which hinder practical deployment. To overcome these challenges, this work proposes an automated modeling framework integrating a multi-agent architecture with retrieval-augmented generation. Leveraging fine-tuning-free LLM agents, the method dynamically retrieves solver documentation via the Model Context Protocol to automatically translate natural language descriptions into executable constraint programming code, effectively bridging modeling and AI-driven decision-making. Experimental results demonstrate that the single-run success rate of generated scripts improves from 14.8% to 59.3%, reaching 80.6% for medium-complexity problems. These findings significantly mitigate hallucinations and validate the feasibility of LLM-assisted modeling.

0 citationsRead paper

Argument Structure Prediction in Online Conversations: A Comparative Study of Modeling Paradigms and Task Architectures

Sep 30, 2026

This study addresses the unclear trade-offs among modeling paradigms and task architectures for online conversational argument structure prediction. We present the first multi-paradigm comparative analysis under strict constraints, developing a dedicated data processing pipeline that transforms Informational Argumentation Trees (IAT) into bipolar argument structures. Furthermore, this work systematically evaluates the end-to-end performance and efficiency of supervised fine-tuning and large language models within both single-step and multi-step architectures. Our findings reveal that relation identification constitutes the primary bottleneck and delineate the respective advantages and limitations of each paradigm. To facilitate future research, we release the complete framework as open source.

0 citationsRead paper

Synthetic Data Characterization via Training Dynamics

Sep 30, 2026

This study addresses the challenge of interpreting the intrinsic properties of large language model (LLM)-generated data, whose utility and limitations in learning tasks remain unclear. To this end, it proposes a synthetic data modeling framework grounded in sample-level learnability representations. By dynamically inferring empirical data distributions through encoder training, the approach systematically compares various LLM families, model scales, and human-written data, while validating cross-encoder robustness and data selection strategies across single- and multi-label classification tasks. The findings reveal fundamental differences in learnability between machine-generated and human-authored data, demonstrating that learnability-driven data selection strategies are effective across diverse data sources. Ultimately, this work establishes a novel paradigm for evaluating synthetic data quality, offering actionable insights into the principled utilization of LLM-generated content for downstream learning applications.

0 citationsRead paper