🤖 AI Summary
This study addresses the scarcity of lexical simplification resources and joint datasets for Romanian by constructing the first annotated dataset encompassing both lexical complexity prediction and simplification, comprising 3,921 contextualized word instances. The authors propose a pairwise ranking approximation method that leverages human judgments of word complexity to rank candidate substitutions. Building upon this, they design a multi-stage simplification pipeline that integrates machine learning models for complexity prediction and lexical replacement, thereby developing the first automatic lexical simplification system for Romanian. This work establishes a foundational baseline and provides essential resources to support readability optimization efforts for the language.
📝 Abstract
We introduce the first dataset that jointly covers both lexical complexity prediction (LCP) annotations and lexical simplification (LS) for Romanian, along with a comparison of lexical simplification approaches. We propose a methodology for ordering simplification suggestions using a pairwise ranking approximation method, arranging candidates from simple to complex based on a separate set of human judgments. In addition, we provide human lexical complexity annotations for 3,921 word samples in context. Finally, we explore several novel pipelines for complexity prediction and simplification and present the first text simplification system for Romanian.