🤖 AI Summary
This study addresses the challenges of modeling sentence-level semantic textual similarity (STS) for Slovak, a low-resource language, by systematically evaluating traditional algorithms, supervised machine learning models, and prominent deep learning approaches—including SlovakBERT, OpenAI embeddings, GPT-4, and CloudNLP. The work innovatively integrates conventional linguistic features with the artificial bee colony optimization algorithm to perform joint feature selection and hyperparameter tuning. It represents the first application of intelligent optimization strategies to Slovak STS tasks, uncovering critical performance trade-offs among diverse methodologies under low-resource conditions. The findings offer an effective solution and establish a foundational benchmark for semantic understanding in Slovak, thereby advancing NLP capabilities for under-resourced languages.
📝 Abstract
Semantic textual similarity (STS) plays a crucial role in many natural language processing tasks. While extensively studied in high-resource languages, STS remains challenging for under-resourced languages such as Slovak. This paper presents a comparative evaluation of sentence-level STS methods applied to Slovak, including traditional algorithms, supervised machine learning models, and third-party deep learning tools. We trained several machine learning models using outputs from traditional algorithms as features, with feature selection and hyperparameter tuning jointly guided by artificial bee colony optimization. Finally, we evaluated several third-party tools, including fine-tuned model by CloudNLP, OpenAI's embedding models, GPT-4 model, and pretrained SlovakBERT model. Our findings highlight the trade-offs between different approaches.