Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers

📅 2026-02-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of modeling sentence-level semantic textual similarity (STS) for Slovak, a low-resource language, by systematically evaluating traditional algorithms, supervised machine learning models, and prominent deep learning approaches—including SlovakBERT, OpenAI embeddings, GPT-4, and CloudNLP. The work innovatively integrates conventional linguistic features with the artificial bee colony optimization algorithm to perform joint feature selection and hyperparameter tuning. It represents the first application of intelligent optimization strategies to Slovak STS tasks, uncovering critical performance trade-offs among diverse methodologies under low-resource conditions. The findings offer an effective solution and establish a foundational benchmark for semantic understanding in Slovak, thereby advancing NLP capabilities for under-resourced languages.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Machine Learning: Semi-Supervised LearningConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Semantics and Knowledge: Methods, algorithms and applications for the development of semantic models, knowledge graphs and other forms of structured data models with machine-interpretable semanticsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Semantic textual similarity (STS) plays a crucial role in many natural language processing tasks. While extensively studied in high-resource languages, STS remains challenging for under-resourced languages such as Slovak. This paper presents a comparative evaluation of sentence-level STS methods applied to Slovak, including traditional algorithms, supervised machine learning models, and third-party deep learning tools. We trained several machine learning models using outputs from traditional algorithms as features, with feature selection and hyperparameter tuning jointly guided by artificial bee colony optimization. Finally, we evaluated several third-party tools, including fine-tuned model by CloudNLP, OpenAI's embedding models, GPT-4 model, and pretrained SlovakBERT model. Our findings highlight the trade-offs between different approaches.
Problem

Research questions and friction points this paper is trying to address.

Semantic Textual Similarity
Slovak language
low-resource languages
sentence-level similarity
natural language processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Textual Similarity
Low-resource Language
Artificial Bee Colony Optimization
SlovakBERT
Feature Selection
💼 Related Jobs
No related jobs found.
L
Lukás Radoský
Department of Applied Informatics, Faculty of Mathematics, Physics and Informatics, Comenius University Bratislava, Bratislava, Slovakia
M
Miroslav Blšták
Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia
M
Matej Krajcovic
Department of Applied Informatics, Faculty of Mathematics, Physics and Informatics, Comenius University Bratislava, Bratislava, Slovakia
Ivan Polasek
Ivan Polasek
Faculty of Mathematics, Physics and Informatics, COMENIUS UNIVERSITY IN BRATISLAVA
computer sciencesoftware engineering