Institution profile

University of Belgrade

Academic institutioneurope · rs
Official website
Research library19linked papers
Opportunities0open roles
Selected work

Representative Papers

Wiki Dumps to Training Corpora: South Slavic Case

Apr 28, 2026

This work proposes a systematic methodology for constructing high-quality training corpora for South Slavic language models from raw Wikimedia data. Starting with multilingual Wiki project texts, the approach first parses Wiki markup to extract natural language content and then employs an n-gram–based redundancy detection mechanism to effectively filter out highly repetitive, low-information articles. This pipeline substantially enhances the linguistic richness and authenticity of the resulting corpus while maintaining cross-lingual applicability. The final resource encompasses seven South Slavic languages, offering a reliable foundation for large language model training and cross-linguistic comparative studies.

0 citationsRead paper
Recent publications

Latest Papers

Wiki Dumps to Training Corpora: South Slavic Case

Apr 28, 2026

This work proposes a systematic methodology for constructing high-quality training corpora for South Slavic language models from raw Wikimedia data. Starting with multilingual Wiki project texts, the approach first parses Wiki markup to extract natural language content and then employs an n-gram–based redundancy detection mechanism to effectively filter out highly repetitive, low-information articles. This pipeline substantially enhances the linguistic richness and authenticity of the resulting corpus while maintaining cross-lingual applicability. The final resource encompasses seven South Slavic languages, offering a reliable foundation for large language model training and cross-linguistic comparative studies.

0 citationsRead paper