A Character-Level Neural Approach to Sinhala Sandhi Splitting

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of ambiguous word boundaries and the absence of neural baselines in Sinhala sandhi splitting. To this end, it introduces the first character-level sequence-to-sequence model built upon SandhiLex. The proposed architecture employs a bidirectional LSTM encoder paired with a unidirectional LSTM decoder, and validates the effectiveness of native Unicode inputs. Experimental results demonstrate that the model achieves an exact match accuracy of 68.4% on a complex subset. This work establishes a neural baseline for the task, reveals the capability of LSTMs in handling intricate sandhi phenomena, and provides clear directions for future optimization.
📝 Abstract
Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, derivational, and etymological Sandhi, where our best model, a bidirectional LSTM encoder with a unidirectional LSTM decoder, reaches only 68.40\% exact-match accuracy (82.08\% character-level accuracy), well below the 94.00\% achieved on the more regular affixational subset. Ablations show that bidirectional encoding is the largest contributor to performance, while native Sinhala script improves exact match accuracy over romanized input. Qualitative analysis indicates that many errors are near misses involving boundary adjacent characters or plausible but incorrect phonological substitutions. These results establish an empirical baseline for Sinhala Sandhi splitting and identify data scale, Sandhi type conditioning, and attention-based decoding as the main directions for future work.
Problem

Research questions and friction points this paper is trying to address.

Sinhala Sandhi splitting
character-level sequence-to-sequence
morphological analysis
natural language processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sandhi Splitting
Character-level Seq2Seq
Bidirectional LSTM
Sinhala NLP
Neural Benchmark
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yasas Ekanayaka
Informatics Institute of Technology, Colombo 06, Sri Lanka
Deshan Sumanathilaka
Deshan Sumanathilaka
PhD Candidate at Swansea University
NLPMachine Translation and TransliterationMLWSD