🤖 AI Summary
This study addresses the scarcity of annotated data and terminological complexity in legal document summarization for Sinhala, a low-resource language. We propose an annotation-free hybrid summarization framework that integrates domain-aware word graph construction with a neural sentence scoring mechanism. By leveraging models such as mBERT and Llama 3.1 alongside domain-adaptive pretraining techniques, the method achieves fully unsupervised training, effectively eliminating the reliance on labeled data prevalent in low-resource legal NLP. Experimental results demonstrate that the generated summaries maintain excellent factual consistency while exhibiting low lexical overlap with source documents. These findings validate the effectiveness of combining unsupervised graph structures with multi-model scoring fusion strategies in low-resource scenarios.
📝 Abstract
Legal document summarisation in low-resource languages presents significant challenges due to the scarcity of annotated data and the complexity of domain-specific terminology. This paper presents SinBrief, a hybrid abstractive summarisation framework for Sinhala legal documents that does not require human-annotated training data. The proposed framework combines domain-aware word graph construction with neural sentence scoring to generate abstractive summaries from Sinhala legal text. Five sentence scoring models are evaluated within the framework: mBert, Llama 3.1, Falcon 7B, Laser, and a continually pre-trained Llama model domain-adapted to Sinhala legal text. The framework is evaluated on a Sinhala legal corpus using reference-free metrics, including Coverage, Density, Compression Ratio, SummaC, and Self-BertScore. Experimental results demonstrate that SinBrief produces summaries with lower lexical overlap than extractive baselines while maintaining factual consistency, demonstrating the viability of hybrid, largely annotation-free abstractive summarisation for low-resource legal NLP tasks.