SinBrief: A Hybrid Framework for Abstractive Text Summarisation of Sinhala Legal Documents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of annotated data and terminological complexity in legal document summarization for Sinhala, a low-resource language. We propose an annotation-free hybrid summarization framework that integrates domain-aware word graph construction with a neural sentence scoring mechanism. By leveraging models such as mBERT and Llama 3.1 alongside domain-adaptive pretraining techniques, the method achieves fully unsupervised training, effectively eliminating the reliance on labeled data prevalent in low-resource legal NLP. Experimental results demonstrate that the generated summaries maintain excellent factual consistency while exhibiting low lexical overlap with source documents. These findings validate the effectiveness of combining unsupervised graph structures with multi-model scoring fusion strategies in low-resource scenarios.
📝 Abstract
Legal document summarisation in low-resource languages presents significant challenges due to the scarcity of annotated data and the complexity of domain-specific terminology. This paper presents SinBrief, a hybrid abstractive summarisation framework for Sinhala legal documents that does not require human-annotated training data. The proposed framework combines domain-aware word graph construction with neural sentence scoring to generate abstractive summaries from Sinhala legal text. Five sentence scoring models are evaluated within the framework: mBert, Llama 3.1, Falcon 7B, Laser, and a continually pre-trained Llama model domain-adapted to Sinhala legal text. The framework is evaluated on a Sinhala legal corpus using reference-free metrics, including Coverage, Density, Compression Ratio, SummaC, and Self-BertScore. Experimental results demonstrate that SinBrief produces summaries with lower lexical overlap than extractive baselines while maintaining factual consistency, demonstrating the viability of hybrid, largely annotation-free abstractive summarisation for low-resource legal NLP tasks.
Problem

Research questions and friction points this paper is trying to address.

Low-resource languages
Legal document summarisation
Abstractive summarisation
Sinhala
Annotated data scarcity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Abstractive Summarisation
Low-resource Language
Legal NLP
Hybrid Framework
Domain Adaptation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Minduli Lasandi
School of Computing, Informatics Institute of Technology, Sri Lanka
Nevidu Jayatilleke
Nevidu Jayatilleke
University of Moratuwa, Sri Lanka
Computational LinguisticsArtificial IntelligenceMachine Learning