Ranathunga, Surangika and Sirithunga, Rumesh and Rathnayake, Himashi and Shekhar, Ravi and et al (2025) SiTSE: Sinhala Text Simplification Dataset and Evaluation. ACM Transactions on Asian and Low-Resource Language Information Processing, 24 (5). pp. 1-19. DOI https://doi.org/10.1145/3723160
Ranathunga, Surangika and Sirithunga, Rumesh and Rathnayake, Himashi and Shekhar, Ravi and et al (2025) SiTSE: Sinhala Text Simplification Dataset and Evaluation. ACM Transactions on Asian and Low-Resource Language Information Processing, 24 (5). pp. 1-19. DOI https://doi.org/10.1145/3723160
Ranathunga, Surangika and Sirithunga, Rumesh and Rathnayake, Himashi and Shekhar, Ravi and et al (2025) SiTSE: Sinhala Text Simplification Dataset and Evaluation. ACM Transactions on Asian and Low-Resource Language Information Processing, 24 (5). pp. 1-19. DOI https://doi.org/10.1145/3723160
Abstract
Text Simplification is a task that has been minimally explored for low-resource languages. Consequently, there are only a few manually curated datasets. In this article, we present a human-curated sentence-level text simplification dataset for the Sinhala language. Our evaluation dataset contains 1,000 complex sentences and 3,000 corresponding simplified sentences produced by three different human annotators. We model the text simplification task as a zero-shot and zero-resource sequence-to-sequence (seq-seq) task on the multilingual language models mT5 and mBART. We exploit auxiliary data from related seq-seq tasks and explore the possibility of using intermediate task transfer learning (ITTL). Our analysis shows that ITTL outperforms the previously proposed zero-resource methods for text simplification. Our findings also highlight the challenges in evaluating text simplification systems and support the calls for improved metrics for measuring the quality of automated text simplification systems that would suit low-resource languages as well. Our code and data are publicly available: https://github.com/brainsharks-fyp17/Sinhala-Text-Simplification-Dataset-and-Evaluation.
| Item Type: | Article |
|---|---|
| Uncontrolled Keywords: | Intermediate Task Transfer Learning; Low-resource languages; Text Simplification; Transfer Learning; mT5; mBART; sequence-sequence pre-trained language models |
| Divisions: | Faculty of Science and Health Faculty of Science and Health > Computer Science and Electronic Engineering, School of |
| SWORD Depositor: | Unnamed user with email elements@essex.ac.uk |
| Depositing User: | Unnamed user with email elements@essex.ac.uk |
| Date Deposited: | 28 Jul 2026 14:55 |
| Last Modified: | 28 Jul 2026 14:55 |
| URI: | http://repository.essex.ac.uk/id/eprint/40490 |
Available files
Filename: SiTSE Sinhala Text Simplification Dataset and Evaluation.pdf
Licence: Creative Commons: Attribution 4.0