Comparative Study of LLMs and Transformers for Bangla Healthcare Paraphrasing
Shifat Islam, Md. Shorif Uddin, Rayhan Uddin Bhuiyan, Azizul Hakim, Faika Fairuj Preotee, Mohammad Shahmidul Islam
2025 28th International Conference on Computer and Information Technology (ICCIT) · IEEE
Abstract
Paraphrase generation plays a vital role in natural language processing (NLP), particularly in specialized domains such as healthcare where semantic accuracy is critical. Despite Bangla being one of the most widely spoken languages globally, it remains a low-resource language in NLP, with limited datasets and models for paraphrasing tasks. This study addresses this gap by evaluating the performance of both generalpurpose Large Language Models (LLMs) and domain-specific transformers on the bangla-health-related-paraphrased-dataset, comprising 200,000 sentence pairs. We compare five state-of-theart LLMs-Deepseek R1, Llama 3.3, Meta-Llama 4 Maverick, GPT-OSS 120B, and GPT-OSS 20B-with BanglaT5-base and BERT-base models. Results show that GPT-OSS 120B achieves the highest zero-shot performance (ROUGE-1: 0.69, ROUGE2: 0.40), while BanglaT5 demonstrates strong results among fine-tuned models (ROUGE-1: 0.5258, ROUGE-2: 0.2675). Our analysis highlights that multilingual LLMs can effectively handle Bangla healthcare paraphrasing without task-specific training, yet domain-specific models remain valuable for resourceconstrained applications. These findings provide benchmark results for Bangla paraphrase generation and open pathways for developing practical healthcare NLP tools in low-resource language settings.
Keywords
- Type
- Conference Paper
- Status
- Published
- Year
- 2025
- Publisher
- IEEE