Improving Bangla Regional Dialect Detection Using BERT, LLMs, and XAI
Bidyarthi Paul, Faika Fairuj Preotee, Shuvashis Sarker, Tashreef Muhammad
2024 IEEE International Conference on Computing, Applications and Systems (COMPAS) · IEEE
Abstract
This research addresses the challenge of identifying and categorizing Bangla regional dialects through the application of sophisticated natural language processing techniques. Automated translation, digital content personalization, and speech recognition systems are all improved by precise dialect detection, which is a result of the linguistic diversity in Bangladesh. Few-shot learning techniques were used to compare the performance of transformer-based models, specifically Bangla BERT, with state-of-the-art large language models such as GPT-3.5 Turbo and Gemini 1.5 Pro, and to fine-tune them. The methodology entailed the utilization of a diverse regional Bangla speech samples from the Vasantor dataset, which encompassed regions including Mymensingh, Chittagong, Barishal, Noakhali, and Sylhet. In order to enhance the interpretability of the model, implementation of Local Interpretable Model-agnostic Explanations (LIME) was done. The Bangla BERT model achieved the highest accuracy of 88.74%. GPT-3.5 Turbo’s few-shot learning exhibited substantial potential, with an accuracy rate of 64%. These results underscore the significance of hyperparameter optimization and fine-tuning in the enhancement of regional dialect detection models.
Citations by Year
Entered manually, not live-tracked.
- Type
- Conference Paper
- Status
- Published
- Year
- 2024
- Publisher
- IEEE