Document Type : Original/Review Paper

Authors

Department of Arabic Language and Literature, Hakim Sabzevari University, Sabzevar, Iran.

10.22044/jadm.2026.17614.2910

Abstract

Machine translation of low-resource and domain-specific texts remains a challenging problem, particularly in the absence of appropriate evaluation methodologies. In this study, we propose a multi-reference data augmentation framework for low-resource text translation. Two publicly available, pre-trained Transformer-based machine translation models, mBART-50 and M2M100, are employed and fine-tuned on multi-reference, domain-adapted data. A comprehensive set of automatic evaluation metrics—including BLEU, chrF, BERTScore, METEOR, and COMET—is used in both single-reference and multi-reference settings to assess translation quality. In addition, the translations are evaluated by two human evaluators with expertise in Qur’anic studies and Persian linguistics. The experimental results demonstrate substantial improvements, particularly in terms of BLEU and COMET scores, indicating improved translation performance for this low-resource, domain-specific task. The results further suggest that multi-reference training can improve translation quality and that multi-reference evaluation provides a broader assessment of model performance. However, automatic metrics cannot fully capture the theological, cultural, and interpretive fidelity required for Qur’anic translation; therefore, their results are complemented by human expert evaluation.

Keywords

Main Subjects

[1] S. Ruder, I. Vulić, and A. Søgaard, "A survey of cross-lingual word embedding models," Journal of Artificial Intelligence Research, vol. 65, pp. 569–631, 2019.
 
[2] M. A. Hedderich, L. Lange, H. Adel, J. Strötgen, and D. Klakow, "A survey on recent approaches for natural language processing in low-resource scenarios," in Proc. 2021 Conf. North American Chapter Association for Computational Linguistics: Human Language Technologies, pp. 2545–2568, 2021.
 
[3] J. Wei and K. Zou, "EDA: Easy data augmentation techniques for boosting performance on text classification tasks," arXiv preprint arXiv:1901.11196, 2019.
 
[4] J. Raiman and J. Miller, "Globally normalized reader," arXiv preprint arXiv:1709.02828, 2017.
 
[5] X. Dai and H. Adel, "An analysis of simple data augmentation for named entity recognition," in Proc. 28th International Conference on Computational Linguistics, pp. 3861–3867, Barcelona, Spain, 2020.
 
[6] K. Gulordava, P. Bojanowski, E. Grave, T. Linzen, and M. Baroni, "Colorless green recurrent networks dream hierarchically," in Proc. 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 1195–1205, New Orleans, LA, 2018.
 
[7] C. Vania, Y. Kementchedjhieva, A. Søgaard, and A. Lopez, "A systematic comparison of methods for low-resource dependency parsing on genuinely low-resource languages," arXiv preprint arXiv:1909.02857, 2019.
 
[8] M. Fadaee, A. Bisazza, and C. Monz, "Data augmentation for low-resource neural machine translation," in Proc. 55th Annual Meeting of the Association for Computational Linguistics, vol. 2, pp. 567–573, Vancouver, Canada, 2017.
 
[9] S. Kobayashi, "Contextual augmentation: Data augmentation by words with paradigmatic relations," in Proc. 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 2, pp. 452–457, New Orleans, LA, 2018.
 
[10] Z. Li, K. Parnow, M. Utiyama, E. Sumita, and H. Zhao, "MiSS: An assistant for multi-style simultaneous translation," in Proc. 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 1–10, 2021.
 
[11] A. M. Mutawa and A. Alrumaih, "Determining the meter of classical Arabic poetry using deep learning: A performance analysis," Frontiers in Artificial Intelligence, vol. 8, p. 1523336, 2025.
 
[12] S. Altammami, E. Atwell, and A. Alsalka, "Towards a joint ontology of Quran and Hadith," International Journal on Islamic Applications in Computer Science and Technology, vol. 10, no. 2, pp. 01–12, 2022.
 
[13] M. Bamoki, S. H. Wady, and S. Badawi, "Holy Quran Kurdish Sorani translation dataset for language modelling," Data in Brief, vol. 60, p. 111533, 2025.
 
[14] D. M. A. A. Danish Khaleeq, M. Tariq, and J. Iqbal, "Machine translation of Quranic verses: A transformer-based approach to Urdu rendering," International Journal of Innovations in Science and Technology, vol. 7, no. 2, pp. 702–717, 2025.
 
[15] A. AlSajri, "Challenges in translating Arabic literary texts using artificial intelligence techniques," EDRAAK, vol. 2023, pp. 5–10, 2023.
 
[16] A. R. Ghasemi and J. Salimi Sartakhti, "Multilingual Language Models in Persian NLP Tasks: A Performance ‎Comparison of Fine-Tuning Techniques," Journal of AI and Data Mining, vol.  13, no. 1, pp. 107-117, 2025.
 
[17] F. Moodi, A. Jahangard-Rafsanjani, and F. Sadri, “Aspect-based sentiment analysis based on users’ comments in an online marketplace,” Journal of Modeling in Engineering, vol. 24, no. 84, pp. 27–46, 2026.
 
[18] F. Saeedi, G. Al Hinai, K. Al Kharusi, and A. A. Abdulsalam, "Effect of context and tokenization on machine translation of Arabic conversations on social media," Procedia Computer Science, vol. 258, pp. 1757–1763, 2025.
 
[19] A. Alabdullah, L. Han, and C. Lin, "Advancing dialectal Arabic to modern standard Arabic machine translation," arXiv preprint arXiv:2507.20301, 2025.
 
[20] M. A. Faheem, K. T. Wassif, H. Bayomi, and S. M. Abdou, "Improving neural machine translation for low resource languages through non-parallel corpora: A case study of Egyptian dialect to modern standard Arabic translation," Scientific Reports, vol. 14, no. 1, p. 2265, 2024.
 
[21] R. Hidayat and S. Minati, "Comparative analysis of text mining classification algorithms for English and Indonesian Qur'an translation," International Journal on Informatics for Development, vol. 8, no. 1, pp. 47–51, 2019.
 
[22] Z. Touati-Hamad, M. R. Laouar, I. Bendib, and S. Hakak, "Arabic Quran verses authentication using deep learning and word embeddings," The International Arab Journal of Information Technology, vol. 19, no. 4, pp. 681–688, 2022.
 
[23] F. Moodi, A. J. Rafsanjani, S. Zarifzadeh, and M. A. Z. Chahooki, "Fusion of technical indicators and sentiment analysis in a hybrid framework of deep learning models for stock price movement prediction, " IEEE Access, vol. 12, pp. 195696–195709, 2024.
 
[24] F. Moodi, A. Jahangard Rafsanjani, S. Zarifzadeh, and M. A. Zare Chahooki, "Advanced stock price forecasting using a 1D-CNN-GRU-LSTM model, " Journal of AI and Data Mining, vol. 12, no. 3, pp. 393–408, 2024.
 
[25] W. Antoun, F. Baly, and H. Hajj, "AraBERT: Transformer-based model for Arabic language understanding," arXiv preprint arXiv:2003.00104, 2020.
 
[26] F. Moodi, A. Jahangard-Rafsanjani, and S. Zarifzadeh, "Improving stock price prediction using technical indicators and sentiment analysis, " Tabriz Journal of Electrical Engineering, vol. 55, no. 2, pp. 357–370, 2025.
 
[27] A. Alsaleh, E. Atwell, and A. Altahhan, "Quranic verses semantic relatedness using AraBERT," in Proc. Sixth Arabic Natural Language Processing Workshop, pp. 185–190, 2021.
 
[28] R. Malhas and T. Elsayed, "Ayatec: Building a reusable verse-based test collection for Arabic question answering on the Holy Qur'an," ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 19, no. 6, pp. 1–21, 2020.
 
[29] D. I. A. Putra and M. Yusuf, "Proposing machine learning of Tafsir al-Quran: In search of objectivity with semantic analysis and natural language processing," in IOP Conference Series: Materials Science and Engineering, vol. 1098, no. 2, p. 022101, 2021.
 
[30] M. S. Hadj Ameur, Y. Moulahoum, and A. Guessoum, "Restoration of Arabic diacritics using a multilevel statistical model," in Proc. IFIP International Conference on Computer Science and Its Applications, pp. 181–192, Cham, Switzerland, 2015.
 
[31] N. Habash and O. Rambow, "Arabic diacritization through full morphological tagging," in Proc. Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Companion Volume, Short Papers, pp. 53–56, 2007.
 
[32] M. Diab, M. Ghoneim, and N. Habash, "Arabic diacritization in the context of statistical machine translation," in Proc. Machine Translation Summit XI, 2007.
 
[33] S. A. Almaaytah and S. A. Alzobidy, "Challenges in rendering Arabic text to English using machine translation: A systematic literature review," IEEE Access, vol. 11, pp. 94772–94779, 2023.
 
[34] Y. Tang, C. Tran, X. Li, P.-J. Chen, N. Goyal, V. Chaudhary, J. Gu, and A. Fan, "Multilingual translation with extensible multilingual pretraining and finetuning," arXiv preprint arXiv:2008.00401, 2020.
 
[35] Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis, and L. Zettlemoyer, "Multilingual denoising pre-training for neural machine translation," Transactions of the Association for Computational Linguistics, vol. 8, pp. 726–742, 2020.
 
[36] A. Fan, S. Bhosale, H. Schwenk, Z. Ma, A. El-Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary, N. Goyal, T. Birch, V. Liptchinsky, S. Edunov, E. Grave, M. Auli, and A. Joulin, "Beyond English-centric multilingual machine translation," Journal of Machine Learning Research, vol. 22, no. 107, pp. 1–48, 2021.
 
[37] K. Papineni, S. Roukos, T. Ward, and W. J. Zhu, "BLEU: A method for automatic evaluation of machine translation," in Proc. 40th Annual Meeting of the Association for Computational Linguistics, pp. 311–318, 2002.
 
[38] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, "BERTScore: Evaluating text generation with BERT," arXiv preprint arXiv:1904.09675, 2019.
 
[39] R. Rei, C. Stewart, A. C. Farinha, and A. Lavie, "COMET: A neural framework for MT evaluation," arXiv preprint arXiv:2009.09025, 2020.