Document Type : Original/Review Paper

Authors

1 Postdoctoral Researcher in Natural Language Processing, Hakim Sabzevari University, Sabzevar, Iran

2 Assoicate Professor of Arabic Language and Literature, Hakim Sabzevari University, Sabzevar, Iran

10.22044/jadm.2026.17614.2910

Abstract

Machine translation of low-resource and domain-specific texts remains a challenging problem, particularly in the absence of appropriate evaluation methodologies. In this study, we propose a multi-reference data augmentation framework for low-resource text translation. Two publicly available, pre-trained Transformer-based machine translation models, mBART-50 and M2M100, are employed and fine-tuned on multi-reference, domain-adapted data. A comprehensive set of automatic evaluation metrics—including BLEU, chrF, BERTScore, METEOR, and COMET—is used in both single-reference and multi-reference settings to assess translation quality. In addition, the translations are evaluated by two human evaluators with expertise in Qur’anic studies and Persian linguistics. The experimental results demonstrate substantial improvements, particularly in terms of BLEU and COMET scores, indicating improved translation performance for this low-resource, domain-specific task. The results further suggest that multi-reference training can improve translation quality and that multi-reference evaluation provides a broader assessment of model performance. However, automatic metrics cannot fully capture the theological, cultural, and interpretive fidelity required for Qur’anic translation; therefore, their results are complemented by human expert evaluation.

Keywords

Main Subjects