Document Type : Original/Review Paper

Authors

Jam Technical Faculty, Persian Gulf University, Jam, Iran.

10.22044/jadm.2026.17608.2908

Abstract

Retrieval-augmented generation (RAG) is commonly evaluated on clean inputs that underrepresent realistic multilingual variation. We present an English-Persian movie-domain robustness benchmark built from a corpus of 31,564 records, 120 clean queries, and 720 aligned perturbations. The benchmark covers six deterministic query types and 14 operational perturbation labels grouped into four families. We compare BM25, multilingual dense retrieval, character n-gram TF-IDF, and hybrid retrieval, and evaluate top-1 deterministic answer extraction against a field-specific top-5 RAG system using Qwen2-7B-Instruct. Hybrid retrieval achieves 81.50 MRR@10 on clean queries and 67.76 under perturbation; field-specific RAG reaches 84.17% and 72.08% accuracy, respectively. Clustered paired-bootstrap 95% confidence intervals exclude zero for all principal system differences. English-title noise is the most damaging family, whereas query-form and punctuation variation is comparatively well tolerated. A 43-case consistency audit verifies implementation of the rule-based failure categories, and full-output analysis shows that retrieval-coverage errors dominate the difficult English-title family. These results support component-level evaluation of multilingual RAG robustness.

Keywords

Main Subjects

[1] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Kuttler, M. Lewis, W.-t. Yih, T. Rocktaschel, S. Riedel, and D. Kiela, "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, " in Advances in Neural Information Processing Systems 33, 2020, pp. 9459-9474.
[2] V. Karpukhin, B. Oguz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W.-t. Yih, "Dense Passage Retrieval for Open-Domain Question Answering, " in Proc. EMNLP, 2020, pp. 6769-6781, doi: 10.18653/v1/2020.emnlp-main.550.
[3] F. Petroni, A. Piktus, A. Fan, P. Lewis, M. Yazdani, N. De Cao, J. Thorne, Y. Jernite, V. Karpukhin, J. Maillard, V. Plachouras, T. Rocktaschel, and S. Riedel, "KILT: A Benchmark for Knowledge Intensive Language Tasks, " in Proc. NAACL-HLT, 2021, pp. 2523-2544, doi: 10.18653/v1/2021.naacl-main.200.
[4] S. Es, J. James, L. Espinosa Anke, and S. Schockaert, "RAGAs: Automated Evaluation of Retrieval Augmented Generation, " in Proc. EACL System Demonstrations, 2024, pp. 150-158, doi: 10.18653/v1/2024.eacl-demo.16.
[5] J. Saad-Falcon, O. Khattab, C. Potts, and M. Zaharia, "ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems, " in Proc. NAACL-HLT, 2024, pp. 338-354, doi: 10.18653/v1/2024.naacl-long.20.
[6] R. Friel, M. Belyi, and A. Sanyal, "RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems, " arXiv:2407.11005, 2024, doi: 10.48550/arXiv.2407.11005.
[7] P. Tasawong, W. Ponwitayarat, P. Limkonchotiwat, C. Udomcharoenchaikit, E. Chuangsuwanich, and S. Nutanong, "Typo-Robust Representation Learning for Dense Retrieval, " in Proc. ACL, vol. 2, 2023, pp. 1106-1115, doi: 10.18653/v1/2023.acl-short.95.
[8] G. Sidiropoulos and E. Kanoulas, "Improving the Robustness of Dense Retrievers Against Typos via Multi-Positive Contrastive Learning, " in Advances in Information Retrieval: ECIR 2024, Part III, Springer, 2024, pp. 297-305, doi: 10.1007/978-3-031-56063-7_21.
[9] N. Thakur, N. Reimers, A. Ruckle, A. Srivastava, and I. Gurevych, "BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models, " in Proc. NeurIPS Datasets and Benchmarks Track, vol. 1, 2021.
[10] X. Zhang, N. Thakur, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, and J. Lin, "MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages, " Trans. Assoc. Comput. Linguistics, vol. 11, pp. 1114-1131, 2023, doi: 10.1162/tacl_a_00595.
[11] S. Percin, X. Su, Q. S. Syed, P. Howard, A. Kuvshinov, L. Schwinn, and K.-U. Scholl, “Investigating the Robustness of Retrieval-Augmented Generation at the Query Level, " in Proc. Fourth Workshop on Generation, Evaluation and Metrics (GEM2), Vienna, Austria and virtual meeting, 2025, pp. 439-457, Association for Computational Linguistics.
[12] I. T. Sorodoc, L. F. R. Ribeiro, R. Blloshmi, C. Davis, and A. de Gispert, "GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation, " in Findings of ACL, 2025, pp. 17030-17049, doi: 10.18653/v1/2025.findings-acl.875.
[13] M. Song, S. H. Sim, R. Bhardwaj, H. L. Chieu, N. Majumder, and S. Poria, "Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse, " in Proc. ICLR, 2025, OpenReview: Iyrtb9EJBp.
[14] O. Khattab and M. Zaharia, "ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT, " in Proc. SIGIR, 2020, pp. 39-48, doi: 10.1145/3397271.3401075.
[15] T. Formal, B. Piwowarski, and S. Clinchant, "SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking, " in Proc. SIGIR, 2021, pp. 2288-2292, doi: 10.1145/3404835.3463098.
[16] M. A. C. Blandon, J. Talur, B. Charron, D. Liu, S. Mansour, and M. Federico, "MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation, " in Proc. ACL, 2025, pp. 22577-22595, doi: 10.18653/v1/2025.acl-long.1101.
[17] B. Li, F. Luo, S. Haider, A. Agashe, S. Li, R. Liu, M. M. Miao, S. Ramakrishnan, Y. Yuan, and C. Callison-Burch, "Multilingual Retrieval Augmented Generation for Culturally-Sensitive Tasks: A Benchmark for Cross-lingual Robustness, " in Findings of ACL, 2025, pp. 4215-4241, doi: 10.18653/v1/2025.findings-acl.219.
[18] H. Hosseini, M. S. Zare, A. H. Mohammadi, A. Kazemi, Z. Zojaji, and M. A. Nematbakhsh, "PersianRAG: A Retrieval-Augmented Generation System for Persian Language, " in Proc. 15th Int. Conf. Information and Knowledge Technology (IKT), 2024, pp. 272-278, doi: 10.1109/IKT65497.2024.10892726.
[19] S. B. Hosseinbeigi, M. H. Shalchian, S. Asghari, M. A. Seif Kashani, and M. A. Abbasi, "Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization, " in Proc. 15th Language Resources and Evaluation Conference (LREC 2026), 2026, pp. 7320-7330.
[20] A. Ghasemi, "Iranian Movies, " Kaggle dataset. [Online]. Available: https://www.kaggle.com/datasets/arianghasemi/iranian-movies. [Accessed: Jul. 11, 2026].
[21] R. Banik, "The Movies Dataset, " Kaggle dataset. [Online]. Available: https://www.kaggle.com/datasets/rounakbanik/the-movies-dataset. [Accessed: Jul. 11, 2026].
[22] A. Yang et al., "Qwen2 Technical Report, " arXiv:2407.10671, 2024, doi: 10.48550/arXiv.2407.10671.