<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>Shahrood University of Technology</PublisherName>
				<JournalTitle>Journal of AI and Data Mining</JournalTitle>
				<Issn>2322-5211</Issn>
				<Volume></Volume>
				<Issue>Articles in Press</Issue>
				<PubDate PubStatus="epublish">
					<Year>2026</Year>
					<Month>07</Month>
					<Day>18</Day>
				</PubDate>
			</Journal>
<ArticleTitle>Robust Multilingual RAG under Query Perturbations: An English-Persian Benchmark</ArticleTitle>
<VernacularTitle></VernacularTitle>
			<FirstPage></FirstPage>
			<LastPage></LastPage>
			<ELocationID EIdType="pii">3856</ELocationID>
			
<ELocationID EIdType="doi">10.22044/jadm.2026.17608.2908</ELocationID>
			
			<Language>EN</Language>
<AuthorList>
<Author>
					<FirstName>Niloofar</FirstName>
					<LastName>Ranjbar</LastName>
<Affiliation>Persian Gulf University</Affiliation>

</Author>
<Author>
					<FirstName>Hamed</FirstName>
					<LastName>Baghbani</LastName>
<Affiliation>Jam Faculty of Engineering, Persian Gulf University, Bushehr, Iran</Affiliation>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2026</Year>
					<Month>04</Month>
					<Day>13</Day>
				</PubDate>
			</History>
		<Abstract>Retrieval-augmented generation (RAG) is commonly evaluated on clean inputs that underrepresent realistic multilingual variation. We present an English-Persian movie-domain robustness benchmark built from a corpus of 31,564 records, 120 clean queries, and 720 aligned perturbations. The benchmark covers six deterministic query types and 14 operational perturbation labels grouped into four families. We compare BM25, multilingual dense retrieval, character n-gram TF-IDF, and hybrid retrieval, and evaluate top-1 deterministic answer extraction against a field-specific top-5 RAG system using Qwen2-7B-Instruct. Hybrid retrieval achieves 81.50 MRR@10 on clean queries and 67.76 under perturbation; field-specific RAG reaches 84.17% and 72.08% accuracy, respectively. Clustered paired-bootstrap 95% confidence intervals exclude zero for all principal system differences. English-title noise is the most damaging family, whereas query-form and punctuation variation is comparatively well tolerated. A 43-case consistency audit verifies implementation of the rule-based failure categories, and full-output analysis shows that retrieval-coverage errors dominate the difficult English-title family. These results support component-level evaluation of multilingual RAG robustness.</Abstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">retrieval-augmented generation</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">multilingual information retrieval</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">robustness evaluation</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">English-Persian retrieval</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">hybrid retrieval</Param>
			</Object>
		</ObjectList>
</Article>
</ArticleSet>
