<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>Shahrood University of Technology</PublisherName>
				<JournalTitle>Journal of AI and Data Mining</JournalTitle>
				<Issn>2322-5211</Issn>
				<Volume>13</Volume>
				<Issue>1</Issue>
				<PubDate PubStatus="epublish">
					<Year>2025</Year>
					<Month>01</Month>
					<Day>01</Day>
				</PubDate>
			</Journal>
<ArticleTitle>DOSTE: Document Similarity Matching considering Informative Name Entities</ArticleTitle>
<VernacularTitle></VernacularTitle>
			<FirstPage>85</FirstPage>
			<LastPage>94</LastPage>
			<ELocationID EIdType="pii">3391</ELocationID>
			
<ELocationID EIdType="doi">10.22044/jadm.2025.15383.2641</ELocationID>
			
			<Language>EN</Language>
<AuthorList>
<Author>
					<FirstName>Milad</FirstName>
					<LastName>Allhgholi</LastName>
<Affiliation>School of Computer engineering, Iran University of Science and Technology, Tehran, Iran.</Affiliation>

</Author>
<Author>
					<FirstName>Hossein</FirstName>
					<LastName>Rahmani</LastName>
<Affiliation>School of Computer engineering, Iran University of Science and Technology, Tehran, Iran.</Affiliation>

</Author>
<Author>
					<FirstName>Amirhossein</FirstName>
					<LastName>Derakhshan</LastName>
<Affiliation>School of Computer engineering, Iran University of Science and Technology, Tehran, Iran.</Affiliation>

</Author>
<Author>
					<FirstName>Saman</FirstName>
					<LastName>Mohammadi Raouf</LastName>
<Affiliation>School of Computer engineering, Iran University of Science and Technology, Tehran, Iran.</Affiliation>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2024</Year>
					<Month>12</Month>
					<Day>14</Day>
				</PubDate>
			</History>
		<Abstract>Document similarity matching is essential for efficient text retrieval, plagiarism detection, and content analysis. Existing studies in this field can be categorized into three approaches: statistical analysis, deep learning, and hybrid approaches. However, to the best of our knowledge, none have incorporated the importance of named entities into their methodologies. In this paper, we propose DOSTE, a method that first extracts name entities and then utilizes them to enhance document similarity matching through statistical and graph-based analysis. Empirical results indicate that DOSTE achieves better results by emphasizing named entities, resulting in an average improvement of 9% in the average recall metric compared to baseline methods. Also, DOSTE unlike LLM-based approaches, does not require extensive GPU resources. Additionally, non-empirical interpretations of the results indicate that DOSTE is particularly effective in identifying similarity in short documents and complex document comparisons.</Abstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">Document similarity</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Name entities</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Entities graph</Param>
			</Object>
		</ObjectList>
<ArchiveCopySource DocType="pdf">https://jad.shahroodut.ac.ir/article_3391_8c45090f97a2eea5635ec2bec2d44866.pdf</ArchiveCopySource>
</Article>
</ArticleSet>
