<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>Shahrood University of Technology</PublisherName>
				<JournalTitle>Journal of AI and Data Mining</JournalTitle>
				<Issn>2322-5211</Issn>
				<Volume></Volume>
				<Issue>Articles in Press</Issue>
				<PubDate PubStatus="epublish">
					<Year>2026</Year>
					<Month>07</Month>
					<Day>18</Day>
				</PubDate>
			</Journal>
<ArticleTitle>HiSGAN: Image-to-Image Translation with Involution based Hybrid-Scale Transformer and Contrastive learning</ArticleTitle>
<VernacularTitle></VernacularTitle>
			<FirstPage></FirstPage>
			<LastPage></LastPage>
			<ELocationID EIdType="pii">3854</ELocationID>
			
<ELocationID EIdType="doi">10.22044/jadm.2026.16833.2815</ELocationID>
			
			<Language>EN</Language>
<AuthorList>
<Author>
					<FirstName>Farzane</FirstName>
					<LastName>Maghsoudi</LastName>
<Affiliation>Department of Electrical and Computer Engineering, Semnan University</Affiliation>

</Author>
<Author>
					<FirstName>Mohammad Javad</FirstName>
					<LastName>FadaeiEslam</LastName>
<Affiliation>Department of Electrical and Computer Engineering, Semnan University</Affiliation>

</Author>
<Author>
					<FirstName>Farzin</FirstName>
					<LastName>Yaghmaee</LastName>
<Affiliation>Department of Electrical and Computer Engineering-Semnan University-</Affiliation>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2025</Year>
					<Month>10</Month>
					<Day>24</Day>
				</PubDate>
			</History>
		<Abstract>Image-to-image translation is a highly challenging task, as it requires an accurate understanding of image details and their consistent transformation across domains. Notably, GANs have achieved remarkable success in this field. In essence, convolutional layers are the primary building blocks of these architectures. However, the limited receptive field in shallow layers makes it difficult to capture long-range spatial dependencies and non-local context. In this paper, the HiSGAN architecture is proposed to address this limitation. It combines deep representations with traditional techniques, such as SVD and Fast Fourier Convolution (FFC), to effectively extract style-related information and establish a global receptive field. Furthermore, we introduce the HiS-Transformer block with an involution operator in the bottleneck of the generator. This proposed block utilizes hybrid-scale self-attention to adaptively preserve the global receptive field and fine-grained information in salient regions while maintaining low computational cost. HiSGAN employs a new loss function based on gradient contrastive learning to improve cross-domain feature alignment. Quantitative and qualitative results on four public datasets demonstrate the superiority of the proposed approach over state-of-the-art methods. Importantly, these performance gains are achieved while reducing the parameter count and accelerating training. The code is available at https://github.com/OliverRensu/SG-Former</Abstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">Image-to-Image Translation</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Fast Fourier Convolution</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Transformers</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Contrastive Learning</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Feature extraction</Param>
			</Object>
		</ObjectList>
</Article>
</ArticleSet>
