Document Type : Conceptual Paper

Authors

Electrical and Computer Engineering Department, Semnan University, Semnan, Iran.

10.22044/jadm.2026.16833.2815

Abstract

Image-to-image translation is a highly challenging task, as it requires an accurate understanding of image details and their consistent transformation across domains. Notably, GANs have achieved remarkable success in this field. In essence, convolutional layers are the primary building blocks of these architectures. However, the limited receptive field in shallow layers makes it difficult to capture long-range spatial dependencies and non-local context. In this paper, the HiSGAN architecture is proposed to address this limitation. It combines deep representations with traditional techniques, such as SVD and Fast Fourier Convolution (FFC), to effectively extract style-related information and establish a global receptive field. Furthermore, we introduce the HiS-Transformer block with an involution operator in the bottleneck of the generator. This proposed block utilizes hybrid-scale self-attention to adaptively preserve the global receptive field and fine-grained information in salient regions while maintaining low computational cost. HiSGAN employs a new loss function based on gradient contrastive learning to improve cross-domain feature alignment. Quantitative and qualitative results on four public datasets demonstrate the superiority of the proposed approach over state-of-the-art methods. Importantly, these performance gains are achieved while reducing the parameter count and accelerating training. The code is available at https://github.com/OliverRensu/SG-Former

Keywords

Main Subjects

[1] P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, "Image-to-image translation with conditional adversarial networks, " in Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), Piscataway, 2017, pp. 1125-1134.
 
[2] I. J. Goodfellow, et al., "Generative adversarial nets, " in Advances in Neural Information Processing Systems (NeurIPS), vol. 27, 2014.
 
[3] H. Tu, W. Wang, J. Chen, F. Wu, and G. Li, "Unpaired image-to-image translation with improved two-dimensional feature, " Multimedia Tools and Applications, vol. 81, no. 30, pp. 43851–43872, Dec. 2022.
 
[4] Z. Cao, W. Wang, L. Huo, and S. Niu, "Unsupervised class-to-class translation for domain variations, " Pattern Recognition, vol. 138, Art. no. 109346, 2023.
 
[5] R. Hadsell, S. Chopra, and Y. LeCun, "Dimensionality reduction by learning an invariant mapping, " in Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit. (CVPR), New York, NY, USA, 2006, pp. 1735–1742.
 
[6] F. M. Ghombavani, M. J. Fadaeieslam, and F. Yaghmaee, "ARDA-UNIT: Recurrent dense self-attention block with adaptive feature fusion for unpaired (unsupervised) image-to-image translation, " IET Image Processing, vol. 17, no. 13, pp. 3746–3758, 2023.
 
[7] D. Li, J. Hu, C. Wang, X. Li, Q. She, L. Zhu, T. Zhang, and Q. Chen, "Involution: Inverting the inherence of convolution for visual recognition, " in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 12316–12325.
 
[8] M. Zhao, G. Feng, J. Tan, N. Zhang, and X. Lu, "CSTGAN: Cycle Swin Transformer GAN for unpaired infrared image colorization, " in Proc. 3rd Int. Conf. Control, Robot. Intell. Syst. (CCRIS), virtual event, China, Aug. 26–28, 2022, pp. 1–7.
 
[9] H. Liu, Y. Xu, and F. Chen, "Sketch2Photo: Synthesizing photo-realistic images from sketches via global contexts, " Engineering Applications of Artificial Intelligence, vol. 117, Art. no. 105608, 2023.
 
[10] L. Chi, B. Jiang, and Y. Mu, "Fast Fourier convolution,” in Advances in Neural Information Processing Systems (NeurIPS), 2020.
 
[11] R. Chen, W. Huang, B. Huang, F. Sun, and B. Fang, "Reusing discriminators for encoding: Towards unsupervised image-to-image translation, " in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2020, pp. 8168–8177.
 
[12] S. Ren, X. Yang, S. Liu, and X. Wang, "SG-Former: Self-guided transformer with evolving token reallocation, " in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2023, pp. 5980–5991.
 
[13] Y.-H. Hung, J. Tan, T.-M. Huang, S.-C. Hsu, Y.-L. Chen, and K.-L. Hua, "Unpaired image-to-image translation using negative learning for noisy patches, " IEEE MultiMedia, vol. 29, pp. 59–68, 2022.
 
[14] Y. Zhang, M. Li, W. Cai, et al., "SARCUT: Contrastive learning for optical-SAR image translation with self-attention and relativistic discrimination, " in Proc. SPIE Int. Workshop Frontiers Graphics Image Process. (FGIP 2022), vol. 12644, Art. no. 126440B, May 3, 2023.
 
[15] M. Mirza and S. Osindero, "Conditional generative adversarial nets, " arXiv preprint arXiv:1411.1784, 2014.
 
[16] G. Wang, H. Shi, Y. Chen, and B. Wu, "Unsupervised image-to-image translation via long-short cycle-consistent adversarial networks, " Applied Intelligence, vol. 53, no. 14, pp. 17243–17259, Jul. 2023.
 
[17] J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, "Unpaired image-to-image translation using cycle-consistent adversarial networks, " in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), Venice, Italy, 2017, pp. 2242–2251.
 
[18] Y. Liu, J. Chen, and J.-a. Hou, "Learning position information from attention: End-to-end weakly supervised crack segmentation with GANs, " Computers in Industry, vol. 149, Art. no. 103921, Aug. 2023.
 
[19] J. Kim, M. Kim, H. Kang, and K. H. Lee, "U-GAT-IT: Unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation, " in Proc. Int. Conf. Learn. Representations (ICLR), 2020.
 
[20] H. Deng, Q. Wu, H. Huang, X. Yang, and Z. Wang, "InvolutionGAN: Lightweight GAN with involution for unsupervised image-to-image translation, " Neural Computing and Applications, vol. 35, no. 22, pp. 16593–16605, Aug. 2023.
 
[21] A. Vaswani, et al., "Attention is all you need, " in Proc. 31st Int. Conf. Neural Inf. Process. Syst. (NeurIPS), Long Beach, CA, USA, 2017, pp. 6000–6010.
 
[22] A. Dosovitskiy, et al., "An image is worth 16 × 16 words: Transformers for image recognition at scale, " arXiv preprint arXiv:2010.11929, 2020.
 
[23] Z. Liu et al., "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows," 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 2021, pp. 9992-10002.
 
[24] D. Torbunov, et al., "UVCGAN: UNet vision transformer cycle-consistent GAN for unpaired image-to-image translation, " in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV), Waikoloa, HI, USA, 2023, pp. 702–712.
 
[25] G. Youk and M. Kim, "Transformer-Based Synthetic-to-Measured SAR Image Translation via Learning of Representational Features", IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1–18, 01 2023.
 
[26] Q. Mao and S. Ma, "Enhancing Style-Guided Image-to-Image Translation via Self-Supervised Metric Learning", IEEE Transactions on Multimedia, vol. 25, pp. 8511–8526, Jan. 2023.
 
[27] B. Zhao, W. Li, and W. Gong, "Real-aware motion deblurring using multi-attention CycleGAN with contrastive guidance", Digital Signal Processing, vol. 135, p. 103953, 2023.
 
[28] D. Hendrycks and K. Gimpel, "Gaussian Error Linear Units (GELUs) ", arXiv [cs.LG]. 2023.
 
[29] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, "A ConvNet for the 2020s, " in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), New Orleans, LA, USA, 2022, pp. 11976–11986.
 
[30] X. Mao, Q. Li, H. Xie, R. Y. K. Lau, Z. Wang, and S. P. Smolley, "Least squares generative adversarial networks, " in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), 2017, pp. 2794–2802.
 
[31] M. Arjovsky, S. Chintala, and L. Bottou, "Wasserstein generative adversarial networks, " in Proc. 34th Int. Conf. Mach. Learn. (ICML), Sydney, NSW, Australia, 2017, pp. 214–223.
 
[32] T. Karras, T. Aila, S. Laine, and J. Lehtinen, "Progressive growing of GANs for improved quality, stability, and variation, " in Proc. 6th Int. Conf. Learn. Represent. (ICLR), Vancouver, BC, Canada, 2018.
 
[33] M. H. Khosravi, "A Siamese Network Based on InceptionV3 with Custom Loss Functions for Document Image Quality Assessment (DIQA) , " Journal of AI and Data Mining, vol. 14, no. 3, pp. 291-299, July 2026.
 
[34] Y. Li, H. Meng, H. Lin, and C. Liu, "AFF-UNIT: Adaptive feature fusion for unsupervised image-to-image translation, " IET Image Process., vol. 15, no. 13, pp. 3172–3188, 2021.