Document Type : Original/Review Paper

Authors

1 Computer Engineering Department, Kurdistan University, Sanandaj, Iran.

2 Computer Engineering Department, Lorestan University, Khorramabad, Iran.

10.22044/jadm.2026.17636.2918

Abstract

Multi-agent reinforcement learning (MARL) is a key paradigm for coordination in robotics, autonomous systems, and distributed control. However, existing MARL methods face fundamental limitations in scalability, adaptability to dynamic environments, and stability under evolving interactions. To address these challenges, we propose Adaptive Graph-Transformer Reinforcement Learning (AGTRL), a framework integrating graph-based relational modelling with transformer attention for adaptive coordination in large-scale multi-agent systems. AGTRL unifies graph-based perception and attention-based coordination in an end-to-end pipeline, encoding role information and adaptively weighting interactions by context. This combination, missing in prior MARL methods, bridges scalability and robustness in dynamic environments. AGTRL constructs a dynamic graph of agent relationships and uses multi-head self-attention to prioritize relevant interactions in real-time, ensuring robust performance under perturbations. We evaluate robustness under communication dropout (up to 40% link removal) and dynamic edge removal, measuring performance via episode reward and win rate. The framework incorporates an adaptive stability-performance trade-off mechanism that maintains learning efficacy in the presence of communication constraints and environmental uncertainty. We introduce a graph-enhanced policy architecture that jointly optimizes individual agent policies and inter-agent coordination through attention-weighted message passing. Comprehensive evaluations on benchmark environments—including StarCraft II micromanagement scenarios, cooperative navigation (Spread), and adversarial tasks (Predator-Prey)—demonstrate that AGTRL achieves superior sample efficiency, scalability, and robustness compared to state-of-the-art MARL baselines. Experimental results show AGTRL improves convergence speed by 32% on average and maintains stable performance with up to 40% communication dropout, establishing its viability for real-world deployment in dynamic multi-agent domains.

Keywords

Main Subjects

[1]   F. L. Lewis, H. Zhang, K. Hengster-Movric, and A. Das, "Cooperative Control of Multi-Agent Systems: Optimal and Adaptive Design Approaches," Springer London, 2014.
[2]   K. Zhang, Z. Yang, and T. Başar, "Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms," Handbook of Reinforcement Learning and Control, vol. 325, no. 1, pp. 321–384, 2021.
[3]   X. Wang et al., "Deep Reinforcement Learning: A Survey," IEEE Transactions on Neural Networks and Learning Systems, vol. 35, pp. 5064–5078, 2024.
[4]   M. S. Ibrahim and M. Hamada, "Adaptive learning framework," 2016 15th International Conference on Information Technology Based Higher Education and Training (ITHET), vol. 15, no. 1, pp. 1–5, 2016.
[5]   H. van Hasselt, A. Guez, and D. Silver, "Deep Reinforcement Learning with Double Q-learning," arXiv:1509.06461, Sep. 2015.
[6]   C. Zhu, M. Dastani, and S. Wang, "A survey of multi-agent deep reinforcement learning with communication," Autonomous Agents and Multi-Agent Systems, vol. 38, no. 1, p. 4, 2024.
[7]   S. Mariani, G. Cabri, and F. Zambonelli, "Coordination of Autonomous Vehicles: Taxonomy and Survey," arXiv:2001.02443, Jan. 2020.
[8]   J. Liang, M. Chen, and J. Liang, "Graph External Attention Enhanced Transformer," arXiv:2405.21061, May 2024.
[9]   S. Yin and G. Zhong, "LGI-GT: Graph Transformers with Local and Global Operators Interleaving," Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence (IJCAI), vol. 32, pp. 4504–4512, 2023.
[10] C. Wang, J. Zhao, L. Li, L. Jiao, F. Liu, and S. Yang, "Automatic Graph Topology-Aware Transformer," IEEE Transactions on Neural Networks and Learning Systems, vol. X, no. Y, pp. 1–15, 2024.
[11]   M. Herrera, M. Pérez-Hernández, A. K. Parlikad, and J. Izquierdo, "Multi-Agent Systems and Complex Networks: Review and Applications in Systems Engineering," Processes, vol. 8, no. 3, p. 312, 2020.
[12]   A. Dorri, S. S. Kanhere, and R. Jurdak, "Multi-Agent Systems: A Survey," IEEE Access, vol. 6, no. 1, pp. 28573–28593, 2018.
[13]   M. Mes and B. Gerrits, "Multi-agent Systems," in Operations, Logistics and Supply Chain Management, vol. X, H. Zijm, M. Klumpp, A. Regattieri, and S. Heragu, Eds., Springer International Publishing, pp. 611–636, 2019.
[14] K. Keogh and L. Sonenberg, "Designing Multi-Agent System Organisations for Flexible Runtime Behaviour," Applied Sciences, vol. 10, no. 15, p. 5335, 2020.
[15] E. Bastami, A. Mahabadi, and E. Taghizadeh, "A gravitation-based link prediction approach in social networks," Swarm and Evolutionary Computation, vol. 44, pp. 176–186, 2019.
[16]   G. Zhang et al., "G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks," arXiv:2410.11782, 2024.
[17]   Z. Fang, F. Ke, J. Y. Han, Z. Feng, and T. Cai, "Graph Enhanced Reinforcement Learning for Effective Group Formation in Collaborative Problem Solving," arXiv:2403.10006, 2024.
[18]   J. N. Foerster, Y. M. Assael, N. de Freitas, and S. Whiteson, "Learning to Communicate with Deep Multi-Agent Reinforcement Learning," arXiv:1605.06676, 2016.
[19] L. Ratnabala, A. Fedoseev, R. Peter, and D. Tsetserukou, "MAGNNET: Multi-Agent Graph Neural Network-based Efficient Task Allocation for Autonomous Vehicles with Deep Reinforcement Learning," arXiv:2502.02311, 2025.
[20]   J Z. Jia, J. Li, X. Qu, and J. Wang, "Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-based Planner and Graph-based Policy," arXiv:2503.10049, 2025.
[21]   F. L. Da Silva and A. H. R. Costa, "Transfer Learning for Multiagent Reinforcement Learning Systems," Springer International Publishing, 2021.
[22]   W. Li, L. Ni, J. Wang, and C. Wang, "Collaborative representation learning for nodes and relations via heterogeneous graph neural network," Knowledge-Based Systems, vol. 255, p. 109673, 2022.
[23]   T. He, Y. Liu, Y.-S. Ong, X. Wu, and X. Luo, "Polarized message-passing in graph neural networks," Artificial Intelligence, vol. 331, p. 104129, 2024.
[24]   C.-W. Leung, S. Hu, and H.-F. Leung, "Self-Play or Group Practice: Learning to Play Alternating Markov Game in Multi-Agent System," 2020 25th International Conference on Pattern Recognition (ICPR), vol. 25, pp. 9234–9241, 2021.
[25]   K. Hu et al., "A review of research on reinforcement learning algorithms for multi-agents," Neurocomputing, vol. 599, p. 128068, 2024.
[26]   J. M. Górriz et al., "Computational approaches to Explainable Artificial Intelligence: Advances in theory, applications and trends," Information Fusion, vol. 100, p. 101945, 2023.
[27]   T. Rashid, M. Samvelyan, C. Schroeder, G. Farquhar, J. Foerster, and S. Whiteson, "QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning," in Proceedings of the 35th International Conference on Machine Learning (ICML), pp. 4295–4304, 2018.
[28]   K. Son, D. Kim, W. J. Kang, D. E. Hostallero, and Y. Yi, "QTRAN: Learning to Factorize with Transformation for Cooperative Multi-Agent Reinforcement Learning," in Proceedings of the 36th International Conference on Machine Learning (ICML), pp. 5887–5896, 2019.
[29]   J. Wang, Z. Ren, T. Liu, Y. Yu, and C. Zhang, "QPLEX: Duplex Dueling Multi-Agent Q-Learning," in International Conference on Learning Representations (ICLR), 2021.
 
[30]   S. Hu, F. Zhu, X. Chang, and X. Liang, "UPDeT: Universal Multi-Agent Reinforcement Learning via Policy Decoupling with Transformers," in International Conference on Learning Representations (ICLR), 2021.
[31] M. Gallici, M. Martin, and I. Masmitja, "TransfQMix: Transformers for Leveraging the Graph Structure of Multi-Agent Reinforcement Learning Problems," Proceedings of the 22nd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2023), vol. 22, pp. 1679–1687, 2023.
[32] J. Huang, J. Su, and Q. Chang, "Graph neural network and multi-agent reinforcement learning for machine-process-system integrated control to optimize production yield," Journal of Manufacturing Systems, vol. 64, pp. 81–93, 2022.
[33] Z. Zhang, B. He, B. Cheng, and G. Li, "Bridging Training and Execution via Dynamic Directed Graph-Based Communication in Cooperative Multi-Agent Systems," in Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2025), vol. 39, no. 22, pp. 23395–23403, 2025.
[34] S. Dolan, S. Nayak, J. J. Aloor, and H. Balakrishnan, "Asynchronous Cooperative Multi-Agent Reinforcement Learning with Limited Communication," Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2025), vol. 24, pp. 1–3, 2025.
[35] Z. Ning and L. Xie, "A survey on multi-agent reinforcement learning and its application," Journal of Automation and Intelligence, vol. 3, no. 2, pp. 73–91, 2024.
[36] M. Tao, Q. Li, and J. Yu, "Multi-Objective Dynamic Path Planning with Multi-Agent Deep Reinforcement Learning," Journal of Marine Science and Engineering, vol. 13, no. 1, p. 20, 2024.
[37]   J. Zhou et al., "Graph neural networks: A review of methods and applications," AI Open, vol. 1, pp. 57–81, 2020.
[38] O. Vinyals et al., "StarCraft II: A New Challenge for Reinforcement Learning," arXiv:1708.04782, Aug. 2017.
[39] K. Wan, D. Wu, Y. Zhai, B. Li, X. Gao, and Z. Hu, "Multi-Agent Actor-Critic for Mixed Cooperative-Competitive," Entropy, vol. 23, no. 11, p. 1433, 2021.
[40]   T. Weng, H. Yang, C. Gu, J. Zhang, P. Hui, and M. Small, "Predator-prey games on complex networks," Communications in Nonlinear Science and Numerical Simulation, vol. 79, p. 104911, 2019.
[41]   R. Hosseinzadeh and M. Sadeghzadeh, "Attention Mechanisms in Transformers: A General Survey," Journal of AI and Data Mining, vol. 13, no. 3, pp. 359–368, Jul. 2025.