MADECM: A Curiosity‐Augmented Evolutionary Algorithm for Multi‐Agent Policy Diversity Optimization.

Saved in:
Bibliographic Details
Title: MADECM: A Curiosity‐Augmented Evolutionary Algorithm for Multi‐Agent Policy Diversity Optimization.
Authors: Wu, Jianyang1 (AUTHOR), Fu, Yv2 (AUTHOR), Wang, Xinning2 (AUTHOR), Yang, Xin2 (AUTHOR) xinyang@dlut.edu.cn
Source: Computer Animation & Virtual Worlds. May/Jun2026, Vol. 37 Issue 3, p1-12. 12p.
Subjects: Evolutionary algorithms, Reinforcement learning, Intrinsic motivation
Abstract: Multi‐agent reinforcement learning (MARL) often suffers from low sample efficiency and limited behavioral diversity, leading to policy homogenization, insufficient exploration, and reduced robustness. To address these challenges, we propose MADECM, a curiosity‐augmented evolutionary framework built upon MADDPG that integrates curiosity‐driven updates with evolutionary quality‐diversity optimization. MADECM employs random network distillation (RND) to estimate the novelty of each agent's local observations and uses the resulting novelty signal to dynamically allocate additional update frequencies, thereby emphasizing exploration‐relevant experience during training. In addition, MADECM combines population‐based diversification with a quality‐diversity (QD) archive through a staged optimization procedure, enabling the joint improvement of task return and policy diversity. We evaluate MADECM on the multi‐agent particle environment (MPE), including Spread and Reference, which capture cooperative and partially observable dynamics, and on google research football (GRF), which emphasizes long‐horizon sequential decision‐making. Results show that MADECM consistently outperforms strong MADDPG‐based baselines. The modular design of MADECM, consisting of RND‐based novelty estimation and staged QD optimization, further supports consistent generalization across these structurally distinct environments without task‐specific hyperparameter tuning. [ABSTRACT FROM AUTHOR]
Copyright of Computer Animation & Virtual Worlds is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Be the first to leave a comment!
You must be logged in first