Projection Exploration for Multi-agent Reinforcement Learning
preprint
OA: closed
Abstract
Abstract In multi-agent reinforcement learning (MARL), complete exploration is difficult to achieve because of the curse of dimensionality and sparse rewards. Existing methods improve the exploration to some extent, such as noise-based exploration methods and count-based exploration methods. But they don’t evaluate the state exploration value, suffering low data efficiency. In this paper, a projection exploration algorithm (PERL) is designed for MARL. In this method, states with high exploration value are identified in the projected state-action space with maximum distribution entropy through the count-based method. Then the agents are trained to reach these states in a coordinated manner. Meanwhile, comparison experiments are conducted in multi-agent environments of Multi-Particle environments (Pass, Secret-Room) and StarCraftII (3m,2s3z and 3s_vs_5z). Comparison results suggest that PERL consistently outperforms the state-of-the-art baselines in environments with high-dimensional state-action space and sparse rewards.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00