A Model-Based Solution to the Offline Multi-Agent Reinforcement Learning Coordination Problem

05/26/2023
by   Paul Barde, et al.
0

Training multiple agents to coordinate is an important problem with applications in robotics, game theory, economics, and social sciences. However, most existing Multi-Agent Reinforcement Learning (MARL) methods are online and thus impractical for real-world applications in which collecting new interactions is costly or dangerous. While these algorithms should leverage offline data when available, doing so gives rise to the offline coordination problem. Specifically, we identify and formalize the strategy agreement (SA) and the strategy fine-tuning (SFT) challenges, two coordination issues at which current offline MARL algorithms fail. To address this setback, we propose a simple model-based approach that generates synthetic interaction data and enables agents to converge on a strategy while fine-tuning their policies accordingly. Our resulting method, Model-based Offline Multi-Agent Proximal Policy Optimization (MOMA-PPO), outperforms the prevalent learning methods in challenging offline multi-agent MuJoCo tasks even under severe partial observability and with learned world models.

READ FULL TEXT

page 21

page 23

research
06/07/2021

Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning

Learning from datasets without interaction with environments (Offline Le...
research
11/28/2022

Learning From Good Trajectories in Offline Multi-Agent Reinforcement Learning

Offline multi-agent reinforcement learning (MARL) aims to learn effectiv...
research
05/27/2023

MADiff: Offline Multi-agent Learning with Diffusion Models

Diffusion model (DM), as a powerful generative model, recently achieved ...
research
06/26/2018

Learning Existing Social Conventions in Markov Games

In order for artificial agents to coordinate effectively with people, th...
research
11/22/2019

Thompson Sampling for Factored Multi-Agent Bandits

Multi-agent coordination is prevalent in many real-world applications. H...
research
11/22/2019

Multi-Agent Thompson Sampling for Bandit Applications with Sparse Neighbourhood Structures

Multi-agent coordination is prevalent in many real-world applications. H...
research
07/19/2021

DeepCC: Bridging the Gap Between Congestion Control and Applications via Multi-Objective Optimization

The increasingly complicated and diverse applications have distinct netw...

Please sign up or login with your details

Forgot password? Click here to reset