Simulated multi-agent RL credit-assignment-error trajectories and coordination outcomes for 1575 configurations across five India-relevant multi-agent tasks (traffic signal networks, warehouse robot fleets, agri-drone swarms, disaster response, power grid control), seven coordination methods, agent count, and communication conditions, grounded in MADDPG, VDN, QMIX, COMA, MAPPO, MAAC, and AlphaStar. Forward-looking task predicts end-of-training coordination failure from only an early-training che
This dataset simulates the self-reinforcing non-stationarity problem in multi-agent RL: from any one agent's perspective the environment keeps changing because every other agent is also learning. The simulation spans 1575 configurations across five India-relevant tasks, seven coordination methods spanning a naive independent-learners baseline through six methods with published credit-assignment or centralized-training mechanisms, and agent counts from 2 to 32. Twelve grounding sources are documented in marl_theory.py and the README Research basis, including QMIX, MAPPO, MAAC, and the AlphaStar Nature paper.
Lets Teams Training Multi-agent Systems For Traffic Control, Drone Swarms, Or Warehouse Fleets Estimate, From An Early-training Checkpoint, Whether A Given Architecture And Agent-count Combination Is Headed For A Coordination Failure.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.