TY - GEN
T1 - Cooperative UAV search and rescue via multi-agent reinforcement learning in simulated wildfire environments
AU - Sharma, Shivani
AU - Tsoumplekas, Georgios
AU - Spyridis, Yannis
AU - Vitzilaios, Nikolaos
AU - Argyriou, Vasileios
PY - 2026/7/14
Y1 - 2026/7/14
N2 - Wildfires pose an increasing global threat, endangering both human and animal lives. Rapid and coordinated search and rescue (SAR) operations are critical to minimizing casualties in such emergencies. This paper investigates the use of Multi-Agent Reinforcement Learning (MARL) to train autonomous unmanned aerial vehicles (UAVs) capable of cooperative SAR in simulated wildfire environments. The task is modeled as a decentralized partially observable Markov decision process (Dec-POMDP) and trained under a Centralized Training with Decentralized Execution (CTDE) paradigm. Two learning configurations are compared: a single-agent baseline using Proximal Policy Optimization (PPO) and a cooperative multi-agent framework based on Multi-Agent Policy Optimization with Credit Assignment (MA-POCA) incorporating posthumous credit assignment. Training employs a three-stage curriculum to progressively increase environmental complexity and enhance policy generalization. Simulations across one to six UAVs demonstrate that multi-agent coordination significantly improves mission efficiency and consistency. Specifically, teams of four to five UAVs achieved the lowest average completion times while maintaining high stability and reliability across trials. These results confirm that MARL-based cooperative control improves scalability, robustness and overall mission performance in UAV-based SAR operations, especially under optimal team sizing, underscoring the potential of decentralized learning for real-world disaster response scenarios.
AB - Wildfires pose an increasing global threat, endangering both human and animal lives. Rapid and coordinated search and rescue (SAR) operations are critical to minimizing casualties in such emergencies. This paper investigates the use of Multi-Agent Reinforcement Learning (MARL) to train autonomous unmanned aerial vehicles (UAVs) capable of cooperative SAR in simulated wildfire environments. The task is modeled as a decentralized partially observable Markov decision process (Dec-POMDP) and trained under a Centralized Training with Decentralized Execution (CTDE) paradigm. Two learning configurations are compared: a single-agent baseline using Proximal Policy Optimization (PPO) and a cooperative multi-agent framework based on Multi-Agent Policy Optimization with Credit Assignment (MA-POCA) incorporating posthumous credit assignment. Training employs a three-stage curriculum to progressively increase environmental complexity and enhance policy generalization. Simulations across one to six UAVs demonstrate that multi-agent coordination significantly improves mission efficiency and consistency. Specifically, teams of four to five UAVs achieved the lowest average completion times while maintaining high stability and reliability across trials. These results confirm that MARL-based cooperative control improves scalability, robustness and overall mission performance in UAV-based SAR operations, especially under optimal team sizing, underscoring the potential of decentralized learning for real-world disaster response scenarios.
U2 - 10.1109/ICUAS69441.2026.11598622
DO - 10.1109/ICUAS69441.2026.11598622
M3 - Conference contribution
AN - SCOPUS:105045631754
SN - 9798331593179
T3 - International Conference on Unmanned Aircraft Systems (ICUAS)
SP - 828
EP - 835
BT - 2026 International Conference on Unmanned Aircraft Systems, ICUAS 2026
PB - Institute of Electrical and Electronics Engineers, Inc.
T2 - 2026 International Conference on Unmanned Aircraft Systems
Y2 - 15 June 2026 through 18 June 2026
ER -