Özet
Reinforcement learning tree-based planning methods have been gaining popularity in the last few years due to their success in single-agent domains, where a perfect simulator model is available, e.g., Go and chess strategic board games. This paper pretends to extend tree search algorithms to the multi-agent setting in a decentralized structure, dealing with scalability issues and exponential growth of computational resources. The N-Step Dynamic Tree Search combines forward planning and direct temporal-difference updates, outperforming markedly state-of-the-art algorithms such as Q-Learning and SARSA. Future state transitions and rewards are predicted with a model built and learned from real interactions between agents and the environment. As an extension of previous work, this paper analyses the developed algorithm in the Hunter-Pursuit cooperative game against intelligent evaders. The N-Step Dynamic Tree Search aims to adapt the most successful single-agent learning methods to the multi-agent boundaries and demonstrates to be a remarkable advance compared to conventional temporal-difference techniques.
| Orijinal dil | İngilizce |
|---|---|
| Ana bilgisayar yayını başlığı | 2022 American Control Conference, ACC 2022 |
| Yayınlayan | Institute of Electrical and Electronics Engineers Inc. |
| Sayfalar | 761-766 |
| Sayfa sayısı | 6 |
| ISBN (Elektronik) | 9781665451963 |
| DOI'lar | |
| Yayın durumu | Yayınlandı - 2022 |
| Harici olarak yayınlandı | Evet |
| Etkinlik | 2022 American Control Conference, ACC 2022 - Atlanta, United States Süre: 8 Haz 2022 → 10 Haz 2022 |
Yayın serisi
| Adı | Proceedings of the American Control Conference |
|---|---|
| Hacim | 2022-June |
| ISSN (Basılı) | 0743-1619 |
???event.eventtypes.event.conference???
| ???event.eventtypes.event.conference??? | 2022 American Control Conference, ACC 2022 |
|---|---|
| Ülke/Bölge | United States |
| Şehir | Atlanta |
| Periyot | 8/06/22 → 10/06/22 |
Bibliyografik not
Publisher Copyright:© 2022 American Automatic Control Council.
Finansman
*Research supported by Engineering and Physical Sciences Research Council (EPSRC) and BAE Systems under the project reference no. 2454254. Marc Espinós Longa is a PhD Researcher in the School of Aerospace, Transport & Manufacturing at Cranfield University, Bedfordshire, MK43 0AL, United Kingdom (e-mail: [email protected]). Antonios Tsourdos is an AIAA Senior Member, Head of Center and Director of Research in the School of Aerospace, Transport & Manufacturing
| Finansörler | Finansör numarası |
|---|---|
| BAE Systems | 2454254 |
| Engineering and Physical Sciences Research Council |
Parmak izi
Swarm Intelligence in Cooperative Environments: N-Step Dynamic Tree Search Algorithm Extended Analysis' araştırma başlıklarına git. Birlikte benzersiz bir parmak izi oluştururlar.Alıntı Yap
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver