arXiv · 2301.03679
Transformers as Policies for Variable Action Environments
Abstract
In this project we demonstrate the effectiveness of the transformer encoder as a viable architecture for policies in variable action environments. Using it, we train an agent using Proximal Policy Optimisation (PPO) on multiple maps against scripted opponents in the Gym-$\mu$RTS environment. The final agent is able to achieve a higher return using half the computational resources of the next-best RL agent, which used the GridNet architecture. The source code and pre-trained models are available here: https://github.com/NiklasZ/transformers-for-variable-action-envs
Explore related subjects
Keep this discovery
Niklas Zwingenberger. 2023-01-09. Transformers as Policies for Variable Action Environments. https://arxiv.org/abs/2301.03679
Cite the original work for its findings. Save a collection to share your selection of sources.