A3C-LSTM

A Python implementation of the A3C reinforcement learning algorithm with an LSTM network, tested on the CartPole environment. The readme notes it does not converge and points to a DDPG version.

Share on XNo license declared

Overview

A Python implementation of the Asynchronous Advantage Actor-Critic (A3C) reinforcement learning algorithm using an LSTM network, tested on the CartPole environment from OpenAI Gym. It builds on a TensorFlow tutorial by Arthur Juliani and the 2016 paper by Mnih et al. The author states it is example code only and that this model does not converge on CartPole.

Key features

  • A3C agent with LSTM layers, built with TensorFlow
  • Trains only on minibatches larger than 30
  • Uses a reward factor to allow faster learning rates
  • Saves models every 100 episodes for reloading or testing

Best for

Readers who want to study how A3C with an LSTM is put together. The readme warns that it does not converge and points to a DDPG version for a working model.

Upstream
liampetti/A3C-LSTM
Fork on GitHub
Guo-astro/A3C-LSTM
Upstream stars
48
Category
Finance, data and research
Language
Python
License
No license declaredWithout a license, the author keeps all rights. Ask the upstream owner before reusing the code.
Forked
2018-05-20
Sync status
In syncLast synced 2026-09-29