Using a paper from Google DeepMind I've developed a new version of the DQN using threads exploration instead of memory replay as explain in here: http://arxiv.org/pdf/1602.01783v1.pdf I used the one-step-Q-learning pseudocode, and now we can train the Pong game in less than 20 hours and without any GPU or network distribution.
Hi zeta:
I saw your reply and came here to have a look.
This is so cool! I really need a method that can avoid memory replay since the memory space is a big problem and time consuming.
Thank you. I will look into it further. Mingyan