Plato Data Intelligence.
Vertical Search & Ai.

Continual Reinforcement Learning with Multi-Timescale Replay. (arXiv:2004.07530v1 [cs.LG])

Date:

[Submitted on 16 Apr 2020]

Download PDF

Abstract: In this paper, we propose a multi-timescale replay (MTR) buffer for improving
continual learning in RL agents faced with environments that are changing
continuously over time at timescales that are unknown to the agent. The basic
MTR buffer comprises a cascade of sub-buffers that accumulate experiences at
different timescales, enabling the agent to improve the trade-off between
adaptation to new data and retention of old knowledge. We also combine the MTR
framework with invariant risk minimization, with the idea of encouraging the
agent to learn a policy that is robust across the various environments it
encounters over time. The MTR methods are evaluated in three different
continual learning settings on two continuous control tasks and, in many cases,
show improvement over the baselines.

Submission history

From: Christos Kaplanis [view email]
[v1]
Thu, 16 Apr 2020 08:47:40 UTC (4,451 KB)

Source: http://arxiv.org/abs/2004.07530

spot_img

Latest Intelligence

spot_img

Chat with us

Hi there! How can I help you?