Brockman , G., Cheung , V., Pettersson , L., Schneider , J., Schulman , J., Tang , J. & Zaremba , W. 2016. Openai gym. arXiv preprint arXiv:1606.01540.

Brys , T., Harutyunyan , A., Suay , H. B., Chernova , S., Taylor , M. E. & Nowé , A. 2015. Reinforcement learning from demonstration through shaping. In Proceedings of the 24th International Conference on Artificial Intelligence, 3352–3358. AAAI Press.

Devlin , S. & Kudenko , D. 2012. Dynamic potential-based reward shaping. In Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems-Volume 1, 433–440. International Foundation for Autonomous Agents and Multiagent Systems.

Dietterich , T. G. 2000. Hierarchical reinforcement learning with the MAXQ value function decomposition. Journal of Artificial Intelligence Research(JAIR) 13, 227–303.

Fernández , F. & Veloso , M. 2006. Probabilistic policy reuse in a reinforcement learning agent. In Proceedings of the Fifth International Joint Conference on Autonomous Agents and Multiagent Systems, 720–727. ACM.

Hester , T., Vecerik , M., Pietquin , O., Lanctot , M., Schaul , T., Piot , B., Horgan , D., Quan , J., Sendonaris , A., Dulac-Arnold , G., Osband , I., Agapiou , J. 2018. Deep Q-learning from demonstrations. In Thirty-Second AAAI Conference on Artificial Intelligence 2018 Apr 29.

Kaelbling , L. P. 1993. Hierarchical learning in stochastic domains: Preliminary results. In Proceedings of the Tenth International Conference on Machine Learning, 951, 167–173.

Kingma , D. P. & Ba , J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.

Mnih , V., Kavukcuoglu , K., Silver , D., Rusu , A. A., Veness , J., Bellemare , M. G., Graves , A., Riedmiller , M., Fidjeland , A. K., Ostrovski , G., Petersen , S. 2015. Human-level control through deep reinforcement learning. Nature 518(7540), 529–533.

Ng , A. Y., Harada , D. & Russell , S. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In ICML, 99, 278–287.

Schaal , S. 1997. Learning from demonstration. In Proceedings of the 1997 Conference on Neural Information Processing Systems (NIPS97). Denver, CO, pp. 1040–1046.

Sutton , R. S. & Barto , A. G. 1998. Reinforcement Learning: An Introduction, 1. MIT press.

Taylor , M. E., Suay , H. B. & Chernova , S. 2011. Integrating reinforcement learning with human demonstrations of varying ability. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2, 617–624. International Foundation for Autonomous Agents and Multiagent Systems.

Wang , Z. & Taylor , M. E. 2017. Improving reinforcement learning with confidence-based demonstrations. In Proceedings of the 26th International Conference on Artificial Intelligence (IJCAI).

Watkins , C. J. & Dayan , P. 1992. Q-learning. Machine Learning 8(3–4), 279–292.

Wiewiora , E., Cottrell , G. W. & Elkan , C. 2003. Principled methods for advising reinforcement learning agents. In Proceedings of the 20th International Conference on Machine Learning (ICML-03), 792–799.