Search
2019 Volume 34
Article Contents
RESEARCH ARTICLE   Open Access    

Team learning from human demonstration with coordination confidence

More Information
  • Abstract: Among an array of techniques proposed to speed-up reinforcement learning (RL), learning from human demonstration has a proven record of success. A related technique, called Human-Agent Transfer, and its confidence-based derivatives have been successfully applied to single-agent RL. This article investigates their application to collaborative multi-agent RL problems. We show that a first-cut extension may leave room for improvement in some domains, and propose a new algorithm called coordination confidence (CC). CC analyzes the difference in perspectives between a human demonstrator (global view) and the learning agents (local view) and informs the agents’ action choices when the difference is critical and simply following the human demonstration can lead to miscoordination. We conduct experiments in three domains to investigate the performance of CC in comparison with relevant baselines.
  • 加载中
  • Argall , B. D., Chernova , S., Veloso , M. & Browning , B. 2009. A survey of robot learning from demonstration. Robotics and Autonomous Systems 57(5), 469–483. http://dx.doi.org/10.1016/j.robot.2008.10.024

    Google Scholar

    Chernova , S. & Veloso , M. 2007. Multiagent collaborative task learning through imitation. In Proceedings of the 4th International Symposium on Imitation in Animals and Artifacts (AIBS-07), Artificial and Ambient Intelligence.

    Google Scholar

    da Silva , F. L., Glatt , R. & Costa , A. H. R. 2017. Simultaneously learning and advising in multiagent reinforcement learning. In Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems (AAMAS-17).

    Google Scholar

    Fernandez , F., Garcia , J. & Veloso , M. 2010. Probabilistic policy reuse for inter-task transfer learning. Robotics and Autonomous Systems 58(7), 866–871.

    Google Scholar

    Fudenberg , D. & Levine , K. 1998. The Theory of Learning in Games. MIT Press.

    Google Scholar

    Kraemer , L. & Banerjee , B. 2016. Multi-agent reinforcement learning as a rehearsal for decentralized planning. Neurocomputing 190, 82–94.

    Google Scholar

    Le , H. M., Yue , Y., Carr , P. & Lucey , P. 2017. Coordinated multi-agent imitation learning. In Proceedings of the 34th International Conference on Machine Learning (ICML-17).

    Google Scholar

    MacGlashan , J. 2014. The Brown-UMBC reinforcement learning and planning (BURLAP) library, http://burlap.cs.brown.edu/

    Google Scholar

    Mnih , V., Kavukcuoglu , K., Silver , D., Rusu , A. A., Veness , J., Bellemare , M. G., Graves , A., Riedmiller , M., Fidjeland , A. K., Ostrovski , G., Petersen , S., Beattie , C., Sadik , A., Antonoglou , I., King , H., Kumaran , D., Wierstra , D., Legg , S. & Hassabis , D. 2015. Human-level control through deep reinforcement learning. Nature 518, 529–533.

    Google Scholar

    Silver , D., Huang , A., Maddison , C. J., Guez , A., Sifre , L., van den Driessche , G., Schrittwieser , J., Antonoglou , I., Panneershelvam , V., Lanctot , M., Dieleman , S., Grewe , D., Nham , J., Kalchbrenner , N., Sutskever , I., Lillicrap , T., Leach , M., Kavukcuoglu , K., Graepel , T. & Hassabis , D. 2016. Mastering the game of Go with deep neural networks and tree search. Nature 529, 484–489.

    Google Scholar

    Song , J., Ren , H., Sadigh , D. & Ermon , S. 2018. Multi-Agent Generative Adversarial Imitation Learning. In Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS 2018).

    Google Scholar

    Sutton , R. & Barto , A. G. 1998. Reinforcement Learning: An Introduction, MIT Press.

    Google Scholar

    Taylor , M. E. & Stone , P. 2009. Transfer learning for reinforcement learning domains: A survey. Journal of Machine Learning Research 10(1), 1633–1685.

    Google Scholar

    Taylor , M. E., Suay , H. B. & Chernova , S. 2011. Integrating reinforcement learning with human demonstrations of varying ability. In Proceedings of the International Conference on Autonomous Agents and Multiagent Systems (AAMAS).

    Google Scholar

    Wang , Z. & Taylor , M. E. 2017. Improving reinforcement learning with confidence-based demonstrations. In Proceedings of the 26th International Conference on Artificial Intelligence (IJCAI).

    Google Scholar

    Wang , Z. & Taylor , M. E. 2019, Interactive reinforcement learning with dynamic reuse of prior knowledge from human/agent’s demonstration. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI).

    Google Scholar

  • Cite this article

    Bikramjit Banerjee, Syamala Vittanala, Matthew Edmund Taylor. 2019. Team learning from human demonstration with coordination confidence. The Knowledge Engineering Review. 34:43 doi: 10.1017/S0269888919000043
    Bikramjit Banerjee, Syamala Vittanala, Matthew Edmund Taylor. 2019. Team learning from human demonstration with coordination confidence. The Knowledge Engineering Review. 34:43 doi: 10.1017/S0269888919000043

Article Metrics

Article views(15) PDF downloads(246)

RESEARCH ARTICLE   Open Access    

Team learning from human demonstration with coordination confidence

The Knowledge Engineering Review  34 Article number: e12  (2019)  |  Cite this article

Abstract: Abstract: Among an array of techniques proposed to speed-up reinforcement learning (RL), learning from human demonstration has a proven record of success. A related technique, called Human-Agent Transfer, and its confidence-based derivatives have been successfully applied to single-agent RL. This article investigates their application to collaborative multi-agent RL problems. We show that a first-cut extension may leave room for improvement in some domains, and propose a new algorithm called coordination confidence (CC). CC analyzes the difference in perspectives between a human demonstrator (global view) and the learning agents (local view) and informs the agents’ action choices when the difference is critical and simply following the human demonstration can lead to miscoordination. We conduct experiments in three domains to investigate the performance of CC in comparison with relevant baselines.

    • We thank the anonymous reviewers for constructive comments and suggestions. This work was supported in part by National Science Foundation grant IIS-1526813.

    • While the original DRoP paper includes two types of temporal difference confidence measures and three types of action decision models, we focus on one, each.

    • This is called the Dynamic Confidence Update method in the original DRoP paper.

    • This is called the soft decision model in the original DRoP paper.

    • © Cambridge University Press, 2019 2019Cambridge University Press
References (16)
  • About this article
    Cite this article
    Bikramjit Banerjee, Syamala Vittanala, Matthew Edmund Taylor. 2019. Team learning from human demonstration with coordination confidence. The Knowledge Engineering Review. 34:43 doi: 10.1017/S0269888919000043
    Bikramjit Banerjee, Syamala Vittanala, Matthew Edmund Taylor. 2019. Team learning from human demonstration with coordination confidence. The Knowledge Engineering Review. 34:43 doi: 10.1017/S0269888919000043
  • Catalog

      /

      DownLoad:  Full-Size Img  PowerPoint
      Return
      Return