<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Combining Semantic Modeling and Deep Reinforcement Learning for Autonomous Agents in Minecraft</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Andrew Melnik, Lennart Bramlage, Hendric Voss, Federico Rossetto, Helge Ritter CITEC, Bielefeld University 33619</institution>
          <addr-line>Bielefeld</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>74</fpage>
      <lpage>77</lpage>
      <abstract>
        <p>Figure 1: RGB observation, object detection, depth estimation, top-down map reconstruction. This work describes our extended solution [MARb] for the MARLO¨ (Multi-Agent Reinforcement Learning in MalmO¨ ) Minecraft competition [MARa, PLHM+19, JHHM, GCH+19], and provides insights about how a semantic model of abstract representations of rich sensory data in an environment influence learning and performance in a task. Reinforcement learning can achieve state-of-the-art performance on a broad range of tasks [KMO+18, SM18], however, facing limitations in generalization and sample efficiency. Abstract latent representations and different levels of control can substantially increase sample efficiency [MFSR19]. Learning of a semantic model requires a structured representation of an environment. For that, a preprocessing pipeline can translate rich sensory information into useful abstractions and an object-based representation. A trained model can be utilized to provide navigational goals to the lower-level controller trained with Hindsight Experience Replay (HER) [AWR+17]. Universal value function approximators (UVFAs) and HER [AWR+17] allow efficient learning when a goal is provided. To facilitate semantic inference we borrow from methods successfully employed in robotics and autonomous driving, namely U-Net [RPB15], SLAM, OctoMap [HWB+13] and an object detection module [Goo19]. This allows the agent to consider the locations and spatial relationships of all elements of the environment based solely on RGB-image data and the agent's actions. In the MARLO¨ environment, any action is elementary and performed in exclusion of any other action (move one cell forward, turn left or right 90 degrees). We trained an open-source Keras implementation of U-Net [RPB15] on 841322 pairs of RGB and ground-truth depth images collected from the Minecraft simulator. We used Tensorflow Object Detection Library [Goo19] for tracking the second player, pet and exits in the challenge. Training data was collected in different game sessions by manually labeling 10000 frames. Combined with our action-based SLAM approach, continuous integration of new distance data from the depth estimator results in robust scene reconstructions as 3D Octomap [HWB+13] voxel volumes.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>Entity
Object B
Object B
Object B
Object B
Object B
Object B
Object B
The Octomap is queried every step in the game environment to update the abstract top-down map representation
(Fig. 1).</p>
      <p>We consider a mixed reinforcement learning and semantic learning framework with a pre-processing pipeline.
The pipeline translates RGB-images st ∈ S into an abstract representation zt ∈ Z of a top-down map suitable
for navigation (Fig.1), and tracks objects o ∈ O with the object recognition module. We translate the abstract
representations zt ∈ Z into a Boolean space of concepts ct ∈ C and determine causality of rewarding events in
C ∈ Bc (c is the number of concepts). A concept checks for a set of relationships of objects in the environment
and provides a boolean test result. Along with the test result, each concept collects a relative position of
the tested object. We predefined a set of object-related spatial concepts which are computed for each pair of
objects in the “CatchTheMob” environment at each time step (Table 1). We used human-players demonstrations
(20 episodes) to collect trajectories (st ∈ S), actions, and rewards. The objective is to leverage human-player
demonstrations to acquire a semantic model of rewarding states that enables the agent to select navigational goals
in the environment. When a set of rewarding concepts is selected, their superposition determines a distribution
of positions related to the reward. We feed a sample from this distribution as a goal for the policy that controls
the agent. The policy is trained with HER [AWR+17] to navigate the agent to a goal position in the top-down
map (zt ∈ Z).
3</p>
    </sec>
    <sec id="sec-3">
      <title>Untangling causality</title>
      <p>Here we provide experimental results which highlight the structure of the problem in the ”CatchTheMob”
environment [MARa]. We trained a DQN using the Rainbow implementation [HMVH+18] to catch the pet with
two agents (black curve in Fig. 2) in the ”CatchTheMob” environment [MARa] on the reconstructed top-down
map representations (Fig. 1). The reward function returns +1 when the two agents are at adjacent positions to
the pet but from two opposite sides. The rewarding condition of being at two opposite sides of the pet can be
learned end-to-end or resolved by a higher-level planning model. To measure the potential benefit of the model
we trained the same DQN [HMVH+18] in a simplified single-agent ”CatchTheMob” environment. The reward
function returns +1 when a single agent is at a position adjacent to the pet (blue curve in Fig. 2). We got more
than 50 times faster convergence than in the two-agent case. The red curve shows training of a single player with
HER [AWR+17] to navigate to a given position. We got similar results to the single-agent case. These results
highlight the benefit of higher-level planning through learning causality in the environment.</p>
      <p>To learn the rewarding causality we select concepts activated at the rewarding events (ct → ct+1, r == 1) as
candidates for the necessary (but not sufficient) condition. Necessary concepts are encapsulated into a context
and therefore not sufficient to fully determine the rewarding causality. To determine the sufficient set of concepts
we optimize search by prioritizing concepts with a difference in rewarding / not rewarding samples and y a
localization loss (1).</p>
      <p>Loss = aP(Cx) + bMSE(D)
(1)
P(C) - reward prediction error of the selected set x of concepts C.</p>
      <p>MSE(D) - mean square error of distribution of possible rewarding positions D. a, b - coefficients.
[AWR+17]
[GCH+19]</p>
      <p>Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob
McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba. Hindsight experience replay.
In Advances in Neural Information Processing Systems, pages 5048–5058, 2017.</p>
      <p>William H Guss, Cayden Codel, Katja Hofmann, Brandon Houghton, Noboru Kuno, Stephanie
Milani, Sharada Mohanty, Diego Perez Liebana, Ruslan Salakhutdinov, Nicholay Topin, et al. The
minerl competition on sample efficient reinforcement learning using human priors. arXiv preprint
arXiv:1904.10079, 2019.</p>
      <p>Google Brain. Tensorflow object detection api. https://github.com/tensorflow/models/tree/
master/research/object_detection, 2019. Accessed: 2019-09-10.
Armin Hornung, Kai M Wurm, Maren Bennewitz, Cyrill Stachniss, and Wolfram Burgard.
Octomap: An efficient probabilistic 3d mapping framework based on octrees. Autonomous robots,
34(3):189–206, 2013.</p>
      <p>Matthew Johnson, Katja Hofmann, Tim Hutton, and David Bignell Microsoft. The Malmo Platform
for Artificial Intelligence Experimentation *. Technical report.
[KMO+18] Lukasz Kidzin´ski, Sharada Prasanna Mohanty, Carmichael F Ong, Zhewei Huang, Shuchang Zhou,
Anton Pechenko, Adam Stelmaszczyk, Piotr Jarosik, Mikhail Pavlov, Sergey Kolesnikov, et al.
Learning to run challenge solutions: Adapting reinforcement learning methods for
neuromusculoskeletal environments. In The NIPS’17 Competition: Building Intelligent Systems, pages 121–153.
Springer, 2018.</p>
      <p>Crowdai marlo: Multi-agent reinforcement learning in minecraft. https://www.crowdai.org/
challenges/marlo-2018.</p>
      <p>Github - marlo solution. http://rebrand.ly/GitHub-MARLO.</p>
      <p>Andrew Melnik, Sascha Fleer, Malte Schilling, and Helge Ritter. Modularization of end-to-end
learning: Case study in arcade games. arXiv preprint arXiv:1901.09895, 2019.
[PLHM+19] Diego Perez-Liebana, Katja Hofmann, Sharada Prasanna Mohanty, Noburu Kuno, Andre Kramer,
Sam Devlin, Raluca D. Gaina, and Daniel Ionita. The Multi-Agent Reinforcement Learning in
Malm\”O (MARL\”O) Competition. jan 2019.
[SM18]</p>
      <p>Malte Schilling and Andrew Melnik. An approach to hierarchical deep reinforcement learning for a
decentralized walking control architecture. In Biologically Inspired Cognitive Architectures Meeting,
pages 272–282. Springer, 2018.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>