<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Avoiding obstacles in a road reinforcement learning methods</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohamed Saber Rais</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rachid Boudour</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khouloud Zouaidia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Badji Mokhtar Annaba</institution>
          ,
          <country country="DZ">Algeria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>21</fpage>
      <lpage>22</lpage>
      <abstract>
        <p>Decision making has always been a challenge in all fields and many methods and approaches were applied to solve different kinds of problems in different situations mainly to reach a high level of autonomy in those tasks. In this paper we addressed the autonomous driving problem in a simple way by defining our own environment which represents a road similar to a maze with static obstacles or they can be seen in another way as cars with stable velocities, furthermore we trained each agent in this environment to be able to learn a policy that can let them map between the input data such as the environment state and the actions taken while achieving the best rewards. Using both rule based restrictions and different reinforcement and deep reinforcement learning algorithms we were able to achieve satisfying results in this matter using all five of our agents: the q network agent, the SARSA agent, the deep q network agent, the double deep q network agent and the deep SARSA agent, and we were able to show that the proposed algorithm “deep SARSA with rule based restrictions” can achieve similar or even better results than the q network related methods that are mainly used in autonomy problems.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Decision making</kwd>
        <kwd>Reinforcement learning</kwd>
        <kwd>Deep learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Controlling an agent to do specific tasks holds many challenges and was a field of research for
many years.</p>
      <p>At first, researchers addressed simpler goals that required just good and clear rules to follow in
different situations, this represents the rule based category of the methods that has been used in the
last years. But this approach reached its limits after turning our sights to harder tasks in many fields.
This is when the notion of learning came to light and we started training those agents to be able to
achieve better results in more complicated environments and restrictions.</p>
      <p>As for the decision making field, approaches such as deep learning and reinforcement learning were
used mainly to be able to reach a human like performance under the “trial and error” method. The last
mentioned method let the agent interacts with its surrounding and learn with each attempt what can be
best to do in a specific situation till it can find a generalization to what can be done in most situations
it will be facing.</p>
      <p>This is pretty much what a human does when he tries to achieve a good result in an unknown mission.
For example what are addressing now “driving a car”, in general with every ride the human takes he
learns to adapt more to control the vehicle and can react better to different situations. Same as in this
paper with using different techniques of reinforcement and deep Reinforcement learning such as Q
network, SARSA, deep Q learning, double deep Q learning and Deep SARSA, each agent learned the
basics of turning left and right when an obstacle is in front to avoid an accident and reach the
destination.</p>
      <p>The upcoming parts will represent our work with all the steps taken. Section 2 informs about some
prior works related to our subject, section 3 is a background about some concepts we used in our
article, section 4 presents a definition to the problem encountered, section 5 contains the results that
we reached for every agent, section 6 involves a comparison between the agents used, section 7 holds
our final results and finally a conclusion about this subject.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Prior Works</title>
      <p>
        The task of decision making was addressed before in many fields and in a lot of dif-ferent ways.
We can mention:
In 2013, Mnih V. et al combined deep learning with reinforcement learning to learn a policy that can
let the agent play the “Atari game” with human level performance [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Another example is the game of “Go” where the authors in 2017 were able to compete against high
level players and a variety of computer opponents in this game-board using a reinforcement learning
agent [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Focusing now on works related to autonomous vehicles:
In 2018, Carl-Johan H. et al used the double deep Q network technique as the base for their agent to
reach the wanted level of automation in controlling the car’s speed alongside changing the lane in a
highway scenario [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        In 2019, Junjie W. et al combined a high level lateral decision making with low level rule based
trajectory modification in a method based on deep Q network to solve the lane changing problem for
the autonomous vehicles [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        In 2019, Laura García C. et al proposed an approach based on using the Q learning algorithm to train
their agent in 2 scenarios: the first one was navigating a roundabout with traffic and the second one
without traffic. The agent was able to learn smooth and efficient driving to perform maneuvers within
roundabouts [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Background</title>
      <p>3.1.</p>
      <p>Reinforcement learning
is one of the three main paradigms of machine learning alongside supervised learning and
unsupervised learning. The goal of reinforcement learning is to learn a policy which maps from the
input state to the actions taken by getting the best reward that can be achieved. To obtain this policy
reinforcement learning uses a method that maximizes a total expected return (called G) which is a
cumulative sum of immediate rewards received over the long run with each action the agent takes. It
is defined as follow:
∞
 =0
  =</p>
      <p>∑     +</p>
    </sec>
    <sec id="sec-4">
      <title>Deep Reinforcement learning</title>
      <p>
        It uses the principles of both deep learning and reinforcement learning to reach solutions to many
problems in different fields but mainly in the autonomy field [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
3.3.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Markov Decision Process</title>
      <p>
        It is a discrete time stochastic control process. It provides a mathematical framework for modeling
decision making in situations where outcomes are partly random and partly under the control of a
decision maker [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
3.4.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Q learning technique</title>
      <p>
        Agents based on this technique use a predefined table which will hold the q values for each (state,
action) pair. The q table is updated after each action the agent takes. When taking the action the agent
always picks the one with the highest q value in the table for each state [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
3.5.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Deep Q learning technique</title>
      <p>
        This technique has the same concept of the previous one when it comes to taking actions. But
instead of a Q table to save the (state, action) pair, a deep learning model is used to save episodes and
let the agent learn which action is best to take in a specific situation after checking the saved episodes
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
3.6.
      </p>
    </sec>
    <sec id="sec-8">
      <title>Double deep Q learning technique</title>
      <p>
        In Double Deep Q Learning the agent uses two neural networks to learn and predict what action to
take at every step [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
3.7.
      </p>
    </sec>
    <sec id="sec-9">
      <title>SARSA technique</title>
      <p>Same as the q learning technique mentioned before agents based on SARSA use a predefined table
that will hold the q values for each (state, action) pair and update this table after each action the agent
applied.</p>
      <p>
        But when taking every action the agent uses a ɛ-greedy policy. This means that it can take either a
random action that can possibly give some new better results or it can pick the action with maximum
q value [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
3.8.
      </p>
    </sec>
    <sec id="sec-10">
      <title>Deep SARSA</title>
      <p>
        Same as the deep q network, agents based on deep SARSA use a deep learning model to save each
episode the agent pass by instead of a Q table and when it comes to taking the actions at each state it
uses the same concept of the SARSA technique [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
    </sec>
    <sec id="sec-11">
      <title>4. Problem definition</title>
      <p>We decided to use a maze representation for our road to let the agents learn the basic actions that
can be taken. Thus the problem can be defined as a Markov Decision Process (MDP) where all states
are known to the agent.</p>
      <p>The agents receive at each discrete time step “t” the state “s” of the road and then select an action
to take “a” that leads to the next state “s+1” and saves the reward “r” for taking this action, this
continues till an agent reaches a final state, either reaching its goal or losing.
4.1.</p>
    </sec>
    <sec id="sec-12">
      <title>State space definition</title>
      <p>At each time step the agent receives the state of the environment, which in our case is represented
with a 2D vector with 3 rows (the figure below shows an example of our environment). It contains the
different obstacles and the current placement of our agent after taking any possible action. The
obstacles cells were represented with 0 and the free spaces where the agent can travels with 1, our
agent was represented with a 0.5 grey color degree to differentiate it from other occupied cells.</p>
    </sec>
    <sec id="sec-13">
      <title>Action space definition</title>
      <p>Action space is a 1D vector containing 3 actions: forward, left and right, so the agent can move
forward or turn left or right at every obstacle. This basically simulates a highway situation where the
agent should not move backwards.
4.3.</p>
    </sec>
    <sec id="sec-14">
      <title>Rules and restriction</title>
      <p>For a faster and smoother learning for each agent, we defined some rules in this environment that
restricts taking some meaningless actions in some situations (getting off the road) and let the agent
focus only on avoiding the obstacles.</p>
      <p>
        Algorithm 1: Rules Function
Actions = [
        <xref ref-type="bibr" rid="ref1 ref2">0, 1, 2</xref>
        ] # actions vector [forward right left]
if row is left line of road:
      </p>
      <p>actions.remove(2)
elif row is right line of the road:</p>
      <p>actions.remove(1)
if reached the end of the road:</p>
      <p>actions.remove(0)
Example: if our agent is localized at the row with index 0 the action turn left will be re-moved from
the actions space so we avoid letting the agent go off road. This will let us skip some episodes where
our agent will try to learn actions that are not related to avoiding obstacles and this will lead to a faster
learning.
4.4.</p>
    </sec>
    <sec id="sec-15">
      <title>Rewards</title>
      <p>We defined a simple reward function based on 5 situations the agent can be in after taking an
action. These situations are: reached the goal, agent is blocked, agent in a visited location, invalid
action taken and valid action taken.</p>
      <p>Algorithm 2: Reward Function
if agent reached its destination :</p>
      <p>return 1.0
if mode == 'blocked': # no more actions can be taken by the agent</p>
      <p>return self.min_reward – 1
if agent in self.visited : # agent keep turning left and right(no advance)</p>
      <p>return -0.25
if mode == 'invalid':</p>
      <p>return -0.75
if mode == 'valid':
return -0.04
# agent picked an action that leads to an accident
# valid action
In our reward function penalizing the agent with “min_reward – 1” when it is in a blocked situation
will end directly the episode and starts a new one.</p>
      <p>We set a penalization to our agent even when it takes a valid action that leads it to follow the longest
road to reach its destination to let the agent knows that it must reach its destination with the least
actions possible. In other words, it should try to pick the smallest series of actions that will lead it to
reaching its destination while avoiding actions that will lead to consuming more time to finish the
episode.</p>
    </sec>
    <sec id="sec-16">
      <title>5. Our reinforcement learning and deep R learning agents 5.1.</title>
    </sec>
    <sec id="sec-17">
      <title>Reinforcement learning agents</title>
    </sec>
    <sec id="sec-18">
      <title>5.1.1. Q learning agent</title>
      <p>In our experiments we let the agent learn for a total number of 500 episodes for 10 rounds to test
the consistency of our results and this led clearly to updating our q table completely with the best
actions that can be taken in each state of the environment. We achieved an average reward equal to:
-1.1896200000000205 while reaching the best reward possible a lot earlier than expected at around
episode number 40 in each round.</p>
    </sec>
    <sec id="sec-19">
      <title>5.1.2. SARSA agent</title>
      <p>Same as the q learning agent e let this agent learn for a total number of 500 episodes for 10 rounds
too. This led to an average reward equal to: -1.13894000000002 while it reached the best reward
possible around the episode 50 each time.</p>
    </sec>
    <sec id="sec-20">
      <title>Deep reinforcement learning agents</title>
      <p>We decided to use a simple deep learning model shared between all algorithms mentioned below
to see the difference between the performances of these algorithms alone without the interference of
the model’s architecture.</p>
      <p>
        The model consisted of an input layer that takes the shape of our road as an input and uses the
“RELU” activation function with 512 nodes. We used 2 hidden layers for our model each one of them
has in order 256 and 64 nodes. The output layer takes the actions vector as an output [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-21">
      <title>5.2.1. Deep Q learning agent</title>
      <p>The agent was trained for a total of 500 episodes for 10 times. It was able to achieve the max
reward that can be reached in an average number of 250 episodes. The time taken for each round of
training was around 9.20 minutes and the agent’s loss improved by time. Same as the reward that kept
getting bigger till around episode 265 where it got nearly stable with achieving the best reward
possible in every run.</p>
    </sec>
    <sec id="sec-22">
      <title>5.2.2. Double deep Q learning agent</title>
      <p>Same number of episodes and rounds were applied to this agent. It was able to reach the max
reward for the first time at the episode 185, with an average of 8.46 minutes for each round. The total
reward started getting stable when reaching episode 300 with a really good loss values around that
period.</p>
    </sec>
    <sec id="sec-23">
      <title>5.2.3. Deep SARSA agent</title>
      <p>Lastly we applied the same training amount to the deep SARSA agent that was also able to reach
the best achievable reward and this was in general reached at episode 170. The time this agent took
for each round of training was close to the previous ones and it was around 7.92 minutes. The loss
values got better and better with every episode passing and we were able to reach a nearly stable
reward with the best value each time when we got passed episode 230.</p>
    </sec>
    <sec id="sec-24">
      <title>6. Comparison between the different agents 6.1.</title>
    </sec>
    <sec id="sec-25">
      <title>Comparison between Reinforcement learning agents</title>
      <p>The comparison between the q network agent and the SARSA agent was based on 3 qualities: the
time consumed for the training in minutes (TT), average reward (Avg-R) and the last one is the
episodes needed to get the best reward (E-BR).</p>
      <p> For the TT quality: whenever the time consumed for the training is smaller it means the
agent performed better.
 For the Avg-R: we got it by summing all the reward of each episode the agent finished
then dividing by the number of episodes. The agent that will reach the biggest average
reward will be the better one.
 For the E-BR: with this metric we want to know which agent is the faster to reach the best
reward that can be achieved.</p>
      <p>Results are in the table below:
While the q network agent was able to reach the best reward faster than the SARSA agent, the later
one was able to achieve a better average reward and got to know the best action in each situation a bit
faster than the q network agent while at same time taking a significantly less period of time to
complete the training.</p>
    </sec>
    <sec id="sec-26">
      <title>Comparison between deep Reinforcement learning agents</title>
      <p>The comparison between the deep q network agent, the double deep q network agent and the deep
SARSA agent was based on 5 criteria: the average reward achieved (Avg-R), the time consumed
needed for training in minutes (TT), the episodes needed to get the best reward (E-BR), number of
episodes to reach a stable reward in each episode (SR), loss bound (LB).</p>
      <p>The two metrics added here are:
 SR: this metric show when no more improvements can be achieved because we already
reached the best possible outcome.</p>
      <p> LB: A value that shows the stability of our training process.
Several information became clear after drawing the table of comparison.</p>
      <p>First, results achieved by the 3 agents are all close to each other but the deep SARSA agent was able
to get past the other 2 clearly in 4 out of 5 of our defined metrics the Avg-R, TT, E-BR and TT at the
cost of a bigger loss interval.</p>
      <p>Second, the deep Q network was able to reach a stable reward values every round of training in a
small interval of numbers of episodes which means the training for this agent was more stable than
the other two.</p>
      <p>Lastly the DDQN agent was able to reach the maximum reward for the first time a lot faster than the
DQN agent and nearly in the same number of episodes as the SARSA agent, and it has the best loss
values of all three agents.</p>
    </sec>
    <sec id="sec-27">
      <title>6.3. Comparison between reinforcement learning and deep reinforcement learning agents</title>
      <p>In our problem the reinforcement learning agents specifically (q network and SARSA agents) were
able to achieve better results mid-training than the 3 deep reinforcement learning agents, due to the
fact that our environment is static and doesn’t hold that much number of states so a tabular agent can
solve this problem faster. But after training our deep learning models they were able to achieve
perfect scores each time in our predefined environments and they were able to adapt without any
problems to the changes we made later on in our roads.</p>
    </sec>
    <sec id="sec-28">
      <title>7. Final results</title>
      <p>From the comparisons above we can say that if the goal was just a simple task with limited states
there is no need to reach using deep learning models and advanced techniques. But because the final
goal is to use each agent in more complicated environments the deep reinforcement learning agents
should be a better option. Especially the proposed algorithm “rules based deep SARSA agent” and the
double deep q network agent. The last one had the best loss value of all agents but the deep SARSA
agent was better in all the rest aspects.</p>
    </sec>
    <sec id="sec-29">
      <title>8. Conclusion</title>
      <p>This paper demonstrates the power of the reinforcement learning and deep reinforcement learning
along-side simple rule based restriction in learning the best policy to navigate a road with a maze
structure. We used 5 different reinforcement and deep reinforcement learning algorithms such as the q
network, the SARSA algorithm, the deep q network and the double deep q network and lastly the
deep SARSA algorithm.</p>
      <p>All of our agents which used the previous mentioned algorithms were able to learn a policy that let
them achieve the max reward and navigate without any problem in our road. The deep reinforcement
learning algorithms also kept a great overall loss values in the middle of the training. Our proposed
algorithm “rules based deep SARSA algorithm” was able to achieve great results compared to the
more used techniques such as deep q network and double deep q network.</p>
      <p>In the future, this work can be extended in several ways such as adding more algorithms for
comparison and using a real simulator without changing the basics of these implementations to see
how effective these algorithms can be in controlling a real car in a highway situation. This is what is
planned for now alongside some other few ideas of combining other algorithms that are not related to
reinforcement learning with the previous mentioned ones.</p>
    </sec>
    <sec id="sec-30">
      <title>9. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>V.</given-names>
            <surname>Mnih</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kavukcuoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Antonoglou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wierstra</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Riedmiller: “Playing atari with deep reinforcement learning</article-title>
          .
          <source>” arXiv preprint arXiv:1312.5602</source>
          (
          <year>2013</year>
          ). arXiv:
          <volume>1312</volume>
          .
          <fpage>5602</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schrittwieser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Antonoglou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hubert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Baker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bolton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lillicrap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sifre</surname>
          </string-name>
          , G. Van den Driessche, T. Graepel,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Hassabis : “Mastering the game of Go without human knowledge</article-title>
          .
          <source>”Nature</source>
          <volume>550</volume>
          .7676, pp.
          <fpage>354</fpage>
          -
          <lpage>359</lpage>
          (
          <year>2017</year>
          ). https://doi.org/10.1038/nature24270
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] https://towardsdatascience.com/introduction-to
          <article-title>-reinforcement-learning-markov-decision-process44c533ebf8da, last-accessed</article-title>
          <volume>14</volume>
          /05/
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ravichandiran</surname>
          </string-name>
          <article-title>: “Hands-On Reinforcement Learning with Python Master reinforcement and deep reinforcement learning using OpenAI Gym</article-title>
          and TensorFlow”, pp.
          <fpage>91</fpage>
          -
          <lpage>111</lpage>
          . Packt Publishing (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5] https://pylessons.com/CartPole-reinforcement-learning/, last-accessed
          <volume>16</volume>
          /05/
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Homepage</surname>
          </string-name>
          ,https://mc.ai
          <article-title>/introduction-to-double-deep-q-learning-ddqn/</article-title>
          ,last-accessed
          <volume>25</volume>
          /05/
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Andrecut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.K.</given-names>
            <surname>Aliy</surname>
          </string-name>
          <article-title>: “Deep-SARSA: a reinforcement learning algorithm for autonomous navigation”</article-title>
          , World Scientific Publishing Company,
          <source>International Journal of Modern Physics C</source>
          , Vol.
          <volume>12</volume>
          ,No.
          <volume>10</volume>
          ,pp:
          <fpage>1513</fpage>
          -
          <lpage>1523</lpage>
          (
          <year>2001</year>
          ). https://www.worldscientific.com/doi/abs/10.1142/S0129183101002851
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>W.</given-names>
            <surname>Junjie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Qichao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dongbin</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Yaran : “Lane Change Decision-making through Deep Reinforcement Learning with Rule-based Constraints”</article-title>
          ,
          <source>International Joint Conference on Neural Networks</source>
          (
          <year>2019</year>
          ). arXiv:
          <year>1904</year>
          .00231
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Carl-Johan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Krister</surname>
          </string-name>
          , L. Leo : “Automated Speed and
          <article-title>Lane Change Decision Making using Deep Reinforcement Learning”</article-title>
          ,
          <source>IEEE International Conference on Intelligent Transportation Systems</source>
          , pp:
          <fpage>2148</fpage>
          -
          <lpage>2155</lpage>
          (
          <year>2018</year>
          ).
          <volume>10</volume>
          .1109/ITSC.
          <year>2018</year>
          .8569568
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Laura García</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Enrique</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Javier</given-names>
            <surname>Fernandez</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Nourdine , :
          <article-title>“Autonomous Driving in Roundabout Maneuvers using Reinforcement Learning with Q-Learning”</article-title>
          ,
          <source>Electronics journal</source>
          (
          <year>2019</year>
          ).
          <volume>10</volume>
          .3390/electronics8121536
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>