<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Deep Reinforcement Learning Approach to Adaptive Traffic Lights Management</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Vidali</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Crociani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Vizzari</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefania Bandini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CSAI - Complex Systems &amp; Artificial Intelligence Research Center, University of Milano-Bicocca</institution>
          ,
          <addr-line>Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>42</fpage>
      <lpage>50</lpage>
      <abstract>
        <p>-Traffic monitoring and control, as well as traffic simulation, are still significant and open challenges despite the significant researches that have been carried out, especially on artificial intelligence approaches to tackle these problems. This paper presents a Reinforcement Learning approach to traffic lights control, coupled with a microscopic agent-based simulator (Simulation of Urban MObility - SUMO) providing a synthetic but realistic environment in which the exploration of the outcome of potential regulation actions can be carried out. The paper presents the approach, within the current research landscape, then the specific experimental setting and achieved results are described.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Index Terms—reinforcement learning, traffic lights control,
traffic management, agent-based simulation</p>
    </sec>
    <sec id="sec-2">
      <title>I. INTRODUCTION</title>
      <p>Traffic monitoring and control and, in general, approaches
supporting the reduction of congestion still represent hot topics
for research of different disciplines, despite the substantial
researches that have been devoted to these topics. The global
phenomenon of urbanization (half of the world’s population
was living in cities at the end of 2008 and it is predicted
that by 2050 about 64% of the developing world and 86% of
the developed world will be urbanized1) is in fact constantly
changing the situation and making it actually harder to manage
such a concentration of population and transportation
demand. Technological developments among which autonomous
driving represents just the most futuristic one (at least from
a popular culture perspective), represent at the same time
attempts to tackle these issues and further challenges, in terms
of potential developments whose introduction requires further
study and analysis of the potential impact and implications.</p>
      <p>
        Artificial Intelligence plays an important role within this
framework; even not considering the obvious relevance to the
autonomous driving initiative, we focus here on two aspects:
(i) the regulation of traffic patterns, especially based on (ii) the
analysis of situations by means of agent-based simulations, in
which the behaviour of drivers and other relevant entities is
modeled and computer within a synthetic environment. The
latter, in particular, have reached a level of sufficient
complexity, flexibility, and they have proven their capability to support
decision makers in the exploration of alternative ways to
manage traffic within urban settings. On the side of regulation
of traffic patterns, the availability of these simulators, coupled
with advances in machine learning, represents an opportunity
for a scientific investigation of the possibility to employ
these virtual environments as tools to explore the outcome of
potential regulation actions within specific situations, within a
Reinforcement Learning [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] framework.
      </p>
      <p>
        This paper represents a contribution within this line of
research and, in particular, we focus on a simple yet still
studied situation: a single four way intersection regulated by
traffic lights, that we want to manage through an autonomous
agent perceiving the current traffic conditions, and exploiting
the experience carried out in simulated situations, possibly
representing plausibile traffic conditions. The simulations are
actually also agent-based, and in particular, for this study,
they have been carried out in a tool for Simulation of Urban
MObility (SUMO) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] providing a synthetic but realistic
environment in which the exploration of the outcome of potential
regulation actions can be carried out. An important aspect is
the fact that SUMO provides an Application Programming
Interface for interfacing with external programs, therefore we
were able to define a plausible set of observable aspects of
the environment, control the traffic lights according to the
decisions of the learning agent, as well as also to exploit some
stastics gathered by SUMO to describe the overall traffic flow
and therefore to define the reward to the actions carried out
by the traffic lights control agent.
      </p>
      <p>The paper breaks down as follows: we first provide a
compact description of the relevant portion of the state of
the art in traffic lights management with RL approaches, then
we introduce the experimental setting we adopted for this
study. The RL approach we defined and adopted will be given
in Section IV, then the achieved results will be described.
Conclusions and future developments will end the paper.</p>
    </sec>
    <sec id="sec-3">
      <title>II. RELATED WORKS</title>
      <sec id="sec-3-1">
        <title>A. Reinforcement Learning</title>
        <p>
          One of the acceptations of the goals of AI is to develop
machines that resemble the intelligent behavior of a human
being. In order to achieve this goal, an AI system should
be able to interact with the environment and learn how
to correctly act inside it. An established area of AI that
has been proved capable of experience-driven autonomous
learning is reinforcement learning [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Several complex tasks
were successfully completed using reinforcement learning in
multiple fields, such as games [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], robotics [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], and traffic
signal control.
In a Reinforcement Learning (RL) problem, an autonomous
agent observes the environment and perceives a state st,
which is the state of the environment at time t. Then the
agent chooses an action at which leads to a transition of the
environment to the state st+1. After the environment transition,
the agent obtains a reward rt+1 which tells the agent how good
at was with respect to a performance measure. The goal of the
agent is to learn the policy π∗ that maximizes the cumulative
expected reward obtained as a result of actions taken while
following π∗. The standard cycle of reinforcement learning is
shown in Figure 1.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>B. Learning in Traffic Signal Control</title>
        <p>
          Traffic signal control is a well suited application context for
RL techniques: in this framework, one or more autonomous
agents have the goal of maximizing the efficiency of traffic
flow that drives through one or more intersection controlled
by traffic lights. The use of RL for traffic signal control is
motivated by several reasons [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]: (i) if trained properly, RL
agents can adapt to different situations (e.g. road accidents,
bad weather conditions); (ii) RL agents can self-learn without
supervision or prior knowledge of the environment; (iii) the
agent only needs a simplified model of the environment
(essentially related to the state representation), since the agent
learns using the system performance metric (i.e. the reward).
        </p>
        <p>
          RL techniques applied to traffic signal control address the
following challenges: [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
• Inappropriate traffic light sequence. Traffic lights
usually choose the phases in a static, predefined policy. This
method could cause the activation of an inappropriate
traffic light phase in a situation that could cause an
increase in travel times.
• Inappropriate traffic light durations. Every traffic light
phase has a predefined duration which does not depend on
the current traffic conditions. This behavior could cause
unnecessary waitings for the green phase.
        </p>
        <p>Although the above are potential advantages of the RL
approach to traffic signal control, not all of them have already
been achieved, and (as we will show in the remander of the
paper) the present approach only represents an initial step in
this overall line of work.</p>
        <p>In order to apply a RL algorithm, it is necessary to define
the state representation, the available actions and the reward
functions; in the following, we will describe the most widely
adopted approaches for the design of these elements within
the context of Traffic Signal Control.</p>
        <p>1) State representation: The state is the agent’s perception
of the environment in an arbitrary step. In literature, state space
representations particularly differ in information density.</p>
        <p>
          In low information density representations, usually the
intersection’s lanes are discretized in cells along the length of the
lane. Lane cells are then mapped to cells of a vector, which
marks 1 if a vehicle is inside the lane cell, 0 otherwise [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
Some approaches include additional information, adopting
such a vector of car presence with the addition of a vector
encoding the relative velocity of vehicles [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The current
traffic light phase could also be added as a third vector [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>
          Regarding state representations with high information
density, usually the agent receives an image of the current situation
of the whole intersection, i.e. a snapshot of the simulator being
used; multiple successive snapshots will be stacked together
to give the agent a sense of the vehicle motion [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>2) Actions representation: In the context of traffic signal
control, the agent’s actions are implemented with different
degrees of flexibility and they are described below.</p>
        <p>
          Among the category of action set with low flexibility, the
agent can choose among a defined set of light combinations.
When an action is selected, a fixed amount of time will
lasts before the agent can select a new configuration [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
Some works gave the agent more flexibility by defining phase
duration with variable length [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. An agent with a higher
flexibility chooses an action at every step of the simulation
from a fixed set of light combinations. However, the selected
action is not activated if the minimum amount of time required
to release at least a vehicle, has not passed [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. A slightly
different approach would be to have a defined cycle of light
combinations activated into the intersection. The agent action
is represented by the choice of when it is time to switch to
the next light combination, and the decision is made at every
step [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>3) Reward representation: The reward is used by the agent
to understand the effects of the latest action taken in the latest
state; it is usually defined as a function of some performance
indicator of the intersection efficiently, such as vehicles’
delays, queue lengths, waiting times or overall throughput.</p>
        <p>
          Most of the works include the calculation of the change
between cumulative vehicle delay between actions, where the
vehicle delay is defined as the number of seconds the vehicles
is steady [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Similarly, the cumulative vehicle staying
time can be used, which is the number of seconds the vehicle
has been steady since his entrance in the environment [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
Moreover, some works combine multiples indicators in a
weighted sum [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>C. Adopted models and learning algorithms</title>
        <p>The most recent reinforcement learning research has
proposed multiple possible solutions to address the traffic signal
control problem, in which it emerges that different algorithms
and neural networks structure can be used, although some
common techniques are necessary but not sufficient in order
to ensure a good performance.</p>
        <p>
          The most widely used algorithm to address the problem is
Q-learning. The optimal behavior of the agent is achieved with
the passage of time is represented in simulation steps. But the
agent only operates at certain steps, after the environment has
evolved enough. Therefore, in this paper every step dedicated
to the agent’s workflow is called agentstep, while the steps
dedicated to the simulation are simply called ”steps”. Hence,
after a certain amount of simulation steps, the agent starts
its sequence of operations by gathering the current state
of the environment. Also, the agent calculates the reward
of the previous selected action, using some measure of the
current traffic situation. The sample of data containing every
information about the latest simulation steps is saved to a
memory and later extracted for a training session. Now the
agent is ready to select a new action based on the current
state of the environment, which will resume the simulation
until the next agent interaction.
the use of neural networks to approximate Q-values given a
state. Often, this approach includes a Convolutional Neural
Network (CNN) to compute the environment state and learn
features from an image [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] or a spatial representation [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>
          Genders and Ravi [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and Gao et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] make use of a
Convolutional Neural Network to learn features from their
spatial representation of the environment. The output of this
network with the current phase is passed to two fully
connected layers that connect to the outputs represented by
Qvalues. This method showed good results in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] work against
different traffic lights policies, such as long-queue-first and
fixed-times, while in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] it is compared to a shallow neural
network, in which (although it shows a good performance) an
evaluation against real-world traffic lights would lead to more
significant results.
        </p>
        <p>
          Mousavi et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] analyzed a double approach to address
the traffic signal control problem. The first approach is
valuebased, while the second is policy-based. In the first approach,
action values are predicted by minimizing the mean-squared
error of Q-values with the stochastic gradient-descent method.
In the alternative approach, the policy is learned by updating
the policy parameters in such a way that the probability of
taking good actions increases. A CNN is used as a
function approximator to extract features from the image of the
intersection, wherein the value-based approach the output is
the value of actions, and in the policy-based approach it is a
probability distribution over actions. Results show that both
the approaches achieve good performance against a defined
baseline and do not suffer from instability issues.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], a deep stacked autoencoders (SAE) neural network
is used to learn Q-values. This approach uses autoencoders
to minimize the error between the encoder neural network
Qvalue prediction and the target Q-value by using a specific
loss function. It is shown that achieves better performance
than traditional RL methods.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>III. EXPERIMENTAL SETTING</title>
      <p>
        The traffic microsimulator used for this research is
Simulation of Urban MObility (SUMO) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. SUMO provides a
software package which includes an infrastructure editor, a
simulator interface and an application programming interface
(API). These elements enable the user to design and implement
custom configurations and functionalities of a road
infrastructure and exchange data during the traffic simulation.
      </p>
      <p>In this research, the chance of improvement in traffic flow
that drives through an intersection controlled by traffic lights
will be investigated using artificial intelligence techniques. The
agent is represented by the traffic light system that interacts
with the environment in order to maximize a certain measure
of traffic efficiency. Given this general premise, the problem A. Training setup and traffic generation
tackled in this paper is defined as follows: given the state of The entire training is divided in multiple episodes .The total
the intersection, what is the traffic light phase that the agent number of episodes is 300. By default, SUMO provides a time
should choose, selected from a fixed set of predefined actions, frequency of 1 second per step, and the period of each episode
in order to maximize the reward and consequently optimize is set at 1 hour and 30 minutes, therefore the total number of
the traffic efficiency of the intersection. steps per episode is equal to 5400. 300 episodes of 1.30 hours</p>
      <p>The typical workflow of the agent is shown in Figure 2. each are equivalent to almost 19 days of continuous traffic, and
It should be underlined that in this application with SUMO, the entire training takes about 6 hours on a high-end laptop.</p>
      <p>The environment where the agent acts is represented in
Figure 3. It is a 4-way intersection where 4 lanes per arm
approach the intersection from the compass directions, leading
to 4 lanes per arm leaving the intersection. Each arm is 750
meters long. On every arm, each lane defines the possible
directions that a vehicle can follow: the right-most lane enable
vehicles to turn right or going straight, the two central lanes
bound the driver to go straight while on the left-most lane
the left turn is the only direction allowed. In the center of
the intersection, a traffic light system, controlled by the agent,
manages the approaching traffic. In particular, on every arm the
left-most lane has a dedicated traffic light, while the other three
lanes share a traffic light. Every traffic light in the environment
operates according to the common european regulations, with
the only exception being the absence of time between the end
of a yellow phase and the start of the next green phase. In this
environment pedestrians, sidewalks and pedestrian crossings
are not included.</p>
      <p>In a simulated intersection, the traffic generation is a crucial
part that can have a big impact on the agents performance. In
order to maintain a high degree of reality, in each episode the
traffic will be generated according to a Weibull distribution
with a shape equal to 2. An example is shown in Figure 4.</p>
      <p>The distribution is presented in the form of a histogram, where
the steps of one simulation episode are defined on the x-axis
and the number of vehicles generated in that step window is
defined on the y-axis. The Weibull distribution approximates
specific traffic situations, where during the early stage the
number of cars is rising, representing a peak hour. Then,
the number of incoming cars slowly decreases describing the
gradual mitigation of traffic congestion. Also, every vehicles
generated has the same physical dimensions and performance.</p>
      <p>
        The traffic distribution described provides the exact step
of the episode when a vehicle will be generated. For every
vehicle scheduled, its source arm and destination arm are
determined using a random number generator which have a
different seed in every episode, so it is not possible to have two
equivalent episodes. In order to obtain a true adaptive agent,
the simulation should include a significant variety of traffic
flows and patterns [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Therefore, four different scenarios are
defined and they are the following.
      </p>
      <p>• High-traffic scenario. 4000 cars approach the intersection
from every arm evenly distributed. Then, 75 % of
generated cars will go straight and 25 % of cars will turn left
or right at the intersection.
• Low-traffic scenario. 600 cars approach the intersection
from every arm evenly distributed. Then, 75 % of
generated cars will go straight and 25 % of cars will turn left
or right at the intersection.
• NS-traffic scenario. 2000 cars approach the intersection,
with 90 % of them coming from the North or South arm.</p>
      <p>Then, 75 % of generated cars will go straight and 25 %
of cars will turn left or right at the intersection.
• EW-traffic scenario. 2000 cars approach the intersection,
with 90 % of them coming from the East or West arm.</p>
      <p>Then, 75 % of generated cars will go straight and 25 %
of cars will turn left or right at the intersection.</p>
      <p>Each scenario corresponds to one single episode and they
cycle during the training always in the same order.</p>
      <p>IV. DESCRIPTION OF THE REINFORCEMENT LEARNING</p>
      <p>APPROACH</p>
      <p>In order to design a system based on the reinforcement
learning framework, it is necessary to define the state
representation, the action set, the reward function and the agent
learning techniques involved. It should be noted that the such
agent’s elements in this paper are easily replaceable with a
traffic monitoring system in a real world appliance, compared
to others relevant studies in this topic which have higher
requirements in terms of technical feasibility.</p>
      <sec id="sec-4-1">
        <title>A. State representation</title>
        <p>The state of the agent describes a representation of the
situation of the environment in a given agentstep t and it
is usually denoted with st. To allow the agent to effectively
learn to optimize the traffic, the state should provide sufficient
information about the distribution of cars on each road.</p>
        <p>
          The objective of the chosen representation is to let the
agent knows the position of vehicles inside the environment
at agentstep t. For this purpose the approach proposed in this
paper is inspired to the DTSE [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], with the difference that less
information is encoded in this state. In particular, this state
design includes only spatial information about the vehicles
hosted inside the environment, and the cells used to discretize
the continuous environment are not regular. The chosen design
for the state representation is focused on realism: recent works
on traffic signal controller proposed information-rich states,
but in reality they hard to implement since the information
used in that kind of representations is difficult to gather.
Therefore, in this paper will be investigated the chance of
obtaining good results with a simple and easy-to-apply state
representation.
        </p>
        <p>Technically, in each arm of the intersection incoming lanes
are discretized in cells that can identify the presence or absence
of a vehicle inside them. In Figure 5 is showed the state
representation for the west arm of the intersection. Between
the beginning of the road and the intersection’s stop line, there
are 20 cells. 10 of them are located along the left-only lane
while the others 10 cover the others three lanes. Therefore, in
the whole intersection there are 80 cells. Not every cell has the
same size: the further the cell is from the stop line, the longer
it is, so more lane length is covered. The choice of the length
of every cell is not trivial: if cells were too long, some cars
approaching the crossing line may not be detected; if cells
were too short, the number of states required to cover the
length of the lane increases, bringing to higher computational
complexity. In this paper, the length of the shortest cells, which
are also the closest to the stop line, is exactly 2 meters longer
than the length of a car.</p>
        <p>In summary, whenever the agent observe the state of the
environment, he will obtain the set of cells that describe the
presence or absence of vehicles in every area of the incoming
roads.</p>
      </sec>
      <sec id="sec-4-2">
        <title>B. Action set</title>
        <p>The action set identifies the possible actions that the agent
can take. The agent is the traffic light system, so doing an
action translates to activate a green phase for a set of lanes
for a fixed amount of time, choosing from a predefined set of
green phases. In this paper, the green time is set at 10 seconds
and the yellow time is set at 4 seconds. Formally, the action
space is defined in the set (1). The set includes every possible
action that the agent can take.</p>
        <p>A = {NSA, NSLA, EWA, EWLA}
(1)
Every action of set (1) is described below.
• North-South Advance (NSA): the green phase is active
for vehicles that are in the north and south arm and wants
to proceed straight or turn right.
• North-South Left Advance (NSLA): the green phase is
active for vehicles that are in the north and south arm
and wants to turn left.
• East-West Advance (EWA): the green phase is active for
vehicles that are in the east and west arm and wants to
proceed straight or turn right.
• East-West Left Advance (EWLA): the green phase is
active for vehicles that are in the east and west arm and
wants to turn left.</p>
        <p>Figure 6 shows a graphical representation of the four
possible actions.</p>
        <p>If the action chosen in agentstep t is the same as the
action taken in the last agentstep t − 1 (i.e. the traffic light
combination is the same), there is no yellow phase and
therefore the current green phase persists. On the contrary,
if the action chosen in agentstep t is not equal to the previous
action, a 4 seconds yellow phase is initiated between the
two actions. This means that the number of simulation steps
between two same actions is 10, since 1 simulation step is
equal to 1 second in SUMO. When the two consecutive actions
are different, the yellow phase counts as 4 extra simulation
steps and therefore the total number of simulation steps in
between actions is 14. Figure 7 shows a brief scheme of this
process.</p>
        <p>
          In reinforcement learning, the reward represents the
feedback from the environment after the agent has chosen an
action. The agent uses the reward to understand the result
of the taken action and improve the model for future action
choices. Therefore, the reward is a crucial aspect of the
learning process. The reward usually has two possible values:
positive or negative. A positive reward is generated as a
consequence of good actions, a negative reward is generated
from bad actions. In this application, the objective is to
maximize the traffic flow through the intersection over time. In
order to achieve this goal, the reward should be derived from
some performance measure of traffic efficiency, so the agent
is able to understand if the taken action reduce or increase the
intersection efficiency. In traffic analysis, several measures are
used [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], such as throughput, mean delay and travel time. In
this paper, two reward functions are presented which use two
slightly different traffic measures, and they are the following.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>1) Literature reward function: The first reward function is</title>
        <p>called literature because it is inspired to similar studies in this
topic. The literature reward function uses as a metric the total
waiting time, defined as in equation (2).</p>
        <p>n
X
veh=1
twtt =
wt(veh,t)
(2)
Where wt(veh,t) is the amount of time in seconds a vehicle veh
has a speed of less than 0.1 m/s at agentstep t. n represents
the total number of vehicles in the environment in agentstep t.
Therefore, twtt is the total waiting time at agentstep t. From
this metric, the literature reward function can be defined as a
function of twtt and is shown in (3)
rt = 0.9 · twtt−1 − twtt
(3)
Where rt represents the reward at agentstep t. twtt and
twtt−1 represent the total waiting time of all the cars in the
intersection captured respectively at agentstep t and t − 1. The
parameter 0.9 helps with the stability of the training process.</p>
        <p>In a reinforcement learning application, the reward usually
can be positive or negative, and this implementation is no
exception. The equation 3 is designed in such a way that
when the agent chooses a bad action it returns a negative
value and when it chooses a good action it returns a positive
value. A bad action can be represented as an action that, in the
current agentstep t, adds more vehicles in queues compared
to the situation in the previous agentstep t − 1, resulting in
higher waiting times compared to the previous agentstep. This
behavior increases the twt for the current agentstep t and
consequently the equation 3 assumes a negative value. The
more vehicles were added in queues for the agentstep t, the
more negative rt will be and therefore the worst the action
will be evaluated by the agent. The same concept is applied
for good actions.</p>
        <p>The problem with this reward function lays inside the
choiche of the metric, and happens when the following
situation arise. During the High-traffic scenario, very long queues
appears. When the agent activate the green phase for a long
queue, the departure of cars creates a wave of movement
that traverse the entire queue. The reward associated to this
phase activation is received not only in the next agentstep,
as it should, but also in very next ones. That is because the
movement wave persists longer compared to the delta step
between actionstep, and the wave triggers the waiting times
of cars in the queue, misleading the agent about the reward
received.</p>
      </sec>
      <sec id="sec-4-4">
        <title>2) Alternative reward function: The alternative reward</title>
        <p>function uses a metric that is slightly different from the former
metric, which is the accumulated total waiting time, defined
in equation (4).</p>
        <p>atwtt =
n
X awt(veh,t)
veh=1
(4)
Where awt(veh,t) is the amount of time in seconds a vehicle
veh has a speed of less than 0.1 m/s at agentstep t, since the
spawn into the environment. n represents the total number of
vehicles in the environment in agentstep t. Therefore, atwtt
is the accumulated total waiting time at agentstep t. With
this metric, when the vehicle departs but it does not manage
to cross the intersection, the value of atwtt does not resets
(unlike the value of twtt), avoiding the misleading reward
associated with the literature reward function, when a long
queue build up at the intersection. Once the metric is set, the
alternative reward function is defined such as in equation (5)
the cars in the intersection captured respectively at agentstep
t and t − 1.</p>
      </sec>
      <sec id="sec-4-5">
        <title>D. Deep Q-Learning</title>
        <p>
          The learning mechanism involved in this paper is called
Deep Q-Learning, which is a combination of two aspects
widely adopted in the field of reinforcement learning: deep
neural networks and Q-Learning. Q-Learning [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] is a form of
model-free reinforcement learning [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. It consists of assigning
a value, called the Q-value, to an action taken from a precise
state of the environment. Formally, in literature, a Q-value is
defined as in equation (6).
        </p>
        <p>Q(st, at) = Q(st, at)+α(rt+1+γ·maxAQ(st+1, at)−Q(st, at))
(6)
where Q(st, at) is the value of the action at taken from state
st. The equation consists on updating the current Q-value
with a quantity discounted by the learning rate α. Inside the
parenthesis, the term rt+1 represents the reward associated to
taking action at from state st. The subscript t + 1 is used to
emphasize the temporal relationship between taking the action
at and receiving the consequent reward. The term Q(st+1, at)
represents the immediate future’s Q-value, where st+1 is next
state in which the environment has evolved after taking action
at in state st. The expression maxA means that, among the
possible actions at in state st+1, the most valuable is selected.
The term γ is the discount factor that assumes a value between
0 and 1, lowering the importance of future reward compared
to the immediate reward.</p>
        <p>In this paper, a slightly different version of the equation (6)
is used and it is presented in equation (7). This will be called
the Q-learning function from this point.</p>
        <p>′
Q(st, at) = rt+1 + γ · maxAQ (st+1, at+1)
(7)
Where the reward rt+1 is the reward received after taking
action at in state st. The term Q′(st+1, at+1) is the Q-value
associated with taking action at+1 in state st+1, i.e. the next
state after taking action at in state st. As seen in equation (6),
the discount factor γ denote a small penalization of the future
reward compared to the immediate reward. Once the agent
is trained, the best action at taken from state st will be the
one that maximize the function Q(st, at). In other words,
maximizing the Q-learning function means following the best
strategy that the agent have learned.</p>
        <p>
          In a reinforcement learning application, often the state space
is so large that is impractical to discover and save every
stateaction pair. Therefore, the Q-learning function is approximated
using a neural network. In this paper, a fully connected deep
neural network is used, which is composed of an input layer of
80 neurons, 5 hidden layers of 400 neurons each with rectified
linear unit (ReLU) [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] and the output layer with 4 neurons
with linear activation function, each one representing the value
of an action given a state. A graphical representation of the
deep neural network is showed in Figure 8
rt = atwtt−1 − atwtt
(5)
        </p>
      </sec>
      <sec id="sec-4-6">
        <title>E. The training process</title>
        <p>
          Where rt represents the reward at agentstep t. atwtt and Experience replay [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] is a technique adopted during the
atwtt−1 represent the accumulated total waiting time of all training phase in order to improve the performance of the
agent and the learning efficiency. It consists of submitting
to the agent the information needed for learning in the form
of a randomized group of samples called batch, instead of
immediately submitting the information that the agent gather
during the simulation (commonly called Online Learning). The
batch is taken from a data structure intuitively called memory,
which stores every sample collected during the training phase.
        </p>
        <p>A sample m is formally defined as the quadruple (8).</p>
        <p>m = {st, at, rt+1, st+1}</p>
        <p>(8)
Where rt+1 is the reward received after taking the action at
from state st, which evolves the environment into the next state
st+1. This technique is implemented to remove correlations in
the observation sequence, since the state of the environment
st+1 is a direct evolution of the state st and the correlation
can decrease the training capability of the agent. In Figure 9
is shown a representation of the data collection task.</p>
        <p>Fig. 9. Scheme of the data collection.</p>
        <p>As stated earlier, the experience replay technique needs
a memory, which is characterized by a memory size and a
batch size. The memory size represents how many samples
the memory can store and is set at 50000 samples. The batch
size is defined as the number of samples that are retrieved
from the memory in one training instance and is set at 100. If
at a certain agentstep the memory is filled, the oldest sample
is removed to make space for the new sample.</p>
        <p>A training istance consists of learning the Q-value function
iteratively using the information contained in the batch of
samples extracted. Every sample in the batch is used for
training. From the standpoint of a single sample, which contains
the elements {st, at, rt+1, st+1}, the following operations are
executed:
1) Prediction of the Q-values Q(st), which is the current
knowledge that the agent has about the action values from
st.
2) Prediction of the Q-values Q′(st+1). These represents the
knowledge of the agent about the action values starting
from the state st+1.
3) Update of Q(st, at) which represents the value of the
particular action at selected by the agent during the
simulation. This value is overwritten using the Q-learning
function described in equation (7). The element rt+1 is the
irsewoabrtdainasesdocuisaitnegd tthoethpereadcitcitoionnato,f
mQa′(xsAt+Q1′)(satn+d1,raetp+r1e)sents the maximum expected future reward i.e. the higher
action value expected by the agent, starting from state
st+1. It will be discounted by a factor γ that gives more
importance to the immediate reward.
4) Training of the neural network. The input is the state st,
while the desired output is the updated Q-values Q(st, at)
that now includes the maximum expected future reward
thanks to the Q-value update.</p>
        <p>Once the deep neural network has sufficiently approximated
the Q-learning function, the best traffic efficiency is achieved
by selecting the action with the highest value given the
current state. A major problem in any reinforcement learning
task is the action-selection policy while learning; whether
to take exploratory action and potentially learn more, or to
take exploitative action and attempt to optimize the current
knowledge about the environment evolution. In this paper the
ǫ-greedy exploration policy is chosen, and it is represented
by the equation (9). It defines a probability ǫ for the current
episode h to choose an explorative action, and consequently a
probability 1 − ǫ to choose an exploitative action.</p>
        <p>(9)
ǫh = 1 −
h</p>
        <p>H
where h is the current episode of training and E is the
total number of episodes. Initially, ǫ = 1, meaning that the
agent exclusively explores. However, as training progresses,
the agent increasingly exploits what it has learned, until it
exclusively exploits.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>V. SIMULATION RESULTS</title>
      <p>The performance of the agents is assessed in two parts:
initially, the reward trend during the training is analyzed. Then,
a comparison between the agents and a static traffic light is
discussed, with respect to common traffic metrics, such as
cumulative wait time and average wait time per vehicle.</p>
      <p>One agent is trained using the literature reward function,
while the other one adopts the alternative reward function.
Figure 10 shows the learning improvement during the training
in the Low-traffic scenario of both agents, in term of
cumulative negative reward i.e the magnitude of actions’ negative
outcomes during each episode. As it can be seen, each agent
has learned a sufficiently correct policy in the Low-traffic
scenario. As the training proceeded, both agents efficiently
explore the environment and learn an adequate approximation
of the Q-values; then, towards the end of the training, they
try to optimize the Q-values by exploiting the knowledge
learnined so far. The fact that the agent with the alternative
reward function has a better reward curve overall is not a
strong evidence of a better performance, since haveing two
different reward functions means that different reward values
are produced. The performance difference will be discussed
later during the static traffic light benchmark.</p>
      <p>Figure 11 shows the same training data as Figure 10, but
referred to the High-traffic scenario. In this scenario, the agent
with the literature reward shows a significantly unstable reward
curve, while the other agent’s trend is stable. This behavior is
caused by the choiche of using the waiting time of vehicles as a
metric for the reward function, which in situations with long
queues causes the aquisition of misleading rewards. In fact,
by using the accumulated waiting time like in the alternative
reward function, vehicles does not resets their waiting times
by simply advancing through the queue. As Figure 11 shows,
the alternative reward function produces a more stable policy.
In the NS-traffic and EW-traffic scenarios, both agents perform
well since it is a simpler task to exploit.</p>
      <p>
        In order to truly analyze which agent achieve better
performance, a comparison between the agents and a Static Traffic
Light (STL) is presented. The STL has the same layout of
the agents and it cycle through the 4 phases always in the
following order: [NSA − NSLA − EWA − EWLA]. Moreover,
every phase has a fixed duration and they are inspired by those
on real-world static traffic lights [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. In particular, the phases
NSA and EWA lasts 30 seconds, the phases NSLA and EWLA
lasts 15 seconds and the yellow phase is the same as the agent,
which is 4 seconds.
      </p>
      <p>In Table I are shown the performance of the two agents,
compared to the STL. The metric used to measure the
performance difference are the cumulative wait time and the
average wait time per vehicle. The cumulative wait time is
defined as the sum of all waiting times of every car during
the episode, while the average waiting time per vehicle is
defined as the average amount of seconds spent by a vehicle
in a steady position during the episode. These measures are
gathered across 5 episodes and then averaged.</p>
      <p>cwt
awt/v
cwt
awt/v
cwt
awt/v
cwt</p>
      <p>Literature reward Alternative reward
agent agent</p>
      <p>Low-traffic scenario
-30
-29
+145
+136
-50
-47
-65</p>
      <p>High-traffic scenario
NS-traffic scenario
EW-traffic scenario
-47
-45
+26
+25
-62
-56
-65</p>
      <p>In general, the alternative reward agent achieves a better
traffic efficiency compared to the literature agent: this is a
consequence of the adoption of a reward function
(accumulated waiting time) that more properly discounts waiting times
exceeding a single traffic light cycle. Considering just the
waiting time starting from the last stop of the vehicles, leads
to not sufficiently emphasize the usefulness of keeping longer
light cycles, introducing too many yellow lights situations
and changes, that are effective in low or medium traffic
situations. The fact that the agent is more effectine in low
to medium traffic situations, leads to think that an easy and
almost immediate opportunity would be to separately develop
agents devoted to different traffic situations, having a sort of
controller that monitors the traffic flow and that selects the
most appropriate agent configuration. This experimentation
also leads to consider that, however, additional improvements
would be possible by (i) improving the learning approach
to achieve a more stable and faster convergence, (ii) further
improving the reward fuction to better describe the desired
behaviour and to influence the average cycle lengths, that is
more fruitfully short in low traffic situation and long whenever
the traffic condition worsens.</p>
    </sec>
    <sec id="sec-6">
      <title>VI. CONCLUSIONS AND FUTURE DEVELOPMENTS</title>
      <p>This work has presented a believable exploration of the
plausibility of a RL approach to the problem of traffic lights
adaptation and management. The work has employed a
realistic and validated traffic simulator to provide an environment
in which training and evaluating a RL agent. Two metrics for
the reward of agent’ actions have been investigated, clarifying
that a proper decription of the application context is just
as important as the competence in the proper application of
machine learning approaches for achieving proper results.</p>
      <p>Future works are aimed at further improving achieved
results, but also, within a longer term, at investigating what
would be the implications of introducing mutiple RL agents
within a road network and what would be the possiblity to
coordinate their efforts for achieving global improvements
over local ones, and also the implications on the vehicle
population, that could perceive the change in the infrastructure
and adapt in turn to exploit additional opportunities and
potentially negating the achieved improvements due to an additional
traffic demand on the improved intersections. It is important
to perform analyses along this line of work to understand the
plausibility, potential advantages or even unintended negative
implications of the introduction in the real world of this form
ofself-adaptive system.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Barto</surname>
          </string-name>
          et al.,
          <article-title>Introduction to reinforcement learning</article-title>
          . MIT press Cambridge,
          <year>1998</year>
          , vol.
          <volume>135</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Behrisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bieker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Erdmann</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Krajzewicz</surname>
          </string-name>
          , “
          <article-title>Sumo - simulation of urban mobility: An overview</article-title>
          ,”
          <source>in SIMUL</source>
          <year>2011</year>
          , S. . U. of Oslo Aida Omerovic,
          <string-name>
            <given-names>R. I. R. T. P. D. A.</given-names>
            <surname>Simoni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. I. R. T. P. G.</given-names>
            <surname>Bobashev</surname>
          </string-name>
          , Eds. ThinkMind,
          <year>October 2011</year>
          . [Online]. Available: https://elib.dlr.de/71460/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Babuschkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mathieu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jaderberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. M.</given-names>
            <surname>Czarnecki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dudzik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Georgiev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Powell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ewalds</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Horgan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kroiss</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Danihelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Agapiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dalibard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sifre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sulsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vezhnevets</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Molloy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Budden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Paine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gulcehre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Pfaff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Pohlen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yogatama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>McKinney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schaul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lillicrap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Apps</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kavukcuoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hassabis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          , “
          <article-title>AlphaStar: Mastering the Real-Time Strategy Game StarCraft II</article-title>
          ,” https://deepmind.com/blog/ alphastar
          <article-title>-mastering-real-time-strategy-game-starcraft-ii/</article-title>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Irpan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pastor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ibarz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Herzog</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Jang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Quillen</surname>
          </string-name>
          , E. Holly,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kalakrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vanhoucke</surname>
          </string-name>
          et al., “
          <article-title>Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation</article-title>
          ,” arXiv preprint arXiv:
          <year>1806</year>
          .10293,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>K.-L. A. Yau</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Qadir</surname>
            ,
            <given-names>H. L.</given-names>
          </string-name>
          <string-name>
            <surname>Khoo</surname>
            ,
            <given-names>M. H.</given-names>
          </string-name>
          <string-name>
            <surname>Ling</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Komisarczuk</surname>
          </string-name>
          , “
          <article-title>A survey on reinforcement learning models and algorithms for traffic signal control</article-title>
          ,
          <source>” ACM Computing Surveys (CSUR)</source>
          , vol.
          <volume>50</volume>
          , no.
          <issue>3</issue>
          , p.
          <fpage>34</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W.</given-names>
            <surname>Genders</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Razavi</surname>
          </string-name>
          , “
          <article-title>Evaluating reinforcement learning state representations for adaptive traffic signal control,” Procedia computer science</article-title>
          , vol.
          <volume>130</volume>
          , pp.
          <fpage>26</fpage>
          -
          <lpage>33</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ito</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Shiratori</surname>
          </string-name>
          , “
          <article-title>Adaptive traffic signal control: Deep reinforcement learning algorithm with experience replay and target network</article-title>
          ,
          <source>” arXiv preprint arXiv:1705.02755</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>W.</given-names>
            <surname>Genders</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Razavi</surname>
          </string-name>
          , “
          <article-title>Using a deep reinforcement learning agent for traffic signal control</article-title>
          ,
          <source>” arXiv preprint arXiv:1611.01142</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Mousavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schukat</surname>
          </string-name>
          , and E. Howley, “
          <article-title>Traffic light control using deep policy-gradient and value-function-based reinforcement learning</article-title>
          ,
          <source>” IET Intelligent Transport Systems</source>
          , vol.
          <volume>11</volume>
          , no.
          <issue>7</issue>
          , pp.
          <fpage>417</fpage>
          -
          <lpage>423</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lv</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.-Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          , “
          <article-title>Traffic signal timing via deep reinforcement learning</article-title>
          ,
          <source>” IEEE/CAA Journal of Automatica Sinica</source>
          , vol.
          <volume>3</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>247</fpage>
          -
          <lpage>254</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wei</surname>
          </string-name>
          , G. Zheng,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yao</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          , “
          <article-title>Intellilight: A reinforcement learning approach for intelligent traffic light control</article-title>
          ,”
          <source>in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining. ACM</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>2496</fpage>
          -
          <lpage>2505</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D.</given-names>
            <surname>Krajzewicz</surname>
          </string-name>
          , G. Hertkorn,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Ro¨ssel, and</article-title>
          <string-name>
            <given-names>P.</given-names>
            <surname>Wagner</surname>
          </string-name>
          , “
          <article-title>Sumo (simulation of urban mobility)-an open-source traffic simulation</article-title>
          ,”
          <source>in Proceedings of the 4th middle East Symposium on Simulation and Modelling (MESM20002)</source>
          ,
          <year>2002</year>
          , pp.
          <fpage>183</fpage>
          -
          <lpage>187</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Rodegerdts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nevers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ringert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Koonce</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>McGill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Stewart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Suggett</surname>
          </string-name>
          et al., “Signalized intersections: informational guide,
          <source>” Tech. Rep.</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Dowling</surname>
          </string-name>
          , “
          <article-title>Traffic analysis toolbox volume vi: Definition, interpretation, and calculation of traffic analysis tools measures of effectiveness,”</article-title>
          <string-name>
            <surname>Tech. Rep.</surname>
          </string-name>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Watkins</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Dayan</surname>
          </string-name>
          , “
          <article-title>Q-learning,” Machine learning</article-title>
          , vol.
          <volume>8</volume>
          , no.
          <issue>3-4</issue>
          , pp.
          <fpage>279</fpage>
          -
          <lpage>292</lpage>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>C. J. C. H. Watkins</surname>
          </string-name>
          , “
          <article-title>Learning from delayed rewards</article-title>
          ,
          <source>” Ph.D. dissertation, King's College</source>
          , Cambridge,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J. N.</given-names>
            <surname>Tsitsiklis</surname>
          </string-name>
          and
          <string-name>
            <surname>B. Van Roy</surname>
          </string-name>
          ,
          <article-title>“Analysis of temporal- diffference learning with function approximation</article-title>
          ,”
          <source>in Advances in neural information processing systems</source>
          ,
          <year>1997</year>
          , pp.
          <fpage>1075</fpage>
          -
          <lpage>1081</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>L.-J. LIN</surname>
          </string-name>
          , “
          <article-title>Reinforcement learning for robots using neural networks</article-title>
          ,
          <source>” Ph.D. thesis</source>
          , Carnegie Mellon University,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P.</given-names>
            <surname>Koonce</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Rodegerdts</surname>
          </string-name>
          , “
          <article-title>Traffic signal timing manual</article-title>
          .
          <source>” United States. Federal Highway Administration, Tech. Rep.</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>