<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>GESA: A GEneral Scenario-Agnostic Reinforcement Learning for Trafic Signal Control ⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Haoyuan Jiang</string-name>
          <email>jianghaoyuan@zju.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ziyue Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhishuai Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lei Bai</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hangyu Mao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wolfgang Ketter</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rui Zhao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>STRL'24: Third International Workshop on Spatio-Temporal Reasoning and Learning</institution>
          ,
          <addr-line>5</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Sensetime Research</institution>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Shanghai AI Lab</institution>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Cologne</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Reinforcement learning (RL) can automatically learn a better policy through a trial-and-error paradigm and has been adopted to revolutionize and optimize traditional trafic signal control systems that are usually based on handcrafted methods. However, most existing RL-based models are either based on a single scenario or multiple independent scenarios, where each scenario has a separate simulation environment with predefined road network topology and trafic signal settings. These models implement training and testing in the same scenario, thus being strictly tied up with the specific setting and sacrificing model generalization heavily. While a few recent models could be trained by multiple scenarios, they require a huge amount of manual labor to label the intersection structure, hindering the model's generalization. In this work, we aim at a general framework that could eliminate heavy labeling and model a variety of scenarios simultaneously. To this end, we propose a general Scenario-Agnostic (GESA) reinforcement learning framework for trafic signal control with: (1) A general plug-in module to map all diferent intersections into a unified structure, freeing us from the heavy manual labor to specify the structure of intersections; (2) A unified state and action space design to keep the model input and output consistently structured; (3) A large-scale co-training with multiple scenarios, leading to a generic trafic signal control algorithm. GESA can automatically handle various structured intersections from various cities without human labeling, and it co-trains a generalist agent to control trafic signals for multiple cities together, which also demonstrates superior transferability in zero-shot settings. In experiments, we demonstrate our algorithm as the first one that can be co-trained with seven diferent scenarios without manual annotation and gets 13.27% higher rewards than baselines. When dealing with a new scenario, our model can still achieve 9.39% higher rewards. The code, scenarios, and demos are available here. The full paper is available at [1].</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Trafic signal control</kwd>
        <kwd>Reinforcement learning</kwd>
        <kwd>A generalist agent</kwd>
        <kwd>Zero-shot transfer</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Reinforcement learning (RL) [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
        ] has been preferably
adopted into the TSC domain since it is a learning-based
method with higher automation. Such a trial-and-error
paradigm based on the trafic simulator has demonstrated
better performance than transport engineering-based
methods [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The recent RL-based TSC models can be roughly
divided into two categories based on the scenarios where
the training and testing are conducted. A scenario is usually
a simulation environment that contains a set of
intersections: (1) Single-scenario RL, as the majority, its training
and testing need to be on the same scenario [
        <xref ref-type="bibr" rid="ref2 ref6 ref7">2, 6, 7</xref>
        ].
However, the model will be unusable or perform badly in a new
scenario. For example, in Fig. 1(b) top, these methods might
be trained and tested in the same scenario with 5 × 5
fourapproach intersections of Fig. 1(a2), but these methods will
ill-perform in another new scenario with mixed
intersections of Fig. 1(a1) and 1(a2). (2) Multi-scenario RL, as
shown in Fig. 1(b) bottom, where training is conducted
in multiple scenarios, and testing could be in diferent
scenarios. For example, [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8, 9, 10</xref>
        ] are proposed to train a TSC
system with multi-scenarios. However, in the training stage,
the existing multi-scenario RL models need heavy manual
labor to annotate the structure of intersections, such as the
direction of each entering approach, the number of entering
lanes of each entering approach, the trafic movement of
each entering lane, etc. Moreover, they either achieve
multiscenario co-training in a sequential manner, one scenario
after another, leading to a rather unstable learning curve
and slower convergence [
        <xref ref-type="bibr" rid="ref10 ref8">8, 10</xref>
        ] or only narrow the scale of
a scenario to only one intersection in one scenario, which
heavily limits the model’s generalization [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>Moreover, current RL-based TSC methods are trained
with several pre-defined and fixed scenarios, whereas they
cannot gain generalization capability without labeling,
which limits the application of RL-based methods in the
real world. These methods can exploit the various trafic
lfows generated by the simulator to make the model
efective in training scenarios, but finding a low-cost universal
method with promising transferability meanwhile is still
a research gap. As a result, the existing methods still face
tremendous challenges in jumping out from the simulation
and implementing them in real cities. This is known as
sim2real challenge. The challenges mainly come from the
wide gap between the real complex cities and the simplified
simulation systems. In the real world, the intersection
structure could be rather versatile in terms of diferent settings
of approaches (i.e., north, south, east, west), movements (i.e.,
left, right, through), and lanes (e.g., two through lanes, one
right-through lane). As shown in Fig. 1(a), an intersection
could have a diferent amount of approaches. Within an
approach, there can be diferent combinations of movements;
A lane could also combine diferent movements. However,
most of the existing methods only consider a standard
simulation intersection with four approaches and three lanes
(right, through, and left) within each approach. This largely
limits the model generalization.</p>
      <p>To conclude, a qualified TSC approach needs high
generalization and efectiveness: it should handle various
intersections and be able to transfer to other unseen targets
easily and with low cost. In this paper, we aim to answer
three questions: (1) How do we co-train an RL with
multiscenarios without labeling, given the diverse intersection
structures? (2) Will multi-scenario co-training improve the
TSC? If yes, why? And the more scenarios, the better? (3)</p>
    </sec>
    <sec id="sec-2">
      <title>North</title>
      <p>lean ltaen
t h
feL igR</p>
    </sec>
    <sec id="sec-3">
      <title>East</title>
      <p>Right lane
Through lane
Through lane
Through lane
Right-Through-Left</p>
    </sec>
    <sec id="sec-4">
      <title>East</title>
      <p>!!</p>
      <p>Left lane
Through lane
Through lane
Through lane
South
(a1) Intersection of three approaches, (a2) Intersection of four approaches,
with 2-4 entering lanes, and 2 with 1 entering lane of
right-throughmovements on each approach left movement on each approach

!
ht-Tohuroguhgh
RigThrLeft
✅</p>
      <sec id="sec-4-1">
        <title>Adapt</title>
        <p>RL
Model ✅</p>
      </sec>
      <sec id="sec-4-2">
        <title>Adapt ✅</title>
        <p>…
(b) Single-scenario
Training v.s.
Multiscenario Training</p>
        <p>Training scenario
Adapting scenario</p>
        <p>Does the co-trained RL model still perform well in the new
scenario?</p>
        <p>
          To narrow the sim2real gap significantly and get more
ready to be deployed in real cities, in this paper, we provide a
GEneral Scenario-Agnostic (GESA) reinforcement learning
framework for the TSC task. To our best knowledge, GESA
is the first work that pursues high generability and co-trains
multiple scenarios without labels: it automatically handles
various scenarios; the reinforcement learning is designed
accordingly to achieve generalization; it is co-trained with
multiple scenarios simultaneously and demonstrates high
transferability. Specifically, to co-train in multiple
scenarios with various intersections, the vectors with approach
spatial information are employed to map shape-odd and
complex intersections into the standard intersection. Then,
the mapped intersections are used to generate the
characteristic information of each trafic movement and the phase
of the trafic lights in a specific order. Finally, we extend
the original FRAP [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] to a policy gradient-based framework,
which can facilitate the model coverage and is compatible
with diferent intersections.
        </p>
        <p>The contributions are summarized in three-fold: (1) We
present a general plug-in module to map the intersections
into a unified structure, freeing us from the heavy
manual labeling work to specify the intersection structure and
enabling large-scale co-training under multiple diferent
scenarios. (2) Accordingly, we design a unified state and action
space to keep the model input and output structure
consistent for more general capabilities. Moreover, the GESA can
adapt to various unseen scenarios and achieve promising
performance without re-training. (3) We build two
realworld scenarios using the real city road map and the real
trafic dynamics, together with five public scenarios, where
we co-train and validate the GESA with prudent
experiments. All these lead us closer to the ultimate goal: to
implement RL-based TSC in real cities.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ketter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>A general scenario-agnostic reinforcement learning for trafic signal control</article-title>
          ,
          <source>IEEE Transactions on Intelligent Transportation Systems</source>
          (
          <year>2024</year>
          )
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Van der Pol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Oliehoek</surname>
          </string-name>
          ,
          <article-title>Coordinated deep reinforcement learners for trafic light control</article-title>
          ,
          <source>Proceedings of Learning, Inference and Control of Multi-agent Systems (at NIPS</source>
          <year>2016</year>
          )
          <volume>1</volume>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wei</surname>
          </string-name>
          , G. Zheng,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Intellilight: A reinforcement learning approach for intelligent trafic light control</article-title>
          ,
          <source>in: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>2496</fpage>
          -
          <lpage>2505</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Zheng,
          <string-name>
            <given-names>K.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gayah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Presslight: Learning max pressure control to coordinate trafic signals in arterial network</article-title>
          ,
          <source>in: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1290</fpage>
          -
          <lpage>1298</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>K.-L. A. Yau</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Qadir</surname>
            ,
            <given-names>H. L.</given-names>
          </string-name>
          <string-name>
            <surname>Khoo</surname>
            ,
            <given-names>M. H.</given-names>
          </string-name>
          <string-name>
            <surname>Ling</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Komisarczuk</surname>
          </string-name>
          ,
          <article-title>A survey on reinforcement learning models and algorithms for trafic signal control</article-title>
          ,
          <source>ACM Computing Surveys (CSUR) 50</source>
          (
          <year>2017</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Learning phase competition for trafic signal control</article-title>
          ,
          <source>in: Proceedings of the 28th ACM International Conference on Information and Knowledge Management</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1963</fpage>
          -
          <lpage>1972</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Tian,</surname>
          </string-name>
          <article-title>MTLight: Eficient multi-task reinforcement learning for trafic signal control</article-title>
          ,
          <source>in: ICLR 2022 Workshop on Gamification and Multiagent Solutions</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yao</surname>
          </string-name>
          , G. Zheng,
          <string-name>
            <given-names>N.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Metalight: Value-based meta-reinforcement learning for trafic signal control</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>34</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>1153</fpage>
          -
          <lpage>1160</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Oroojlooy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nazari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hajinezhad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Silva</surname>
          </string-name>
          , Attendlight:
          <article-title>Universal attention-based reinforcement learning model for trafic signal control</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>4079</fpage>
          -
          <lpage>4090</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>O. Pun</surname>
          </string-name>
          ,
          <article-title>Adlight: A universal approach of trafic signal control with augmented data using reinforcement learning</article-title>
          ,
          <source>arXiv preprint arXiv:2210.13378</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>