<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SERVICE RELIABILITY IN THE CLOUD OF DATA CENTERS UNDER OPENSTACK</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>P.V. Fedchenkov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>S.E. Khoruzhnikov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>V.A. Grudinin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>O.L. Sadov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A.E. Shevel</string-name>
          <email>shevel.andrey@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A.B. Kairkanov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>O.I. Lazo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A.A. Oreshkin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Network and Cloud Technologies, ITMO University</institution>
          ,
          <addr-line>St.-Petersburg, 197101</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Research Centre "Kurchatov Institute" PETERSBURG NUCLEAR PHYSICS INSTITUTE</institution>
          ,
          <addr-line>Gatchina, 188300</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>2017 Petr V. Fedchenkov, Sergey E. Khoruzhnikov, Vladimir A. Grudinin, Oleg L. Sadov, Andrey E. Shevel</institution>
          ,
          <addr-line>Arsen B. Kairkanov, Oleg I. Lazo, Anatoly A. Oreshkin</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>282</fpage>
      <lpage>287</lpage>
      <abstract>
        <p>University ITMO (ifmo.ru) is developing the cloud of geographically distributed data centers under Openstack. The term “geographically distributed” in our proposal means data centers (DC) located in different places far from each other at least by many tens of kilometers. Authors follow the conception of “dark” DC, i.e. the DC has to perform normal operation without permanent maintainers even with minor problems (single machine or a number of disk drives went down). In such “dark” DC the staff might visit DC if required to fix appeared problems. Authors describe thoughts and experiments with service reliability for cloud of DCs under Openstack.</p>
      </abstract>
      <kwd-group>
        <kwd>Openstack</kwd>
        <kwd>service reliability</kwd>
        <kwd>cloud of data centers</kwd>
        <kwd>SDN</kwd>
        <kwd>CEPH</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Cloud of Data Centers</title>
      <p>
        The main cloud services to be provided in Data Centers (DCs) include Computing (VMs),
Storage, Data transfer channels. Here it is considered several globally dispersed DCs which are
functioning like united cloud of DCs. There are several advantages for such the architecture [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]:
improved reliability of the data storage due to existing several replicas of data on several DCs,
improved reliability of computing due to possibility to continue the user computing when one DC
went down. We can eliminate a single point of failure and achieve High Availability (HA)
deployment. HA is the ability for the cloud to continue functioning after failure of one or more of the
hardware or software components [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        All DCs are running under Openstack (ocata) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. It means that users have ability to be
logged to Openstack panel Horizon and use computing and other Openstack services divided by
several zones.
      </p>
      <p>
        Obviously DCs need the data links in between each other. In the project it is distinguished
two types of virtual channels: regular channels (encrypted) for data transfer and control channels
(with OpenVPN). Also the cloud needs distributed storage which is implemented with CEPH [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
Further discussion is concentrated to computing service reliability. The conditions where the
reliability is planned to be increased in this project are following:
 User can be logged into any Openstack instance in the cloud of DCs and start up Virtual
      </p>
      <p>Machine (VM) on any DC in the cloud;
 If DC went down where the user uses Openstack and user’s VM was running in the same
DC, then user can log into another Openstack instance in different DC in the cloud and start
new user’s VM instance. Presumably all intermediate results from broken DC are available
in distributed cloud storage.</p>
      <p>In the table 1 several options shown when one DC in the cloud went down. From the table
can be seen that even with cloud consisting of two DCs the reliability to get the computing service
for user is significantly higher to compare with one DC.</p>
      <sec id="sec-1-1">
        <title>Number</title>
        <p>1
2
3
4</p>
      </sec>
      <sec id="sec-1-2">
        <title>Situation</title>
        <p>User was logged into
Openstack instance in
broken DC and user’s VM
was running in the same
DC.</p>
        <p>User was logged into
Openstack instance in
broken DC but user’s VM
is running on different
DC.</p>
        <p>User was logged into
Openstack instance in
running DC but user’s
VM was running in
broken DC
One of the DCs became
broken but user is logged
in Openstack instance in
different DC and user’s
VM is running in running
DC.</p>
        <p>It is assumed that user’s VM is dead but intermediate
results are saved in distributed storage and available in
different DCs. User might start up new VM instance in
different DC and use saved earlier data from distributed
storage.</p>
        <p>No affection to user and user’s VM.</p>
        <p>
          To increase computing service reliability it is used at least one independent Openstack
deployment in each DC. To make united cloud with decentralized administration [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], which is
available for users as integrated resource it is, required to build up the special interconnection of the
all Openstack deployments in different DCs. With all measures being discussed further in the paper it
is planned to increase significantly the computing service reliability in according to the above table.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Cloud Services</title>
      <p>As a storage subsystem, the CEPH platform is chosen to provide long-term distributed
storage. It is configured and deployed the CEPH storage cluster to geographically distributed servers
using the Internet line with nominal capacity 1 Gbit. Servers are divided into two groups (two sites)
which are far from each other at a distance of about 40 Km. In other words it has been configured the
testbed for geographically distributed storage. The only disadvantage of this configuration is the
relatively low speed of replication. At the same time this testbed is quite enough for reliable
archiving purposes for relatively long time (years). Clients might use the virtual storage when part of
data are located on one group of storage servers and another part is in another group in the manner
shown on the fig.1. All the data transfers between remote sites are performed over network tunnel
with OpenVPN.</p>
      <sec id="sec-2-1">
        <title>Clients-1</title>
      </sec>
      <sec id="sec-2-2">
        <title>Clients-2</title>
        <p>CEPH</p>
        <p>Virtual Storage
DC-1
DC-2</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The computing reliability testbed</title>
      <p>In the proposed testbed (fig.2), we focused on providing fault tolerance for computing.
Components of the OpenStack do use common MySql database to store information about project
settings, created virtual entities, provide message delivery. In order to do that database available on
both DCs it were created MySql cluster in Master-Master mode. To make possible control
interconnection by message quesues between sites it was created RabbitMQ. Both keystone database
content are available in each DCs, and therefore you can make changes from any working DC.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Near future testbed improvement</title>
      <p>With number of regions more than two, the cloud configuration with Pacemaker and
HAProxy might give more advantages. Pacemaker can make automation of Active/Passive cluster
availability. HAProxy can be used in this case not only for load balancing, but also to ensure routing
of packets.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Future project developments</title>
      <p>As feature development the reliable data links between DCs are required. It is assumed that
in between DCs more than one data link is to be used. The data links are planned to be configured
with failover option with SDN approach.</p>
      <p>
        To guarantee strong encryption of transferred data in this project is planned use Quantum
Key Distribution [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The network side developments are running with extensive use of approach
NFV.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgement</title>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>This work was financially supported by the Ministry of Education and Science of Russian
Federation (project No. 02.G25.31.0229).</p>
      <p>It was described a range of tests on developed testbed when one DC is down to verify the
reliability of computing services. It is shown that it is possible to increase cloud computing service
reliability under specially configured Openstack in the cloud of geographically distributed data
centers significantly.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Justin</given-names>
            <surname>Riley</surname>
          </string-name>
          , John Noss, Wes Dillingham, James Cuff,
          <string-name>
            <surname>Ignacio M. Llorente</surname>
          </string-name>
          ,
          <article-title>"A High-Availability Cloud for Research Computing"</article-title>
          ,
          <source>Computer</source>
          , vol.
          <volume>50</volume>
          , issue 6 , pp.
          <fpage>92</fpage>
          -
          <lpage>95</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nabi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Toeroe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Khendek</surname>
          </string-name>
          ,
          <article-title>"Availability in the Cloud"</article-title>
          ,
          <source>J. Network and Computer Applications</source>
          , vol.
          <volume>60</volume>
          , pp.
          <fpage>54</fpage>
          -
          <lpage>67</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Sefraoui</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aissaoui</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Eleuldj</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>OpenStack: toward an open-source solution for cloud computing</article-title>
          .
          <source>International Journal of Computer Applications</source>
          ,
          <volume>55</volume>
          (
          <issue>3</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Weil</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brandt</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>E. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>D. D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Maltzahn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2006</year>
          , November).
          <article-title>Ceph: A scalable, high-performance distributed file system</article-title>
          .
          <source>In Proceedings of the 7th symposium on Operating systems design and implementation</source>
          (pp.
          <fpage>307</fpage>
          -
          <lpage>320</lpage>
          ). USENIX Association.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Almutairi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarfraz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basalamah</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aref</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ghafoor</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>A distributed access control architecture for cloud computing</article-title>
          .
          <source>IEEE software</source>
          ,
          <volume>29</volume>
          (
          <issue>2</issue>
          ),
          <fpage>36</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Chistyakov</surname>
            ,
            <given-names>V. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sadov</surname>
            ,
            <given-names>O. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vasiliev</surname>
            ,
            <given-names>A. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Egorov</surname>
            ,
            <given-names>V. I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kompaniets</surname>
            ,
            <given-names>M. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fedchenkov</surname>
            ,
            <given-names>P. V.</given-names>
          </string-name>
          , ... &amp;
          <string-name>
            <surname>Khoruzhnikov</surname>
            ,
            <given-names>S. E.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Software-defined subcarrier wave quantum networking operated by OpenFlow protocol</article-title>
          .
          <source>arXiv preprint arXiv:1709</source>
          .
          <fpage>09081</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>