<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Optimizing Semantic Data Transformation using High Performance Computing Techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jose Antonio Bernabe-D az</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mar a del Carmen Legaz-Garc a</string-name>
          <email>mcarmen.legaz@ffis.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose M. Garc a</string-name>
          <email>jmgarciag@um.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jesualdo Tomas Fernandez-Breis</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Optimizing SWIT</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Biomedical Informatics and Bioinformatics Platform, IMIB-Arrixaca</institution>
          ,
          <addr-line>Calle Luis Fontes Pagan, n 9, 30003 Murcia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Informatics, University of Murcia</institution>
          ,
          <addr-line>30100 Murcia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The growth of the Life Science Semantic Web is illustrated by the increasing number of resources available in the Linked Open Data Cloud. Our SWIT tool supports the generation of semantic repositories, and it has been successfully applied in the eld of orthology resources, helping to achieve objectives of the Quest for Orthologs consortium. However, our experience with SWIT reveals that the time required for the generation of datasets is longer than desired. In this work we present the application of High Performance Computing techniques, mainly memory optimization and parallelization, to speed up SWIT.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic Web</kwd>
        <kwd>Data transformation</kwd>
        <kwd>High Performance Computing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The application of SWIT to the transformation of large datasets has revealed
that despite the computational complexity increases linearly with the number
of entities to be transformed, it is slower than expected, and we have identi ed
some limitations to the performance: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) use of an interpreted language; (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
inefcient memory management; (3) the execution of identity rules using SPARQL
queries create an execution bottleneck; and (4) the sequential execution of the
transformation. Hence, a code modernization process was carried out3:
1. The SWIT kernel has been re-implemented in C/C++.
      </p>
      <p>2. We have optimized the way SWIT manages the ontology individuals during
the transformation process. We use now two hash maps composed of pointers
to individuals and not copies, so keeping the coherence and reducing memory
consumption. One map grants that no failure happens when searching for an
individual, if this exists, while the other map could miss in some searches since
it acts as a greedy algorithm. The speed up of the execution of searches with the
two maps is up to 2x.</p>
      <p>3. The optimization of memory management also a ects the process of
identifying equivalent individuals using identity rules. Identity rules are de ned using
AND and OR conditions, and the new method uses one hash map of vectors
for each type, where pointers to hashed individuals are stored. The hashing is
carried out di erently depending on the map. For AND conditions, a hash is
performed along all the properties of the entity. Contrariwise, for OR conditions,
the entity properties are hashed separately, having several pointers to the same
individual along the hash map.</p>
      <p>4. The SWIT parallelization is done by using gnu parallel tool4. The
parallel design consists in setting one input le and one SWIT instance per core, so
the parallelization only works when multiple les are established. When
transforming a single large le, we need to split it in several smaller les to enable
parallelization. It provides the largest speed-up.
2</p>
      <p>Results
Our tests show a speed-up of 1000x, 4100x and 7800x in the InParanoid [2]
datasets E.coli, H.arabidopsidis and H.sapiens respectively. The executions were
tested on a high performance server that provides 2 chips of Intel R Xeon R
E52698 v4 with 20 cores each (2 hyper-threading), making a total of 40 physical
cores or 80 virtual cores, running at 2,2 GHz and 128 GB RAM DDR4.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Legaz-Garc</surname>
            <given-names>a</given-names>
          </string-name>
          , M.D.C.,
          <string-name>
            <surname>Min</surname>
          </string-name>
          <article-title>~arro-</article-title>
          <string-name>
            <surname>Gimenez</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tortosa</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fernandez-Breis</surname>
            ,
            <given-names>J.T.</given-names>
          </string-name>
          :
          <article-title>Generation of open biomedical datasets through ontology-driven transformation and integration processes</article-title>
          .
          <source>J. Biomedical Semantics</source>
          <volume>7</volume>
          (
          <year>2016</year>
          )
          <fpage>32</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>O</given-names>
            <surname>'brien</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.P.</given-names>
            ,
            <surname>Remm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Sonnhammer</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.L.</surname>
          </string-name>
          :
          <article-title>Inparanoid: a comprehensive database of eukaryotic orthologs</article-title>
          .
          <source>Nucleic acids research 33(suppl 1)</source>
          (
          <year>2005</year>
          )
          <article-title>D476</article-title>
          {
          <fpage>D480</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>