<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Alignment as a Service</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rob Upson</string-name>
          <email>robert.upson@city.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ernesto Jiménez-Ruiz</string-name>
          <email>ernesto.jimenez-ruiz@city.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Knowledge Graph Alignment, Cloud Computing, REST API,</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>City, University of London</institution>
          ,
          <addr-line>London</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SIRIUS, University of Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <fpage>23</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>Researchers and practitioners in the field of knowledge graphs rely on a variety of tools to perform knowledge graph alignment. In order to broaden the domain of a knowledge scientist, the process of aligning knowledge graphs must be accessible to a broader range of users, and in a way that can scale to meet the needs of a growing userbase. This paper presents a cloud-based, language agnostic approach that aims at providing knowledge graph alignment tooling at scale.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>CEUR
Workshop
Proceedings
1The term knowledge graph alignment supersedes the notion of ontology alignment or matching to emphasise the
need of aligning not only the terminological part (i.e., concepts and relationships), but also the data (i.e., instances
of the concepts).
with the MELT platform, one must first use Maven to produce Docker images, then provide the
images to the MELT client. This is useful because any matching system could be swapped into
the MELT client as a Docker image, provided it meets the contract, but it does not solve the
problem of a zero-dependency alignment system.</p>
      <p>This paper presents a proof of concept system that provides the matching functionality of
L o g M a p over a HTTP RESTful API. This gives the consumer absolute control and agency over
how to integrate with L o g M a p in a way that best fits their use case - decoupling L o g M a p from any
programming language or paradigm. A Web app has also been produced that consumes the new
API, meaning that a user can use a UI on the Web to perform alignments if this is preferable to
a programmatic approach.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Architecture Overview</title>
      <p>The proposed Knowledge Graph Alignment as a Service (K G A S ) system is comprised of 3
individual containerised applications (see Figure 1), deployed onto cloud infrastructure using Azure.
Azure is the cloud provider chosen to host the working proof of concept on the Web, but the
system is designed such that it could be deployed anywhere, be that to another cloud provider
or run on a local server. It was important not to change L o g M a p ’s source code directly for the
project. The purpose of the proof of concept is to show that the functionality of alignment tools
can be exposed without impacting the implementation. This way L o g M a p can still be developed
as a standalone tool, and K G A S will continue to work alongside it without needing to make
changes. This was achieved by creating a Docker image that clones L o g M a p ’s source code from
Git using the latest commit hash, and builds the application from source. This means that if a
new release of L o g M a p were to be made, the app could simply rebuild with the new commit hash
and users get the latest version of L o g M a p immediately. Once there is a dockerised version of
L o g M a p running in a container, the REST API needs to be able to use that app. Since the API
runs in a diferent container to L o g M a p , there needs to be an interface for them to communicate.
The L o g M a p container exposes a port accepting SSH connections for this, which enables the CLI
commands to be run by the REST server via SSH.
2.1. Cloud Implementation
The containerised applications are platform-agnostic, however, in order to serve the KG
alignment outputs as zip files, some external storage is required, which is where the proof of concept
becomes coupled to Azure. When a user starts an alignment via the API, a unique ID is generated
for the process. This is the ID a user will query against to retrieve the results of the alignment,
and also represents the file name used when writing to the remote storage.</p>
      <p>For the project to be a success, the solution must be repeatable. To ensure this, all of the
cloud resources were created using Infrastructure as Code (IaC); specifically, Terraform. 3 This
means that with access to the codebase anyone could spin up their own version of K G A S on their
own infrastructure. This provides a wide range of benefits to maintainability of the project as
well as scalability.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Web Interface</title>
      <p>The K G A S Web app created for the proof of concept is a small React application4 with a UI
allowing a user to upload two KGs for alignment, begin the alignment, and query for results of
ongoing or previously completed alignments by their id (see Figure 2). The UI will store the
alignment ID in the browser’s storage, so there is no requirement for a user to remember them.
The Web app’s purpose is twofold. Firstly, to present a zero-dependency, out of the box way to
perform KG alignment. Secondly, it serves as an exemplar piece on how to integrate with the
REST API to access L o g M a p ’s functionality and retrieve alignment results. This will be helpful
for those wishing to integrate with the API programmatically, for example when running an
alignment in a Python script.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Impact</title>
      <p>The working concept shows that KG alignment can be decoupled from its implementation and
provided to consumers. This iteration is limited to L o g M a p specifically, but the same process of
containerising and wrapping in a REST API layer can be applied to any alignment system. L o g M a p
itself has more functionality than that which is exposed by K G A S , including the repair of a given
alignment or the division of the KG alignment process into subtasks. Exposing these additional
behaviours only requires a more mature API, where various parameters or configurations can
be passed as arguments. The most significant room for expansion is to decouple from L o g M a p
altogether. Currently K G A S relies on ssh-ing into a machine with the L o g M a p CLI, then running
L o g M a p specific commands inside that machine. If the interface through which alignments were
performed could be generalised, then K G A S would not need to know specific commands for any
3https://www.terraform.io/
4https://reactjs.org/
of the systems. The MELT platform provides exactly this. MELT relies on being provided a
Docker image for an alignment tool, which it will use for alignment using its common interface.
A container could be created ro run MELT and include pre-made images saved to the disk that
are compatible with MELT (many of these already exist). Then the API layer only needs to
know the CLI commands to interact with MELT, and it can use any one of the alignment tools
baked into the image. This will enable a comprehensive KG alignment API where a user can
pick and choose the best tool for their use case on demand without having to learn or use any
additional technologies. It will also provide a platform where the various tools can be compared
with one another, since they will all be in one place and running on the same hardware. The
working concept in this demo shows that such a system is within reach.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Demonstration</title>
      <p>For the demonstration, we will upload to the Web app KGs from the Ontology Alignment
Evaluation Initiative5 like the h u m a n . o w l and m o u s e . o w l anatomy ontologies.</p>
      <p>
        The Web app will begin the alignment and display the id. It will then begin polling for results
until the alignment completes, at which point the UI will present a button to download the
alignment output as a zip file, which can be extracted to view the individual files. In the case of
L o g M a p , the alignment output is composed by the following files:
• l o g m a p 2 _ m a p p i n g s in table format (.txt and .tsv), OWL format (.owl), and RDF alignment6
format (.rdf): final output mappings selected by L o g M a p .
• l o g m a p _ a n c h o r s (same formats as above): mappings computed by L o g M a p with very high
confidence.
• l o g m a p _ d i s c a r d e d _ m a p p i n g s . t x t and l o g m a p _ h a r d _ d i s c a r d e d _ m a p p i n g s . t x t : set of
mapping that were considered by LogMap but eventually discarded.
• l o g m a p _ l o g i c a l l y _ c o n f l i c t i n g _ m a p p i n g s . t x t : mappings that were discarded during the
logic-based repair process [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
• l o g m a p _ o v e r e s t i m a t i o n . t x t : set of all mappings considered by L o g M a p , which represents
a manageable subset of the Cartesian product between the entities of the input KGs.
• m o d u l e 1 _ o v e r l a p p i n g _ l o g m a p 2 . o w l and m o d u l e 2 _ o v e r l a p p i n g _ l o g m a p 2 . o w l : represent the
locality modules of the input KGs taking as signature the entities within L o g M a p ’s output
mappings (see context of an alignment in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]).
      </p>
      <p>The demo will be performed within a Web browser, and no additional dependencies will be
required. The K G A S source code, API documentation and links to the Web app and REST API are
available on GitHub: https://github.com/rupson/knowledge-graph-alignment-as-a-service.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Euzenat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shvaiko</surname>
          </string-name>
          , Ontology matching, Springer,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Cuenca Grau</surname>
          </string-name>
          ,
          <article-title>LogMap: Logic-Based and Scalable Ontology Matching</article-title>
          , in: 10th
          <source>International Semantic Web Conference</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>273</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Cuenca</given-names>
            <surname>Grau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Horrocks</surname>
          </string-name>
          ,
          <article-title>Large-scale interactive ontology matching: Algorithms and implementation</article-title>
          ,
          <source>in: 20th European Conference on Artificial Intelligence (ECAI)</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>444</fpage>
          -
          <lpage>449</lpage>
          .
          <source>doi:1 0 . 3 2</source>
          <volume>3 3 / 9 7 8 - 1 - 6 1 4 9 9 - 0 9 8 - 7 - 4 4</volume>
          <fpage>4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hertling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Portisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          , MELT - Matching EvaLuation Toolkit, in: 15th
          <source>International SEMANTiCS Conference</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>231</fpage>
          -
          <lpage>245</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agibetov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Samwald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Cross</surname>
          </string-name>
          ,
          <article-title>Dividing the ontology alignment task with semantic embeddings and logic-based modules</article-title>
          ,
          <source>in: 24th European Conference on Artificial Intelligence (ECAI)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>784</fpage>
          -
          <lpage>791</lpage>
          .
          <source>doi: 1 0 . 3 2 3 3 / F A I A 2 0</source>
          <volume>0 1 6 7 .</volume>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>