<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Semantic Alignment of Heterogeneous Structures and Its Application to Digital Humanities</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Renata Vieira</string-name>
          <email>renatav@uevora.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cassia Trojahn</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CIDEHUS, University of E ́vora</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IRIT, UMR 5505</institution>
          ,
          <addr-line>1118 Route de Narbonne, F-31062 Toulouse</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The field of Digital Humanities comprises the use of technology within arts, heritage
and humanities research. This brings new methods of inquiry, new means of
dissemination, but also constitute a new core of investigation in itself. Not only creation and access
to collections of interest for these areas have improved with digitalization of research
material, but further use of computing technology is being proposed and discovered [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
The primary source of information for humanities researchers comes from free,
unstructured sources in written language, that is ambiguous and context-dependent. Also the
humanities might face difficulties due to the particularities of the source of information,
that might be available in ancient forms of registration. For instance, there is a need
for identifying specific vocabulary of a historical period and also align non uniform
spelling which was usual in old publications [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In this perspective, the ability to
establish a relationship between different forms of expression of knowledge (from structured
and unstructured sources) and its meaning or intent is crucial [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This scenario reflects
a unifying framework of a wide range of solutions from a variety of domains, including
NLP and semantic web.
      </p>
      <p>
        Different variants of the notion of ‘alignment’ have been adopted in a range of areas,
focusing on homogeneous structures (e.g., text alignment [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], database alignment [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] or
ontology alignment [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) or heterogeneous structures (e.g., annotation of text with
ontologies [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], alignment of dictionaries and ontologies [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], alignments between relational
databases and ontologies [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]). These alignment approaches, however, take little account
of the alignment of multiple structures. This type of approach is becoming increasingly
necessary to manage the growing volume of unstructured information sources available
on the Web (encyclopedias such as Wikipedia, social media data, etc.) and LOD
knowledge bases. In addition, the approaches are mostly developed for the English language.
These needs have to be addressed through a global vision of alignment that takes into
account a multiplicity of structures in which knowledge can be expressed. This paper
seeks a holistic approach to semantic computing and alignment, when considering
heterogeneous structures in which knowledge is represented.
? Copyright c 2020 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).
      </p>
    </sec>
    <sec id="sec-2">
      <title>Proposal</title>
      <p>The approach consists of two main steps. First, knowledge extraction approaches will
be applied to extract the terminology of the relevant corpora. We plan to specialise
general language models, since the corpora present distinctive language characteristics
due to scope and time. We also plan to make use of techniques for the recognition of
named entities which might help finding important relations and events. On the basis of
the models and recognised entities we plan to extract other information with the help
of semantic alignment methods. Second, the extracted terminology will be aligned to
existing sources of knowledge (available dictionaries, lexicons, corpora and ontologies).
In particular, there are basic ontological concepts describing fundamental elements such
as persons, places, periods, and that have to be anchored to what is extracted. Ontologies
will be the central focus for semantic alignment of textual occurrences of concepts,
and its relations with other semantic sources. The alignment may consider previous
semantic knowledge, or might be inferred trough semantic similarity analysis.</p>
      <p>
        We plan to apply our approach on current projects such as the Curvo Semedo’s
works [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This is a corpus integrated by six works published between 1707 and 1727,
authored by Alentejo doctor Joa˜o Curvo Semedo (1635-1719), containing medical and
pharmacological knowledge constituted and published in Portuguese. The focus reader
of his works, at the time they were recorded, was a less educated person, little affected
by the materials available only in Latin. The six works gathered include a collection of
about 2,150 pages, which are treated and offered in the form of transcripts, in different
formats, in original spelling and reproduced, accompanied by descriptions of their
terminologies and representations of the content of each one, generated with the support
of computational tools. The evaluation phase will be carried out with he help of
humanities expert. The proposed methodology has potential utility for other projects with a
variety of history and linguistic inquiries.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Cole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          , et al.
          <article-title>The ribosomal database project: improved alignments and new tools for rrna analysis</article-title>
          .
          <source>Nucleic acids research</source>
          ,
          <volume>37</volume>
          :
          <fpage>D141</fpage>
          -
          <lpage>D145</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>B.</given-names>
            <surname>Dalvi</surname>
          </string-name>
          , E. Minkov,
          <string-name>
            <given-names>P. P.</given-names>
            <surname>Talukdar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W. W.</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <article-title>Automatic gloss finding for a knowledge base using ontological constraints</article-title>
          .
          <source>In 8th Conf. WSDM</source>
          , pages
          <fpage>369</fpage>
          -
          <lpage>378</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>M.</given-names>
            <surname>Erdmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maedche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.-P.</given-names>
            <surname>Schnurr</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          .
          <article-title>From manual to semi-automatic semantic annotation</article-title>
          .
          <source>In COLING Workshop on Semantic Annotation</source>
          , pages
          <fpage>79</fpage>
          -
          <lpage>85</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>J.</given-names>
            <surname>Euzenat</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Shvaiko</surname>
          </string-name>
          . Ontology matching. Springer-Verlag, Heidelberg, Germany,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>M.</given-names>
            <surname>Matuschek</surname>
          </string-name>
          and
          <string-name>
            <given-names>I.</given-names>
            <surname>Gurevych</surname>
          </string-name>
          .
          <article-title>Dijkstra-wsa: A graph-based approach to word sense alignment</article-title>
          .
          <source>TACL</source>
          ,
          <volume>1</volume>
          :
          <fpage>151</fpage>
          -
          <lpage>164</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>P.</given-names>
            <surname>Quaresma and M. J. B. Finatto</surname>
          </string-name>
          .
          <article-title>Information extraction from historical texts: a case study</article-title>
          .
          <source>In DHandNLP@PROPOR</source>
          , pages
          <fpage>49</fpage>
          -
          <lpage>56</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S.</given-names>
            <surname>Schreibman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Siemens</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Unsworth</surname>
          </string-name>
          .
          <article-title>A new companion to digital humanities</article-title>
          . John Wiley &amp; Sons,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. D. Tufis¸,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Barbu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Ion</surname>
          </string-name>
          .
          <article-title>Extracting multilingual lexicons from parallel corpora</article-title>
          .
          <source>Computers and the Humanities</source>
          ,
          <volume>38</volume>
          (
          <issue>2</issue>
          ):
          <fpage>163</fpage>
          -
          <lpage>189</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. D. Un˜a, N. Ru¨mmele, G. Gange, et al.
          <article-title>Machine learning and constraint programming for relational-to-ontology schema mapping</article-title>
          .
          <source>In 27th IJCAI</source>
          , pages
          <fpage>1277</fpage>
          -
          <lpage>1283</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>