<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extracting Information from Medieval Notarial deeds?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Charlene Ellul</string-name>
          <email>charlene.ellul@um.edu.mt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joel Azzopardi</string-name>
          <email>joel.azzopardi@um.edu.mt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Charlie Abela</string-name>
          <email>charlie.abela@um.edu.mt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Malta</institution>
          ,
          <country country="MT">Malta</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Notarial Archives in Valletta houses a collection of Latin Notarial deeds that has not been exploited yet. In this paper, Machine Learning techniques are proposed and implemented to extract entities such as people, place names, dates, deed types and keywords from these historical texts. Both supervised and unsupervised techniques are considered and compared with baseline models. Experimental results on a subset of these documents are already showing results that outperform the baselines for Latin text such as those from CLTK. Evaluation was carried out using indexes of four published Notarial Registers.</p>
      </abstract>
      <kwd-group>
        <kwd>Named Entity Recognition</kwd>
        <kwd>Keyphrase Extraction</kwd>
        <kwd>Classi cation</kwd>
        <kwd>Latin Text</kwd>
        <kwd>Historical Texts</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Archives around the world are a source of hidden information. One of these
archival gems is found at the Notarial Archives in Valletta, Malta and houses
around 20,000 notarial deeds dating back to the 13th century. It is typically the
case that archives publish high quality images and metadata of the structure
of historical documents, but the content itself is not exposed in a meaningful
way to aid historians. In literature, some researchers [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] dedicated their e orts
to mine data from medieval Latin documents. The extraction of named entities
such as dates, places and people can be used to aid historical research in areas
such as genealogy and toponymy.
      </p>
      <p>
        Most of the notarial deeds found in the Valletta archives fall under categories
such as wills, dowry and transfer of land. Although some notaries used to write
the deed type, this was not a requirement. Thus, text classi cation can be used
to maximize the use of the remaining words to predict the presented deed type.
Automatic keyphrase extraction can be used to express a document as a set of
keywords/keyphrases[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A notarial deed can be represented in a similar way to
shed light on its content and avoid unnecessary handling whilst also reducing
the time required for archival researchers and enthusiasts to nd what they are
looking for.
? This work is partially funded by project E-18LO28-01 as part of the collaboration
between the Notarial Archives in Valletta and the University of Malta.
      </p>
      <p>
        In our research we use four Latin notarial transcribed registers entitled
'Documentary Sources of Maltese History' and compiled by Professor Stanley
Fiorini[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. These are the only existent transcribed documents from the collection
dating back to the 15th century with 981 deeds. Our main goal is to extract
entities such as dates, people, places, deed types and keyphrases. Annotating
a large corpus of data requires expertise and time, fortunately, however, these
publications include indexes for place names, persons and subjects which could
be used to annotate the deeds and also for evaluation.
      </p>
      <p>In the rest of the paper, we discuss some related work in Section 2 and present
the adopted methodology in the following section. In Section 3 we present some
initial evaluation which is followed by some future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Latin poses a great challenge for Named Entity Recognition (NER). Annotated
corpora that can be used for training are scarce and most of them focus on
classical texts. Conditional Random Fields (CRF) have been used successfully for
training Latin models as in Aguilar et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] who applied CRF on a database of
Burgundy cartularies which were manually annotated. Text classi cation
techniques based on supervised machine learning have also been used in the context
of di erent languages[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Both supervised and unsupervised models for Keyphrase detection and
extraction have been used successfully[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A keyphrase extraction model is usually
based on a list of extracted candidate words and some heuristic such as stopword
removal through which candidate keywords are ltered out. RAKE1, TextRank2
and TF-IDF are three popular unsupervised approaches that have been applied
on generic languages [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A more domain-speci c keyphrase extraction method
was developed by Witten et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] who designed the KEA supervised algorithm.
Candidate keyphrases (up to 3 words) are ltered before computing TF-IDF
and the distance from the start of document for each candidate as features and
then fed into a Naive Bayes Model. CRF were used by Zhang et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] using a
number of features among which are length of word, POS tag, previous words,
next words and TF-IDF. This was tested on a Chinese text and yielded the best
F1 score compared to SVM and other baseline models.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>
        We used the indexes found in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to annotate the text with people's names
and places for the NER, and keyphrases for keyphrase extraction. Typically a
keyphrase index has the following form, Coquine domus/domuncula 226,
241242, 396, with the term and the deeds containing the term. There exists however
an electronic version of a single index, while the other copies are available as hard
      </p>
      <sec id="sec-3-1">
        <title>1 https://github.com/fabianvf/python-rake</title>
      </sec>
      <sec id="sec-3-2">
        <title>2 https://github.com/davidadamojr/TextRank</title>
        <p>Extracting Information from Medieval Notarial deeds
copies. Furthermore, there is no index available for dates. A dictionary of all
possible mentioned entities was compiled using the Ratcli -Obershelp distance
algorithm3 to annotate the text with the relevant tag to be used for evaluation.</p>
        <p>
          Dates are presented using indictions for the year and thus we had to work
out the indiction cycle of each act. Notaries tend to use shorthand when writing
dates such as eodem (same date as before) and penultimo (day before the last).
For this reason a rule based entity extraction was implemented to convert the
dates to modern dates. Extraction of person and place entities was performed
using a trained CRF model based on su xes, uppercase, title, digit, POS, lemma,
next and previous words. The POS and lemma tags were derived using Schmidt's
treetagger4 using parameters for Latin to improve accuracy. We used the Fiorini's
register [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to train and test our model which was then compared with existing
libraries such as the Classical Language Toolkit (CLTK) 5, Spacy (multilingual
model)6 and Stanford's NER (Spanish Model - Latin derivative)7.
        </p>
        <p>Deeds can have a variety of categories and these were generalized for text
classi cation. In total there are 981 deeds and 79 di erent categories with an
average of 12 deeds per category. Some of the categories include only one deed,
making the training set highly imbalanced. Di erent feature vectors were tried
out including count vectors, word level TF-IDF, n-grams and character level
vectors. Di erent models were trained for deed classi cation using Naive Bayes,
Linear Classi er, SVM and Random Forest. The model with the highest accuracy
was saved.</p>
        <p>
          The index of keyphrases was used to annotate the corpus for keyphrase
extraction. Lemmas were used for comparisons as Latin often uses declensions. The
index was merged with the deed text using exact, lemma and stem matches. A
list of annotated/non-annotated words was kept for each deed to be used for
evaluation. Generic unsupervised approaches were used for keyphrase extraction
including TextRank, RAKE and TF-IDF, however these yielded unsatisfactory
results with RAKE giving the best results as shown in Table 1. We then used
a variant of the supervised approach presented by [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] called KEA which due to
its candidate phrases ltering did not yield good results. A CFR algorithm was
implemented using the same technique to extract entities for people and places,
with the addition of TF-IDF and distance features giving the best results as
shown in Table 1.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>The results achieved for the extracted entities are already very promising. Both
NER and keyphrase extraction were done using 100 acts (indexes of other
registers are yet to be used for annotations) with a 70%-30% split (results in Table 1).</p>
      <sec id="sec-4-1">
        <title>3 https://docs.python.org/2/library/di ib.html</title>
      </sec>
      <sec id="sec-4-2">
        <title>4 http://www.cis.uni-muenchen.de/ schmid/tools/TreeTagger/</title>
      </sec>
      <sec id="sec-4-3">
        <title>5 http://docs.cltk.org/en/latest/index.html</title>
      </sec>
      <sec id="sec-4-4">
        <title>6 https://spacy.io/</title>
      </sec>
      <sec id="sec-4-5">
        <title>7 https://nlp.stanford.edu/software/CRF-NER.shtml</title>
        <p>Since we used a domain speci c corpus, supervised models gave better results.
The dataset was highly imbalanced and the text classi cation was performed on
the whole corpus of 981 records with a 75%-25% split. A Linear classi er was
used that leveraged on CLTK's stopword list and the count vector features giving
an accuracy of 72%. The removal of stopwords improved slightly the achieved
results across all trained models.
Purpose Method Precision Recall F1 score
NER People/Places CLTK 0.339 0.113 0.17
NER People/Places Spacy with multilingual model 0.414 0.922 0.572
NER People/Places Stanford NER with Spanish model 0.152 0.947 0.263
NER People/Places Conditional Random Fields Model 0.956 0.957 0.956</p>
        <p>RAKE with CLTK and</p>
      </sec>
      <sec id="sec-4-6">
        <title>Keyphrase extraction Voyant tools stop words1 0.257 0.189</title>
        <p>Purpose Method
Keyphrase extraction Conditional Random Fields</p>
      </sec>
      <sec id="sec-4-7">
        <title>1 https://github.com/aurelberra/stopwords</title>
        <p>4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Future Work</title>
      <p>In this paper, we presented our initial research on extracting entities and
keyphrases from historical Latin texts. The results are very encouraging even though
the datasets are fairly small. We plan to digitize the indexes of the other registers
in Fiorini's collection so that we can train the models with more data. We will
furthermore be using the extracted information to create a knowledge graph for
the Notarial Archives.</p>
      <p>O-KEY
0.982
0.15
F1 score
B-KEY I-KEY
0.751 0.465</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Aguilar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tannier</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Chastang</surname>
          </string-name>
          ,
          <article-title>Named entity recognition applied on a database of Medieval Latin charters. The case of chartae burgundiae</article-title>
          . M.
          <string-name>
            <surname>Dring</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Jatowt</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Preiser-Kapeller</surname>
          </string-name>
          , A. van den Bosch,
          <year>2016</year>
          , pp.
          <volume>67</volume>
          {
          <fpage>71</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>K.</given-names>
            <surname>Saidul Hasan</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <article-title>Automatic Keyphrase Extraction: A Survey of the State of the Art</article-title>
          .
          <source>Association for Computational Linguistics</source>
          ,
          <year>2014</year>
          , pp.
          <volume>1262</volume>
          {
          <fpage>1273</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>S.</given-names>
            <surname>Fiorini</surname>
          </string-name>
          ,
          <article-title>Documentary Sources of Maltese History Part I Notarial Documents No 3 Notary Paulo Bonello</article-title>
          , Notary Giacomo Zabbara, 1st ed. University of Malta,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>A.</given-names>
            <surname>Al-Thubaity</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Abanumay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Al-Jerayyed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alrukban</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Mannaa</surname>
          </string-name>
          , \
          <article-title>The e ect of combining di erent feature selection methods on arabic text classi cation," in 2013 14th IEEE/ACIS</article-title>
          , SNPD,
          <year>2013</year>
          , pp.
          <volume>211</volume>
          {
          <fpage>216</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>I. H</given-names>
            <surname>Witten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. W</given-names>
            <surname>Paynter</surname>
          </string-name>
          , E. Frank,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gutwin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. G</given-names>
            <surname>Nevill-Manning</surname>
          </string-name>
          , \Kea:
          <article-title>Practical automatic keyphrase extraction</article-title>
          ,"
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liao</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          , \
          <article-title>Automatic keyword extraction from documents using conditional random elds,"</article-title>
          <source>Journal of Computational Information Systems</source>
          , pp.
          <volume>1169</volume>
          {
          <issue>1180</issue>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>