<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CoToHiLi: computational tools for historical linguistics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alina Maria Cristea</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anca Dinu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liviu P. Dinu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simona Georgescu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ana Sabina Uban</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laurent</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>iu Zoicas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Research group: Human Language Technologies Research Center, University of Bucharest</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Bucharest</institution>
        </aff>
      </contrib-group>
      <fpage>31</fpage>
      <lpage>34</lpage>
      <abstract>
        <p>This project represents a computational framework for historical linguistics. The general purpose of the CoToHiLi project is to integrate expert knowledge and computational power to address cognate identification, cognate-borrowing discrimination, Latin protoword reconstruction and semantic divergence. The goal of the project is twofold: 1) to automate certain parts of the traditional work-flow of the comparative method (such as the collection of data or the automatic alignment based on predefined or inferred rules), and 2) to bring new insights or avenues of investigation, which might not be easily accessible otherwise (e.g., the automatic identification of patterns and regularities in large amounts of data). The project will provide tools for the main Romance kernel group (French, Italian, Portuguese, Romanian, Spanish), as well as Latin. The methodologies and computational tools proposed could also serve as a basis for further development for other comparable language families, including less studied languages, with scarce resources available.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Historical linguistics</kwd>
        <kwd>cognates</kwd>
        <kwd>semantic divergence</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>2018) explored the potential correlation of genetic
and linguistic distances, starting from what he called
The general purpose of the CoToHiLi1 project is Darwin’s last challenge: “If we possessed a perfect
to integrate expert knowledge and computational pedigree of mankind, a genealogical arrangement of
power to address the following topics: cognate iden- the races of man would aford the best classification
tification, cognate-borrowing discrimination, Latin of the various languages now spoken throughout
proto-word reconstruction and semantic divergence. the world; and if all extinct languages, and all
inOur project is focused on the Romance languages termediate and slowly changing dialects, were to
(French, Italian, Portuguese, Romanian, Spanish), be included, such an arrangement would be the
and will provide tools for the main Romance kernel only possible one” (see also [2]). Given that the
group and for Latin. The duration of the project is socio-economical and cultural factors are some of
3 years, starting from January 2021. the motivations for borrowing from one language</p>
      <p>The research problems that we address are sig- to another [1, 3], the topic of this research project
nificant on multiple levels. From a scientific point facilitates reconstructing certain aspects related to
of view, any advance in historical linguistics is of society and culture for groups of people speaking
paramount cultural importance, being inherently a given proto-language, and gaining insights into
connected with human history (“each word a his- their past social interactions and into their social
tory”, cf. [1]). Longobardi (LanGeLin project, 2012- and cultural practices [3]. Moreover, establishing
the direction and source of borrowing is important
to our understanding of the social relations between
the groups involved. From a technological
perspective, as linguistic change is the most visible at the
lexical and semantic level, computational tools can
be designed to serve both aspects, for instance to
automatically identify related words and to assess
the semantic change. Even though historical
lexicology has leveraged technological advances, and
some pioneering work was initiated on various steps
of the work-flow (cognate identification, proto-word
where reconstruction), historical semantics has not
sufiupdates: ciently benefited from the advances in computer
science. Yet, by drawing special attention to the English. Our aim is to track the semantic change of
semantic divergence occurring in pairs of cognates, words in Latin and across multi-languages, in the
we could both take a few steps forward towards Romance language family, for the first time, with
a unitary theory of semantic change, and improve the substantial purpose of looking for common
patpractical applications such as automatic translation terns characterizing the overall semantic divergence
systems or language e-learning systems, aware of cases. Additionally, we intend to explore the
stafalse friends and related phenomena. tistical properties of the word embedding vectorial
spaces [9, 10].</p>
    </sec>
    <sec id="sec-2">
      <title>2. Objectives</title>
      <p>The innovation of the project consists in integrating
linguists’ knowledge with new computational meth- The methodologies and computational tools we
proods in a unified framework, to address important pose could extend their applicability not only to
problems from historical linguistics, enabling ex- various linguistic branches in the Indo-European
perts to provide input and feedback throughout the family, but also to less studied languages or
linwhole development process, in the pre-processing, guistic families. Such advances could provide new
annotation, feature engineering, training and evalu- answers in historical and social sciences, given that
ation phases: lexical and semantic change is a key source of clues
regarding both the dynamics of cultural interactions
2.1. Identification of related words between groups in the past, and the technological
innovations and exchanges that have taken place
We aim at going one step further than the cur- across space and time [3]. Moreover, the semantic
rent state-of-the-art methods by: a) proposing a data provided by the CoToHiLi project could be
more in-depth analysis, by identifying the direction of great help for the cognitive sciences and
neuroof the borrowings and b) automatizing the whole sciences [11], to the extent that they can ofer a new
process as a pipeline that, given a pair of input perspective on our brain mechanisms.
words, provides an automatic analysis regarding the As for the socio-economic impact of the
CoTorelationship between them [4, 5, 6]. HiLi project, in the context of the increasing
number of attempts to create automatic tools designed
2.2. Latin proto-word reconstruction for linguistic comprehension, our computational
devices could support Romance intercomprehension
To improve previous results, we intend to use more by bringing into light the common linguistic
fearecent techniques [7], such as conditional random tures, as well as the semantic relations between the
ifelds (CRF) for sequence labelling and deep learn- Romance cognates or borrowings. Such an advance
ing, in particular character-level neural networks. can prove its usefulness in the constant eforts to
The alignment technique, which stands at the foun- improve the automatic translation systems.
dation of our approach, will be improved by an
heuristic for choosing the best alignment. We
also address the challenging problem of multiple 4. Methodology
alignment (finding an alignment for more than two
words), in order to be able to extract knowledge
from cognate sets in multiple languages. Another
promising line of research is to make use of the more
recent Latin resources, such as The Latin Diachronic
Database [8].</p>
      <p>
        For the first two objectives, our methodology is
focused on two main aspects: creating clean datasets
and developing computational methods for
achieving the proposed research tasks. For the Romance
languages there are already some existing resources
(for cognates, for borrowings and for proto-word
reconstruction), but they are scattered, incomplete,
2.3. Diachronic semantic divergence or with uncertain availability (cf. [12, 13]). Thus,
Semantic change is a continuous and complex pro- datasets do not have to be built from scratch, but
cess ([1] presents no less than 11 types of semantic the data need to be harmonized, verified and
enchange), which has been recently studied in the hanced where necessary, in order to become a
benchcontext of distributional semantics theory. Vec- mark in the domain. By using computational tools,
torial representations of word meaning (word em- corroborated by the direct intervention of classical
beddings) have been used for tracking semantic linguists, we have already built a significant part
shifts across diferent time periods, especially for of the database, representing the starting point for
3. Impact
the computational methods that are being devel- automatically discriminating between inherited and
oped. We have continued with the alignment of borrowed Latin words. We have introduced a new
word pairs. Given the lack of an unanimously ac- dataset and investigated the case of Romance
lancepted alignment method [
        <xref ref-type="bibr" rid="ref3">13, 14</xref>
        ], we confront a guages - where words directly inherited from Latin
semi-automatic manner of choosing the alignment coexist with words borrowed from Latin -, and
exwith the knowledge of classical linguists, in order plored whether automatic discrimination between
to establish an heuristic capable of making the best them was possible. An initial trial was to
autochoice. From the alignment, we extract features for matically predict whether a word was inherited or
machine learning models. We improve current exist- borrowed by simply taking into account its intrinsic
ing computational methods with linguistic features structure, given that borrowed words are
presumprovided by experts. We develop a machine-learning ably less eroded than inherited ones, subject to
classifiers (using support vector machines), sequen- historical sound shifts. We then took a step farther
tial models (using CRF and neural networks) and and employed n-gram character features extracted
ensemble techniques. Moreover, we experiment with from the word-etymon pairs and from their
alignnew ensemble techniques, to improve the overall per- ment, which led to considerably better results [6].
formance by combining results from multiple sister For the third objective, a first step has been taken
languages. We are currently working with the or- with the investigation of the semantic divergence of
thographic form of the words, while for Romanian, cognate pairs in English and Romance languages.
Spanish and Italian we are planning to also use the To this end, we introduced a new curated dataset
phonetic transcription. of cognates in all pairs of those languages. We
de
      </p>
      <p>For the third objective, in order to identify seman- scribed the types of errors that occurred during
tic shifts across time periods as well as languages, the automated cognate identification process and
we leverage vector space representations of meaning, manually corrected them. Additionally, we labeled
or word embeddings, relying on traditional models the English cognates according to their etymology,
such as word2vec and FastText [15, 16], as well as separating them into two groups: old borrowings
experimenting with state-of-the-art language mod- and recent borrowings. On this curated dataset,
els such as BERT [17]. The method consists of we analysed word properties such as frequency and
building vectorial semantic representations for the polysemy, and the distribution of similarity scores
words in each of the target languages, based on the between cognate sets in diferent languages. We
multilingual corpora, and then obtaining a shared automatically identified diferent clusters of English
multilingual semantic space. This will allow us to cognates, setting a new direction of research in
cogcompute semantic distances between cognates as nates, borrowings and possibly false friends analysis
well as analyze the statistical and the linguistic in related languages [10, 19].
properties of words whose meanings have diverged.</p>
      <p>The available corpora are unequal from one
language to another; for instance, the Royal Spanish 6. Conclusions
Academy provides an exhaustive diachronic corpus
of its language, whereas for Romanian we only have Drawn within a computational framework, the
Coaccess to a scarce data-base, composed of a fairly ToHiLi project addresses key concerns of
historilimited number of old texts. In order to ensure the cal linguistics centered on the Romance languages,
accuracy of our analysis, in this stage of the project, such as cognate identification, cognate-borrowing
we use mainly lexicographic resources, as well as discrimination, Latin protoword reconstruction and
data-bases built for the contemporary stage of each semantic divergence, towards which we have taken
language (such as multilingual Wikipedia2). a few steps forward by performing various
experiments. At this stage of the project, we analyze
only the main five Romance languages (French,
Ital5. Current Results ian, Portuguese, Romanian, Spanish), but as we
advance we intend to include other Romance idioms
For the first two objectives, we have started building as well. We predict that the methodologies and
datasets of cognates and borrowed words for the computational tools proposed will also serve as a
Romance languages [18]. This first step relies on basis for further development for other comparable
dictionaries that contain etymological information language families, including less studied languages,
(e.g., for Romanian we use 13 dictionaries available with scarce resources available.
in digital format). We have proposed a new method</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>cient Languages Using Probabilistic Models of Research supported by the Ministry of Research, Sound Change</article-title>
          , PNAS
          <volume>110</volume>
          (
          <year>2013</year>
          )
          <fpage>4224</fpage>
          -
          <lpage>4229</lpage>
          . Innovation and Digitization, CNCS/CCCDI UEFIS- [13]
          <string-name>
            <surname>A. M. Ciobanu</surname>
            ,
            <given-names>L. P.</given-names>
          </string-name>
          <string-name>
            <surname>Dinu</surname>
          </string-name>
          , Automatic idenCDI,
          <source>project number 108/2021</source>
          ,
          <article-title>Romania. tification and production of related words for</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>historical linguistics</article-title>
          ,
          <source>Computational LinguisReferences tics 45</source>
          (
          <year>2019</year>
          )
          <fpage>667</fpage>
          -
          <lpage>704</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kondrak</surname>
          </string-name>
          ,
          <article-title>A new algorithm for the align</article-title>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Campbell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Historical</given-names>
            <surname>Linguistics</surname>
          </string-name>
          .
          <article-title>An Intro- ment of phonetic sequences</article-title>
          , in: Proceedings
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          duction, MIT Press,
          <year>1998</year>
          .
          <source>of ANLP</source>
          <year>2000</year>
          ,
          <year>2000</year>
          , pp.
          <fpage>288</fpage>
          -
          <lpage>295</lpage>
          . [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ritt</surname>
          </string-name>
          , Selfish Sounds and Linguistic Evo- [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado,
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Change</surname>
          </string-name>
          , Cambridge University Press,
          <year>2004</year>
          .
          <article-title>and phrases and their compositionality</article-title>
          , in: [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Epps</surname>
          </string-name>
          ,
          <article-title>Historical linguistics and socio-</article-title>
          <source>Proceedings of NIPS</source>
          <year>2013</year>
          ,
          <year>2013</year>
          , pp.
          <fpage>3111</fpage>
          <lpage />
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>cultural reconstruction</article-title>
          ,
          <source>in: The Routledge</source>
          <volume>3119</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>Handbook of Historical Linguistics</source>
          , London: [16]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Grave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , T. Mikolov,
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Routledge</surname>
          </string-name>
          ,
          <year>2014</year>
          , pp.
          <fpage>579</fpage>
          -
          <lpage>597</lpage>
          .
          <article-title>Enriching word vectors with subword informa[4]</article-title>
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Ciobanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          , Automatic detec- tion
          <source>, TACL</source>
          <volume>5</volume>
          (
          <year>2016</year>
          )
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>tion of cognates using orthographic alignment</article-title>
          , [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>in: Proceedings of ACL 2014</source>
          , Volume
          <volume>2</volume>
          ,
          <year>2014</year>
          , BERT:
          <article-title>Pre-training of deep bidirectional trans-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          pp.
          <fpage>99</fpage>
          -
          <lpage>105</lpage>
          .
          <article-title>formers for language understanding</article-title>
          , in: Pro[5]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Ciobanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          , Automatic discrim- ceedings
          <source>of NAACL</source>
          <year>2019</year>
          ,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <article-title>ination between cognates and borrowings</article-title>
          , in: [18]
          <string-name>
            <surname>A. M. Cristea</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Dinu</surname>
            ,
            <given-names>L. P.</given-names>
          </string-name>
          <string-name>
            <surname>Dinu</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>Proceedings of ACL</source>
          <year>2015</year>
          ,
          <year>2015</year>
          , pp.
          <fpage>431</fpage>
          -
          <lpage>437</lpage>
          . S. Georgescu,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Uban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zoicas</surname>
          </string-name>
          , Towards [6]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Cristea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Georgescu</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Mi- an Etymological Map of Romanian</article-title>
          , in: Pro-
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>hai</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          <string-name>
            <surname>Uban</surname>
          </string-name>
          ,
          <source>Automatic discrimination ceedings of RANLP</source>
          <year>2021</year>
          ,
          <year>2021</year>
          , pp.
          <fpage>315</fpage>
          -
          <lpage>323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <article-title>between inherited and borrowed latin</article-title>
          words in [19]
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Uban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          , Automatically building
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <year>2021</year>
          ,
          <year>2021</year>
          , pp.
          <fpage>2845</fpage>
          -
          <lpage>2855</lpage>
          . supervision,
          <source>in: Proceedings of LREC</source>
          <year>2020</year>
          , [7]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Ciobanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          , Ab initio: Au- 2020, pp.
          <fpage>3001</fpage>
          -
          <lpage>3007</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <source>Proceedings of COLING</source>
          <year>2018</year>
          ,
          <year>2018</year>
          , pp.
          <fpage>1604</fpage>
          -
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          1614. [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Spinelli</surname>
          </string-name>
          ,
          <article-title>The latin diachronic database: a</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          2022, forthcoming. [9]
          <string-name>
            <given-names>A.-S.</given-names>
            <surname>Uban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Ciobanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          , Cross-
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <article-title>tional approaches to semantic change 6 (</article-title>
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          219. [10]
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Uban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cristea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dinu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Dinu</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <year>2021</year>
          ,
          <year>2021</year>
          , pp.
          <fpage>64</fpage>
          -
          <lpage>74</lpage>
          . [11]
          <string-name>
            <given-names>D.</given-names>
            <surname>Poeppel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Embick</surname>
          </string-name>
          , Defining the relation
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Cornerstones</surname>
          </string-name>
          , New York, Routledge,
          <year>2017</year>
          , pp.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          103-
          <fpage>118</fpage>
          . [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bouchard-Côté</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hall</surname>
          </string-name>
          , T. L. Grifiths,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>