<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Digital Geolinguistics: On the use of Linked Open Data for Data-Level Interoperability Between Geolinguistic Resources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giorgio Maria Di Nunzio</string-name>
          <email>dinunzio@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Information Engineering</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Padua</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <abstract>
        <p>The Open Language Archives Community which recently celebrated its rst 10 years of activity, is a worldwide network dedicated to collecting information on language resources and developing standard protocols for interoperability. In this context, Linked Open Data paradigm is very promising, because it eases interoperability between di erent systems by allowing the de nition of data-driven models and applications. In this talk, we give an overview of present geolinguistics projects and an approach which moves the focus from the systems handling the linguistic data to the data themselves. As a concrete example, we present a geolinguistic application build upon a real linguistic dataset which provides linguists with a system for investigating variations among closely related languages.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The research eld of linguistics studies all aspects of human language, including
morphology (the formation and composition of words), syntax (the formation
and composition of phrases and sentences from these words) and phonology
(sound systems) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Research in the variations in languages allows linguists
to understand the fundamental principles that underlie language di erences,
language innovation and language variation in time and space.
      </p>
      <p>
        Geolinguistics is an interdisciplinary eld that incorporates language maps
depicting spatial patterns of language location or the results of processes that
lead to language change [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In this context, the linguistic atlas has proved to
be a vital tool and product of geolinguistics since the earliest stages of the eld,
and it has provided a stage for the incorporation of modern GIS.
      </p>
      <p>
        In the last two decades, several large-scale databases of linguistic material
of various types have been developed worldwide. The Open Language Archives
Community,1 which recently celebrated its rst 10 years of activity, is a
worldwide network dedicated to collecting information on language resources ( eld
notes, grammars, audio/video recording, descriptive papers, and so on) and
developing standard protocols for interoperability. GOLD2 was the rst ontology
to be designed speci cally for linguistic description on the Semantic Web [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. It
proposes a solution to the lack of interoperability between linguistic projects and
projects designed speci cally for NLP applications. It can act as a kind of lingua
franca for the linguistic data community, provided that data providers are
willing to map their data to GOLD or to some similar resource. In [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], the authors
present a framework for producing multi-layer annotated corpora: a pivot format
serving as \interlingua" between annotation tools, an ontology-based approach
for mapping between tag sets, and an information system that integrates the
various annotations and allows for querying the data either by posing simple
queries or by using the ontology.
      </p>
      <p>
        Language resources that have been made publicly available can vary in the
richness of the information they contain: on the one hand, a corpus typically
contains at least a sequence of words, sounds or tags; on the other hand, a
corpus may contain a large amount of information about the syntactic structure,
morphology, prosody and semantic content of every sentence, plus annotations
of discourse relations or dialogue acts [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. However, the quality of such corpora
may have been reduced by the intense, and often poorly controlled, usage of
automatic learning algorithms [
        <xref ref-type="bibr" rid="ref13 ref6">6</xref>
        ].
      </p>
      <p>
        The heterogeneity of linguistic projects has been recognized as a key problem
limiting the reusability of linguistic tools and data collections [
        <xref ref-type="bibr" rid="ref14 ref7">7</xref>
        ]. The rate of
re-use for linguistic database technology together with related processing tools
and environments is still too low. For example, the Edisyn search engine { the
aim of which was to make di erent dialectal databases comparable { \in practice
has proven to be unfeaseable".3 In order to nd common ground where
linguistic material can be shared and re-used, the methodological and technological
boundaries existing in each research linguistic project needs to be overcome.
2
      </p>
      <p>
        Linked Open Data for Geolinguistic Resources
The research direction we want to discuss in this talk is to move the focus from
the systems handling the linguistic data to the data themselves. For this
purpose the LOD paradigm [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is very promising, because it eases interoperability
between di erent systems by allowing the de nition of data-driven models and
applications. LOD is based on the de nition of real-world objects, identi ed by
means of a dereferenceable URI4. Objects are related to one another by means
of typed links. Interoperability is achieved by a unifying data model (i.e. RDF5),
a standardized data access mechanism (i.e. HTTP), hyperlink-based data
discovery (i.e. URI), and self-descriptive data (based on shared open vocabularies
from di erent namespace) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
2 http://linguistics-ontology.org/
3 http://www.dialectsyntax.org/
4 http://tools.ietf.org/html/rfc3986
5 http://www.w3.org/RDF/
In this context, a relevant initiative is ISOcat,6 a linguistic concept database
developed by ISO Technical Committee 37, Terminology and other language
and content resources, to provide reference semantics for annotation schemata.
The goal of the project is to create a universally available resource for
languagerelated metadata that can be used in a variety of applications and environments.
It also provides uniform naming and semantic principles to facilitate the
interoperability of language resources across applications and approaches [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>In this talk, we discuss the steps of a possible approach for exposing
geolingistic data into LOD [10{12] by presenting:
{ the ASIt7 linguistic project which is based on micro-variations of
Italo</p>
      <p>Romance dialects;
{ a geolinguistic Web application that provides functionalities for accessing,
browsing, searching the linked open data by means of linguistic features, and
visualizing the data on dynamically generated maps.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Akmajian</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demers</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farmer</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harnish</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Linguistics</surname>
          </string-name>
          ,
          <article-title>Sixth Edition - An Introduction to Language and Communication</article-title>
          . The MIT Press (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hoch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          :
          <source>Geolinguistics: The Incorporation of Geographic Information Systems and Science. The Geographical Bulletin</source>
          <volume>51</volume>
          (
          <issue>1</issue>
          ) (
          <year>2010</year>
          )
          <volume>23</volume>
          {
          <fpage>36</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Farrar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langendoen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>A Linguistic Ontology for the Semantic Web</article-title>
          .
          <source>Glot International</source>
          <volume>7</volume>
          (
          <issue>3</issue>
          ) (
          <year>March 2003</year>
          )
          <volume>97</volume>
          {
          <fpage>100</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chiarcos</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dipper</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Gotze,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Leser</surname>
          </string-name>
          ,
          <string-name>
            <surname>U.</surname>
          </string-name>
          , Ludeling,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Ritz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Stede</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>A Flexible Framework for Integrating Annotations from Di erent Tools and Tag Sets</article-title>
          .
          <source>TAL</source>
          <volume>49</volume>
          (
          <issue>2</issue>
          ) (
          <year>2008</year>
          )
          <volume>217</volume>
          {
          <fpage>246</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loper</surname>
            ,
            <given-names>E.: Natural</given-names>
          </string-name>
          <string-name>
            <surname>Language Processing with Python. O'Reilly Media</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Sparck Jones,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Computational linguistics: What about the linguistics?</article-title>
          <source>Computational Linguistics</source>
          <volume>33</volume>
          (
          <issue>3</issue>
          ) (
          <year>2007</year>
          )
          <volume>437</volume>
          {
          <fpage>441</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Chiarcos</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Interoperability of corpora and annotations</article-title>
          . In Chiarcos,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Nordho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Hellmann</surname>
          </string-name>
          , S., eds.: Linked Data in Linguistics. Springer Berlin Heidelberg (
          <year>2012</year>
          )
          <volume>161</volume>
          {
          <fpage>179</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Linked Data: Evolving the Web into a Global Data Space</article-title>
          .
          <article-title>Synthesis Lectures on the Semantic Web</article-title>
          . Morgan &amp; Claypool Publishers (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kemps-Snijders</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Windhouwer</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wittenburg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wright</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>ISOcat: Remodelling Metadata for Language Resources</article-title>
          .
          <source>IJMSO</source>
          <volume>4</volume>
          (
          <issue>4</issue>
          ) (
          <year>2009</year>
          )
          <volume>261</volume>
          {
          <fpage>276</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>Di</given-names>
            <surname>Buccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.M.</given-names>
            ,
            <surname>Silvello</surname>
          </string-name>
          , G.:
          <article-title>A system for exposing linguistic linked open data</article-title>
          . In Zaphiris,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Buchanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Rasmussen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Loizides</surname>
          </string-name>
          , F., eds.
          <source>: TPDL</source>
          . Volume
          <volume>7489</volume>
          of Lecture Notes in Computer Science., Springer (
          <year>2012</year>
          )
          <volume>173</volume>
          {
          <fpage>178</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Di</given-names>
            <surname>Buccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.M.</given-names>
            ,
            <surname>Silvello</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          :
          <article-title>An open source system architecture for digital geolinguistic linked open data</article-title>
          . In Aalberg, T.,
          <string-name>
            <surname>Papatheodorou</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dobreva</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsakonas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farrugia</surname>
          </string-name>
          , C.J., eds.
          <source>: TPDL</source>
          . Volume
          <volume>8092</volume>
          of Lecture Notes in Computer Science., Springer (
          <year>2013</year>
          )
          <volume>438</volume>
          {
          <fpage>441</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>Di</given-names>
            <surname>Buccio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Di Nunzio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.M.</given-names>
            ,
            <surname>Silvello</surname>
          </string-name>
          , G.:
          <article-title>A curated and evolving linguistic linked dataset</article-title>
          .
          <source>Semantic Web</source>
          <volume>4</volume>
          (
          <issue>3</issue>
          ) (
          <year>2013</year>
          )
          <volume>265</volume>
          {
          <fpage>270</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>6 http://www.isocat.org/</mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>7 http://svrims2.dei.unipd.it:8080/asit-enterprise/</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>