<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>M. Tamper)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Extending the Finnish Linked Data Infrastructure with Natural Language Processing Services in FIN-CLARIAH</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Minna Tamper</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jouni Tuominen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eero Hyvönen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Helsinki Centre for Digital Humanities (HELDIG), University of Helsinki</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Helsinki Institute for Social Sciences and Humanities (HSSH), University of Helsinki</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Semantic Computing Research Group (SeCo), Aalto University</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The DARIAH-EU infrastructure for Digital Humanities (DH) is often focusing on using structured data for quantitative studies, while the EU-CLARIN infrastructure deals primarily with unstructured natural language texts. However, in DH research both texts and structured data are often needed. It therefore makes sense to develop and use both infrastructures together, as suggested in the Dutch CLARIAH programme and the corresponding FIN-CLARIAH initiative in Finland, a new part of the Finnish research infrastructure road map of the Academy of Finland. This poster paper introduces work in FINCLARIAH relating to the idea of integrating natural language processing (NLP) tools with the Linked Open Data (LOD) Infrastructure for Digital Humanities in Finland (LODI4DH). We present a plan for NLP services to be opened as part of the Linked Data Finland (LDF.fi) platform. The new services are used for knowledge extraction from Finnish texts for weaving LOD, and on the other hand for language DH data analyses of the published datasets in applications in many domains, such as political culture. The extended LDF.fi platform will provide users with documented APIs for NLP services using unified output formats as well as software delivery as Docker containers, to lower the bar for deployment.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;knowledge extraction</kwd>
        <kwd>natural language processing</kwd>
        <kwd>linked data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>FIN-CLARIAH (2022–) is the premier Finnish digital research infrastructure development
initiative for Social Sciences and Humanities (SSH) comprising two components:
1. FIN-CLARIN: Finnish dimension of the pan-European CLARIN1 infrastructure
2. DARIAH-FI: Finnish dimension of the pan-European DARIAH2 infrastructure
1. Reach beyond processing of spoken standard Finnish into colloquial speech
2. Cater to a broad range of SSH research needs for processing unstructured text
3. Facilitate research based on metadata</p>
      <p>FIN-CLARIAH involves Finnish universities with research in SSH, including the coordinator
University of Helsinki (Faculty of Arts, Faculty of Social Sciences, and National Library), CSC –
IT Center for Science Ltd., Aalto University, Tampere University, University of Eastern Finland,
University of Jyväskylä, and University of Turku. In addition, FIN-CLARIAH has as project
collaborators University of Vaasa, University of Oulu, the Institute for the Languages of Finland,
and the National Archives of Finland.</p>
      <p>
        This paper introduces development work of the Semantic Computing Research Group (SeCo)
in FIN-CLARIAH at the Aalto University and University of Helsinki, Helsinki Centre for Digital
Humanities (HELDIG)3, on integrating existing natural language processing (NLP) tools with
the Linked Open Data (LOD) Infrastructure for Digital Humanities in Finland (LODI4DH)4 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
In conclusion, related works are discussed and contributions of our work summarized.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Services for Weaving Linked Data from Texts</title>
      <p>
        The NLP services to be opened to public use are targeted to knowledge extraction from Finnish
texts, based on named entity recognition (NER), linking, and relation extraction [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The
same NLP technology can also be used for pseudonymizing texts [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] needed due to the GDPR
regulations5 of EU.
      </p>
      <p>
        The idea in our work is to reuse various existing NLP tools from other developers and
repurpose and re-package them for extracting Linked Data from unstructured texts. For example,
for NER a pretrained FinBERT NER model6 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is used. The rule-based methods to be used
include regular expressions (to find prescribed surface forms, e.g., property codes and vehicle
registration plates) and dictionaries, such as the Finnish person name ontology [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In addition,
the Turku Neural Parser [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is used to perform grammatical and morphological analysis on
the whole document. NLP tools developed in our own earlier projects for the various Sampo
portals [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] will also be re-used here.
      </p>
      <p>
        The NLP tools will be used in FIN-CLARIAH for language analyses of texts and for DH
analyses regarding their subject matter content. Application case study areas here include
analysing the nearly million speeches 1907—2021 of the Parliament of Finland7, biographical
collections, such as the National Biography of Finland8, and Finnish legislation and case law9
published by the Ministry of Justice. Texts for these application areas are already available
through the Sampo series of LOD services and portals10 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] but also other datasets from the
participants of FIN-CLARIAH will be used on as-needed basis.
      </p>
      <p>3https://seco.cs.aalto.fi/projects/fin-clariah/
4https://seco.cs.aalto.fi/projects/lodi4dh/
5https://europa.eu/youreurope/business/dealing-with-customers/data-protection/data-protection-gdpr/
6https://turkunlp.org/fin-ner.html
7https://seco.cs.aalto.fi/projects/semparl/
8https://seco.cs.aalto.fi/projects/biografiasampo/
9https://seco.cs.aalto.fi/projects/lakisampo/
10https://seco.cs.aalto.fi/applications/sampo/</p>
      <p>
        The NLP services will be provided using the existing Linked Data Finland platform11 that is
in use in Finland for publishing Linked Data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The forth-coming NLP services provide users
with demo applications and web APIs, unified output formats (e.g., JSON, RDF), documentation,
and software delivery as Docker container images, which lowers the bar for deployment. The
containerized tools can be deployed as components in data processing workflows, and be scaled
to support varying workloads. In addition to using and providing tools as containers, the portal
enables testing the tools with custom input.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Discussion</title>
      <p>
        Portals for NLP web services have been created before for many languages, e.g., the GATE
Cloud [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In addition to web services, this kind of tools have also been packaged at into
useful modules available for programming languages, such as Python, e.g., EstNLTK12 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
UralicNLP13 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], StanfordNLP14 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], and NLTK15 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. However, the modules or packages for
specific programming languages not only tie their users to the language but also require more
technical skills and understanding. A benefit of service portals is that it can provide users with
demo UIs and APIs through which the users can use the services more easily using the HTTP
protocol and regardless of the programming language they use.
      </p>
      <p>
        Methods for information extraction are surveyed in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A novelty of our work is its focus
on Finnish and on knowledge extraction from texts for linked data to be used in research for
language analyses and studies in Digital Humanities. The tools and services planned to be
included in the portal come from various NLP and DH projects that can be utilized to extract
knowledge from diferent types of source texts. The services are distributed as containers and
their results are converted into well-known output formats to ease deployment and improve
usability in other DH projects.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>Our work is funded by the Academy of Finland as part of the FIN-CLARIAH program for
national research infrastructures. CSC – IT Center for Science provides computational resources
our project.
11https://ldf.fi
12https://github.com/estnltk/estnltk
13https://github.com/mikahama/uralicNLP
14https://github.com/stanfordnlp/stanfordnlp
15https://github.com/nltk/nltk</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hyvönen</surname>
          </string-name>
          ,
          <article-title>Linked open data infrastructure for Digital Humanities in Finland</article-title>
          ,
          <source>in: Proceedings of Digital Humanities in Nordic Countries (DHN</source>
          <year>2020</year>
          ),
          <article-title>CEUR-WS Proceedings</article-title>
          , Vol.
          <volume>2612</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>254</fpage>
          -
          <lpage>259</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2612</volume>
          /short10.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Martinez-Rodriguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Lopez-Arevalo, Information extraction meets the semantic web: A survey</article-title>
          ,
          <source>Semantic Web - Interoperability, Usability, Applicability</source>
          <volume>11</volume>
          (
          <year>2020</year>
          )
          <fpage>255</fpage>
          -
          <lpage>335</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Oksanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hietanen</surname>
          </string-name>
          , E. Hyvönen,
          <article-title>Anoppi: A pseudonymization service for Finnish court documents</article-title>
          ,
          <source>in: Legal Knowledge and Information Systems. JURIX</source>
          <year>2019</year>
          :
          <article-title>The Thirty-second Annual Conference</article-title>
          , IOS Press,
          <year>2019</year>
          , pp.
          <fpage>251</fpage>
          -
          <lpage>254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Luoma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Oinonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pyykönen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Laippala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pyysalo</surname>
          </string-name>
          ,
          <article-title>A broad-coverage corpus for Finnish named entity recognition</article-title>
          ,
          <source>in: Proceedings of the 12th Language Resources and Evaluation Conference</source>
          , European Language Resources Association, Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>4615</fpage>
          -
          <lpage>4624</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>567</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Leskinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          , E. Hyvönen,
          <article-title>Modeling and publishing finnish person names as a linked open data ontology</article-title>
          ,
          <source>in: 3rd Workshop on Humanities in the Semantic Web (WHiSe</source>
          <year>2020</year>
          ),
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>2695</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>14</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2695</volume>
          /paper1.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kanerva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ginter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Miekka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Leino</surname>
          </string-name>
          , T. Salakoski,
          <article-title>Turku neural parser pipeline: An end-to-end system for the CoNLL 2018 shared task</article-title>
          ,
          <source>in: Proceedings of the CoNLL</source>
          <year>2018</year>
          <article-title>Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, Association for Computational Linguistics</article-title>
          ,
          <year>2018</year>
          , pp.
          <fpage>133</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hyvönen</surname>
          </string-name>
          ,
          <article-title>Digital humanities on the Semantic Web: Sampo model</article-title>
          and portal series, Semantic Web - Interoperability, Usability, Applicability (
          <year>2021</year>
          ). URL: http://www.semantic
          <article-title>-web-journal.net/content/ digital-humanities-semantic-web-sampo-</article-title>
          <string-name>
            <surname>model-</surname>
          </string-name>
          and
          <string-name>
            <surname>-</surname>
          </string-name>
          portal-series, submitted.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hyvönen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alonen</surname>
          </string-name>
          , E. Mäkelä,
          <article-title>Linked Data Finland: A 7-star model and platform for publishing and re-using linked datasets, in: The Semantic Web: ESWC 2014 Satellite Events</article-title>
          ,
          <source>Revised Selected Papers</source>
          , Springer,
          <year>2014</year>
          , pp.
          <fpage>226</fpage>
          -
          <lpage>230</lpage>
          . URL: https: //doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -11955-7_
          <fpage>24</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>V.</given-names>
            <surname>Tablan</surname>
          </string-name>
          , I. Roberts,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          , Gatecloud. net:
          <article-title>a platform for largescale, open-source text processing on the cloud</article-title>
          ,
          <source>Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences</source>
          <volume>371</volume>
          (
          <year>2013</year>
          )
          <fpage>20120071</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Laur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Orasmaa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Särg</surname>
          </string-name>
          , P. Tammo, EstNLTK
          <volume>1</volume>
          .
          <article-title>6: Remastered Estonian NLP pipeline</article-title>
          ,
          <source>in: Proceedings of the 12th Language Resources and Evaluation Conference</source>
          , European Language Resources Association, Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>7152</fpage>
          -
          <lpage>7160</lpage>
          . URL: https: //aclanthology.org/
          <year>2020</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>884</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hämäläinen</surname>
          </string-name>
          ,
          <article-title>UralicNLP: An NLP library for Uralic languages</article-title>
          ,
          <source>Journal of Open Source Software</source>
          <volume>4</volume>
          (
          <year>2019</year>
          )
          <article-title>1345</article-title>
          . doi:
          <volume>10</volume>
          .21105/joss.01345.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dozat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , C. D. Manning,
          <article-title>Universal dependency parsing from scratch, in: Proceedings of the CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, Association for Computational Linguistics</article-title>
          , Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>160</fpage>
          -
          <lpage>170</lpage>
          . URL: https://nlp.stanford.edu/pubs/qi2018universal.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Klein</surname>
          </string-name>
          , E. Loper,
          <article-title>Natural language processing with Python: analyzing text with the natural language toolkit,</article-title>
          <string-name>
            <surname>O'Reilly Media</surname>
          </string-name>
          , Inc.,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>