<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Uppsala, Sweden, March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>WarMemoirSampo: A Semantic Portal for War Veteran Interview Videos</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rafael Leal</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heikki Rantala</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mikko Koho</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Esko Ikkala</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minna Tamper</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Markus Merenmies</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eero Hyvönen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Helsinki Centre for Digital Humanities (HELDIG), University of Helsinki</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Archives of Finland</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Semantic Computing Research Group (SeCo), Aalto University</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>1</volume>
      <fpage>5</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>This paper presents WarMemoirSampo, a portal that provides semantic search and navigation of video interviews with Finnish World War II veterans. The portal associates video fragments with contextual data extracted from the video transcriptions, enabling users to find suitable video segments via faceted search and highlighting relevant content in the video being watched. This is carried out by processing natural language texts in order to extract named entities, keywords and lemmas. The result is a Linked Data Knowledge Graph that underpins the portal. We describe the collaboration between Natural Language Processing and Semantic Web technologies used in order to produce these results.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Linked Open Data</kwd>
        <kwd>Named entity recognition</kwd>
        <kwd>Named entity linking</kwd>
        <kwd>Military history</kwd>
        <kwd>War veterans</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>in order to produce a harmonized knowledge graph (KG) – including named entities, keywords
and lemmas – with timestamp information. This enables the videos to be accessed at diferent
points according to their semantic content, and additional contextual information to be provided
while the video is being watched. In order to achieve this, we created a KG containing enriched
data for the interviews as well as information about the interviewees. The data enrichment
and lemmatization is carried out using Natural Language Processing (NLP) techniques and
knowledge extraction, resulting in new metadata about keywords and mentioned named entities
– e.g., people, places, and events. These are integrated into the KG as entities and properties.
The data is then stored in a SPARQL endpoint for open access.</p>
      <p>The interview videos were edited by adding titles and copyright notes in the beginning
and by merging multi-part interviews. They were then published on the YouTube platform
for streaming, as well as archived in a repository at the National Archives for later use. The
WarMemoirSampo Portal2 was published on December 3, 2021, at the National Archives of
Finland.</p>
      <p>
        This system is related to previously published works of knowledge extraction from text to
linked data [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The idea of providing contextual information while watching videos has been
suggested already in the 80’s in systems such as Hypersoap3 that demonstrated the possibility
of interactive product placement in a broadcast setting. There are also existing linked
databased metadata models for videos and video segmentation [
        <xref ref-type="bibr" rid="ref3 ref4 ref6">3, 4, 6</xref>
        ]. However, there does not
seem to exist linked data-based systems for publishing and viewing videos while showing
time-segmented contextual metadata.
      </p>
      <p>The paper is organized as follows: first, the knowledge extraction and transformation pipeline
for the underlying data service is explained. After this, functionalities of the portal application
on top of the data service are briefly delineated. In conclusion, the main results of the work are
discussed and directions for future research outlined.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Text Processing and Data Enrichment</title>
      <p>The interviews were carried out in Finnish, which presents additional dificulties for natural
language processing in comparison to better-researched European languages: its rich
morphology4 complicates lemmatization and Named Entity Recognition (NER). Moreover, Finnish NLP
models are not as varied as for example those for English or French, and since the interviews
were realized in colloquial, dialectal Finnish, not even state-of-the-art speech-to-text algorithms
proved robust enough to handle them: at this point in time, only the standard variant of the
language (known as yleiskieli) leads to acceptable results in speech-to-text tasks.</p>
      <p>As an alternative, rough transcriptions made post-hoc by the interviewers were used for each
of the 159 interviews. These vary widely in length and level of detail, with an average character
count of 5599 ± 2458, which corresponds to a correlation of 0.718 with the interview lengths.
Due to their nature, these notes are ultimately not meant to be used as a source of information
2The portal is in use at: https://sotamuistot.arkisto.fi .
3www.media.mit.edu/hypersoap/
4Finnish contains around 15 cases and various sufixes, which result in a number of surface forms that may reach
well over 2000 for a single word. For example, the webpage http://www.ling.helsinki.fi/~fkarlsso/genkau2.html lists
2253 automatically generated forms for the word kauppa ’shop’
for others than the interviewers themselves: they do not provide full-fledged texts – or even
sentences necessarily – and contain abbreviations as well as orthographic and grammatical
errors, which compounds the challenge posed by the Finnish language. Nevertheless, we found
the notes mostly adequate for the purposes of this project, although we have not thoroughly
assessed their accuracy.</p>
      <p>The interviewers’ notes were provided as a spreadsheet document containing also other
types of metadata, such as interviewees’ name, date and place of birth, length and place of
the interviews, as well as links for the interview videos uploaded to the YouTube platform.
These notes are divided into rows, containing each one roughly a sentence. Each row also
has a timestamp related to their location in the YouTube video. However, these timestamps
are not unique: usually several consecutive rows have the same timestamp. These rows were
grouped together, resulting in video segments of variable length but typically several minutes
long, which became our data unit for this project.</p>
      <p>An overview of the subsequent data processing and transformation process is shown in
Figure 1. The original spreadsheet – containing one tab per interview, as well as metadata
for all of them – is initially split into multiple CSV files. Each interview is then divided into
shorter segments, as explained above. Three diferent procedures are carried out: the notes
related to these segments are lemmatized in order to facilitate searching and other tasks; their
keywords are generated; and their named entities are extracted and linked. The result of the
data transformation process5 are RDF files, which contain the KG that portrays all the diferent
entities found in the data, their relations and properties.</p>
      <p>
        The lemmatization of the texts is carried out by our Secompling library6, while keywords are
obtained via Annif [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], a subject indexing tool developed by the National Library of Finland.
An Annif pre-trained model is used, and from the results obtained all named entities and
keywords below a certain threshold are discarded. Three diferent NER/NEL (named entity
linking) systems are used in this project: Secompling-NER, which uses the TurkuNLP Finnish
NER tool [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to extract a large quantity of named entities but does not provide entity linking;
Nelli, which in this project is performing linking to Wikidata; and Warsa-linkers, which links
to entities in WarSampo. These tools compliment each other by extracting entities and linking
them to two diferent knowledge bases. The process and configurations for each tool will be
explained in the next sections, as well as the procedure adopted to harmonize them and the
overarching data model.
      </p>
      <p>
        Lemmatization and NER with Secompling. Secompling is a library that aims at
combining various Finnish NLP tools, including third-party ones, in order to provide an easy and
integrated interface for methods such as lemmatization, NER, relevance feedback and
unsupervised classification [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and keyword extraction. The module responsible for lemmatization
and NER relies on two tools developed by the TurkuNLP research group7: the Finnish NER
tool for named entity recognition and the Neural Parser pipeline [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for tokenization and
lemmatization [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Both are based on FinBERT [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], a deep learning Finnish language model
the TurkuNLP group has trained from scratch following the BERT guidelines [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Secompling
5WarMemoirSampo data transformation process: https://version.aalto.fi/gitlab/seco/veterans_manager
6https://version.aalto.fi/gitlab/seco/secompling
7https://turkunlp.org/
complements the Neural parser with the third-party tools Voikko8 and uralicNLP [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], which
are used to check and correct the lemmas based on the part-of-speech (POS) tags provided by
the parser itself. Some heuristics and automatic document editing are also applied in order
to fix errors commonly made by the parser, such as wrong tokenization of strings containing
punctuation marks, which may result in erroneous POS tags and lemmas.
      </p>
      <p>
        The results of the NER tool and the parser are then aligned. Since the former does not provide
lemmatization, this step provides the means for obtaining basic word forms for the named
entities. Moreover, this procedure helps to fix errors in both tools, such as combining tokens
that should be split, or assigning diferent named entity classes to parts of the same entity. The
results of a brief examination with 100 random entities indicate that the Neural parser achieved
a lemmatization accuracy of 0.84, while Secompling raised it to 0.96 by also applying named
entity-specific heuristics [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>Only a part of the eighteen entity categories recognized by the NER tool were judged to be
of interest in WarMemoirSampo: Person, Product, Organization, Event, and a combination of
Location (LOC), Facilities (FAC) and Geopolitical entities (GPE) as Place. Warsa-linkers adds
Military Units to this list.</p>
      <p>Since at this point Secompling-NER does not perform entity linking, it does not have the
means to tell an accurate named entity from an error in either recognition or lemmatization. As
a temporary solution, manual correction is applied to the candidates, so that it is possible to
remove an entity, or edit its category or lemma. As a result, a total of 336 entities were ignored
and around 270 had their category and/or lemma changed.</p>
      <p>
        NEL Using Nelli. The named entity linking tool Nelli [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ] is used to identify and link
entities to diferent knowledge bases. In WarMemoirSampo this tool is configured to use
the Turku neural parser pipeline to lemmatize the text prior to linking, and ARPA [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] to
identify and then link named entities to the given vocabularies. ARPA is a configurable entity
linking tool that consists of a list of linking configurations for diferent services; in the case of
WarMemoirSampo the configurations are used to link entities to Wikidata.
      </p>
      <p>In the interviews, the veterans mentioned some of their comrades in arms and superiors.
Correctly identifying the most notable oficers is often easy; disambiguating regular soldiers
is not. The latter were often mentioned only by first name and/or surname, so that more
information would have been necessary to diferentiate them. Hence, the linking strategy
centered around notable figures present in Wikidata.</p>
      <p>Regarding place linkage, several Wikidata ARPA configurations are used to link places such
as continents, countries, cities, as well as smaller places such as towns and villages. It can
be challenging to match former Finnish place names to Wikidata due to varying practices in
classifying place entities (e.g., towns or urban settlements) or missing labels, e.g., Enso or Viipuri.
Moreover, places that were annexed to Russia often changed their names from Finnish to Russian,
and some place names in Wikidata were lacking their original Finnish names mentioned in the
interviews.</p>
      <p>Warsa-linkers. WarSampo contains a data and ontology infrastructure for Finland in World
War II. The WarSampo KG consists of around 100 000 persons, 51 000 places, 16 000 military
units, 166 000 photographs, 26 000 war diaries, and 3000 war veteran memoir articles, among
others, with large amounts of links between entities.</p>
      <p>
        The data processing infrastructure of WarSampo contains a NEL process for linking
descriptions of events and photographs to persons, places, and military units mentioned in texts [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
The same process is reused in WarMemoirSampo to link the interviewers’ notes to persons,
places, and military units in WarSampo. After this NEL phase, the linked entities are pulled with
SPARQL queries from the WarSampo KG, and new entities are created in WarMemoirSampo
based on their metadata, with links to the original entities.
      </p>
      <p>Entity Reconciliation. The generated named entities from diferent steps are disambiguated
between the diferent data sources based on 1) matching the sets of URIs received for them from
LOD data sources, and 2) matching entity types and names. WarSampo and Wikidata often
contain alternative labels which are helpful in entity matching. In Wikidata, it is also possible to
manually add missing labels in order to improve entity matching. After matching, the entities
are reconciled by merging the duplicates found. The reconciled named entities consist of 310
persons, 1840 places, 66 events, 51 products, 117 military units, and 610 other organizations.
These include linked entities as well as those unlinked which were deemed important, which
also receive their own reference page (as explained in section 4).</p>
      <p>
        We evaluated the outcome of the entity recognition pipeline by inspecting twenty random
interview segments containing a total of 99 recognized entities [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]: 91 entities were correctly
identified, 2 entities were mistakenly linked, 2 mistakenly recognized, 4 wrongly categorized,
and 5 were not recognized. The final Precision is 0.919 and the Recall is 0.875, producing an F1
score of 0.897.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. WarMemoirSampo: Knowledge Graph</title>
      <p>The WarMemoirSampo KG, which is the result of the data transformation pipeline, is contained
in RDF files and served via a SPARQL endpoint 9. It is the only point of contact between the
system’s NLP backend and its user interface. It contains 323 371 triples, each one describing
an entity or linking it to another. The main classes are (with the : namespace referring to
http://ldf.fi/schema/warmemoirsampo/ ).</p>
      <p>• :Interview (159 instances), corresponding to one video interview;
• :PersonRecord (159 instances), which collect the interviewee’s personal information, such
as given name, family name, gender and date of birth;
• :TimeSlice (2417): a video segment from start timestamp to end timestamp. An :Interview
is divided into :TimeSlices, as explained in section 2;
• :NamedEntity (2994): Named Entity instances. Each named entity type also receives its
own subclass;
• skos:Concept (3127): keywords as provided by Annif. They are instances of another KG,
the General Finnish Ontology (YSO);</p>
      <p>Both Interviews and TimeSlices are indexed with named entities and keywords. The named
entities and keywords contain links to the same resources in other information sources.</p>
    </sec>
    <sec id="sec-4">
      <title>4. WarMemoirSampo Portal</title>
      <p>
        In order to provide an intuitive and user-friendly access to the WarMemoirSampo KG, a new
web-based user interface had to be created. To avoid starting from scratch, the Sampo-UI
JavaScript framework [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], which has been used for the creation of user interfaces for several
domain-specific semantic portals in recent years 10, was chosen as the basis.
      </p>
      <p>
        The structure of the WarMemoirSampo Portal11 is based on the three main classes in the
KG: :Interview, :TimeSlice, and :NamedEntity. Guided by principle 4. of the Sampo model [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
a separate faceted search perspective was created for each of these classes. Thus, instead of
providing the end-user with only a single text search field to start with, these perspectives ofer
three complementary ways to search and browse the KG:
1. The Interviews perspective is meant for faceted searching of full interview videos using
a combination of the following facets: text search, name and gender of the interviewee,
and a named entity (place, person, military unit, organization, event, or product).
2. The Parts of interviews perspective ofers similar search functionalities as the Interviews
perspective, but in this case the result set consists of interview segments (TimeSlices).
The search result links can be used to access the interview at the specific point that meets
the search criteria. It is possible to for example list all interview segments where the
best-known Finnish military leader, Carl Gustaf Emil Mannerheim, is mentioned.
9The WarMemoirSampo SPARQL endpoint is published at https://ldf.fi/warmemoirsampo/sparql
10See the full list of semantic portals powered by Sampo-UI at https://seco.cs.aalto.fi/applications/sampo
11The source code of the user interface is available on GitHub: https://github.com/SemanticComputing/
veterans-web-app
Related places on map
      </p>
      <p>Word cloud
Embedded YouTube</p>
      <p>video player
General interview</p>
      <p>metadata
All automatically</p>
      <p>recognized
named entities
3. The Index perspective includes faceted search and browsing functionalities for the nearly
3000 named entities mentioned in the summary notes of the interviews. For each named
entity, the search results include links to specific part(s) of the interview where this entity
is mentioned.</p>
      <p>Each search result in the aforementioned faceted search perspectives acts as link to an instance
page, which aggregates information about the particular instance (e.g. an interview or a place)
and ofers links to related instances in WarSampo and Wikidata. These pages ofer the end-user
the possibility to learn more about e.g. the military units, persons or locations mentioned.</p>
      <p>For the purposes of interactive viewing of interview videos, the Sampo-UI framework’s
instance page component was extended for embedded videos with dynamic table of contents
functionalities using the YouTube IFrame player API12. Figure 2 shows an instance page of an
interview. On the top left there is an embedded YouTube video player, while on the bottom left
there is metadata related to the interview. On the right there is a dynamic table of contents
(TOC) of the interview, which can be used to navigate to a specific segment of the interview
video. The time position of the YouTube video is synchronized with the TOC, so that the active
segment is expanded automatically. The expanded TOC part contains the interviewer notes as
well as the recognized named entities related to the current part of the interview, which are
also links to instance pages.</p>
      <p>12https://developers.google.com/youtube/iframe_api_reference</p>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>WarMemoirSampo is the outcome of a collaboration between Linked Open Data and NLP
techniques. Its backend capitalizes on various language-processing tools in order to handle texts
and extract information from them, while the result of this process, a linked data knowledge
graph, is able to store the information and swiftly deliver relevant parts to a faceted search-based
user interface. The resulting WarMemoirSampo Portal ofers the general public as well as
specialists a way to grasp the contents of the interviews through keywords and named entities
and the means to easily search and access relevant parts of the videos.</p>
      <p>The portal is also provided with a feedback button so that the developers can gain insights
based on user experience, since the public did not have access to it during development.
Nevertheless, more features are planned for the future: the links to the entities will be integrated
into the text instead of being listed separately. An event detection tool is to be developed which
extracts event entities based on times and places mentioned in the interviews. Moreover, when
Finnish speech-to-text technology advances to the point that everyday dialectal speech can be
reliably transcribed, the same tools could be used on these new transcriptions.</p>
      <p>Acknowledgements Ilpo Murtovaara provided the original transcripts of the videos and was
influential in filming the videos. Kare Salonvaara edited the videos. Tammenlehvän Perinneliitto
ry funded our work. CSC – IT Center for Science provided computational resources.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Hyvönen</surname>
          </string-name>
          ,
          <article-title>Digital humanities on the semantic web: Sampo model</article-title>
          and portal series, Semantic Web - Interoperability, Usability, Applicability (
          <year>2022</year>
          ). Submitted.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Koho</surname>
          </string-name>
          , E. Ikkala,
          <string-name>
            <given-names>P.</given-names>
            <surname>Leskinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          , E. Hyvönen,
          <article-title>Warsampo knowledge graph: Finland in the second world war as linked open data</article-title>
          ,
          <source>Semantic Web - Interoperability, Usability, Applicability</source>
          <volume>12</volume>
          (
          <year>2021</year>
          )
          <fpage>265</fpage>
          -
          <lpage>278</lpage>
          . doi:
          <volume>10</volume>
          .3233/SW-200392.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hunter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Iannella</surname>
          </string-name>
          ,
          <article-title>The application of metadata standards to video indexing</article-title>
          , in: C.
          <string-name>
            <surname>Nikolaou</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          Stephanidis (Eds.),
          <source>Research and Advanced Technology for Digital Libraries</source>
          , Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>1998</year>
          , pp.
          <fpage>135</fpage>
          -
          <lpage>156</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Mucheroni</surname>
          </string-name>
          ,
          <article-title>Dynamic indexation in video metadata</article-title>
          ,
          <source>Procedia-Social and Behavioral Sciences</source>
          <volume>73</volume>
          (
          <year>2013</year>
          )
          <fpage>551</fpage>
          -
          <lpage>555</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Martinez-Rodriguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hogan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Lopez-Arevalo, Information extraction meets the semantic web: A survey</article-title>
          ,
          <source>Semantic Web Journal</source>
          <volume>11</volume>
          (
          <year>2020</year>
          )
          <fpage>255</fpage>
          -
          <lpage>335</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H. Q.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pedrinaci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dietze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Domingue</surname>
          </string-name>
          ,
          <article-title>Using Linked Data to annotate and search educational video resources for supporting distance learning</article-title>
          ,
          <source>IEEE Transactions on Learning Technologies</source>
          <volume>5</volume>
          (
          <year>2012</year>
          )
          <fpage>130</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>O.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <article-title>Annif: DIY automated subject indexing using multiple algorithms</article-title>
          ,
          <source>LIBER Quarterly 29</source>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          . doi:
          <volume>10</volume>
          .18352/lq.10285.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Luoma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Oinonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pyykönen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Laippala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pyysalo</surname>
          </string-name>
          ,
          <article-title>A broad-coverage corpus for Finnish named entity recognition</article-title>
          ,
          <source>in: Proceedings of the 12th Language Resources and Evaluation Conference</source>
          , European Language Resources Association, Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>4615</fpage>
          -
          <lpage>4624</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Leal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kesäniemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koho</surname>
          </string-name>
          , E. Hyvönen,
          <source>Relevance Feedback Search Based on Automatic Annotation and Classification of Texts</source>
          , in: D.
          <string-name>
            <surname>Gromann</surname>
            , G. Sérasset,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gracia</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Bosque-Gil</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Bobillo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          Heinisch (Eds.),
          <source>3rd Conference on Language, Data and Knowledge (LDK</source>
          <year>2021</year>
          ), Schloss Dagstuhl - Leibniz-Zentrum für Informatik, Dagstuhl, Germany,
          <year>2021</year>
          , pp.
          <volume>18</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          :
          <fpage>15</fpage>
          . doi:
          <volume>10</volume>
          .4230/OASIcs.LDK.
          <year>2021</year>
          .
          <volume>18</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kanerva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ginter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Miekka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Leino</surname>
          </string-name>
          , T. Salakoski,
          <article-title>Turku neural parser pipeline: An end-to-end system for the CoNLL 2018 shared task</article-title>
          ,
          <source>in: Proceedings of the CoNLL</source>
          <year>2018</year>
          <article-title>Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, Association for Computational Linguistics</article-title>
          , Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>133</fpage>
          -
          <lpage>142</lpage>
          . URL: https://aclanthology.org/K18-2013. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>K18</fpage>
          -2013.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kanerva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ginter</surname>
          </string-name>
          , T. Salakoski,
          <article-title>Universal lemmatizer: A sequence to sequence model for lemmatizing universal dependencies treebanks</article-title>
          ,
          <source>Natural Language Engineering</source>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          . doi:
          <volume>10</volume>
          .1017/S1351324920000224.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Virtanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kanerva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ilo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Luoma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Luotolahti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Salakoski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ginter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pyysalo</surname>
          </string-name>
          , Multilingual is not enough: Bert for Finnish,
          <year>2019</year>
          . arXiv:
          <year>1912</year>
          .07076.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Vol.
          <volume>1</volume>
          , Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>N19</fpage>
          -1423.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hämäläinen</surname>
          </string-name>
          ,
          <article-title>UralicNLP: An NLP library for Uralic languages</article-title>
          ,
          <source>Journal of Open Source Software</source>
          <volume>4</volume>
          (
          <year>2019</year>
          )
          <article-title>1345</article-title>
          . doi:
          <volume>10</volume>
          .21105/joss.01345.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Koho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Leal</surname>
          </string-name>
          , E. Ikkala,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rantala</surname>
          </string-name>
          , E. Hyvönen,
          <article-title>Building lightweight ontologies for faceted search with named entity recognition: Case WarMemoirSampo</article-title>
          , in: International Workshop on
          <article-title>Knowledge Graph Generation from Text (TEXT2KG</article-title>
          <year>2022</year>
          ), Proceedings, Springer,
          <year>2022</year>
          . Forthcoming.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oksanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hietanen</surname>
          </string-name>
          , E. Hyvönen,
          <article-title>Automatic Annotation Service APPI: Named Entity Linking in Legal Domain</article-title>
          ,
          <source>in: The Semantic Web: ESWC 2020 Satellite Events</source>
          , Springer-Verlag,
          <year>2020</year>
          , pp.
          <fpage>110</fpage>
          -
          <lpage>114</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamper</surname>
          </string-name>
          , E. Hyvönen,
          <string-name>
            <given-names>P.</given-names>
            <surname>Leskinen</surname>
          </string-name>
          ,
          <article-title>Visualizing and analyzing networks of named entities in biographical dictionaries for digital humanities research</article-title>
          ,
          <source>in: Proceedings of the 20th International Conference on Computational Linguistics and Intelligent Text Processing (CICling</source>
          <year>2019</year>
          ), Springer-Verlag,
          <year>2019</year>
          . Forthcoming.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>E.</given-names>
            <surname>Mäkelä</surname>
          </string-name>
          ,
          <article-title>Combining a REST Lexical Analysis Web Service with SPARQL for Mashup Semantic Annotation from Text</article-title>
          ,
          <source>in: The Semantic Web: ESWC 2014 Satellite Events</source>
          , Springer International Publishing,
          <year>2014</year>
          , pp.
          <fpage>424</fpage>
          -
          <lpage>428</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -11955-7_
          <fpage>60</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>E.</given-names>
            <surname>Heino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tamper</surname>
          </string-name>
          , E. Mäkelä,
          <string-name>
            <given-names>P.</given-names>
            <surname>Leskinen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Ikkala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tuominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koho</surname>
          </string-name>
          , E. Hyvönen,
          <article-title>Named entity linking in a complex domain: Case second world war history</article-title>
          , in: J.
          <string-name>
            <surname>Gracia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Bond</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Chiarcos</surname>
          </string-name>
          , S. Hellmann (Eds.), Language, Data, and
          <string-name>
            <surname>Knowledge</surname>
          </string-name>
          (LDK
          <year>2017</year>
          ), Springer International Publishing, Cham,
          <year>2017</year>
          , pp.
          <fpage>120</fpage>
          -
          <lpage>133</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -59888-8_
          <fpage>10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ikkala</surname>
          </string-name>
          , E. Hyvönen,
          <string-name>
            <given-names>H.</given-names>
            <surname>Rantala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Koho</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sampo-UI</surname>
          </string-name>
          :
          <article-title>A full stack JavaScript framework for developing semantic portal user interfaces</article-title>
          ,
          <source>Semantic Web - Interoperability, Usability, Applicability</source>
          <volume>13</volume>
          (
          <year>2022</year>
          )
          <fpage>69</fpage>
          -
          <lpage>84</lpage>
          . doi:
          <volume>10</volume>
          .3233/SW-210428.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>