<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Large-scale Data Harvesting for Biographical Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alistair Plum⇤</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcos Zampieri⇤</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Constantin Or a˘san⇤</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eveline Wandl-Vogt</string-name>
          <email>eveline.wandl-vogt@oeaw.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ruslan Mitkov⇤</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Austrian Centre for Digital Humanities Austrian Academy of Sciences</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper explores automatic methods to identify relevant biography candidates in large databases, and extract biographical information from encyclopedia entries and databases. In this work, relevant candidates are defined as people who have made an impact in a certain country or region within a pre-defined time frame. We investigate the case of people who had an impact in the Republic of Austria and died between 1951 and 2019. We use Wikipedia and Wikidata as data sources and compare the performance of our information extraction methods on these two databases. We demonstrate the usefulness of a natural language processing pipeline to identify suitable biography candidates and, in a second stage, extract relevant information about them. Even though they are considered by many as an identical resource, our results show that the data from Wikipedia and Wikidata differs in some cases and they can be used in a complementary way providing more data for the compilation of biographies.</p>
      </abstract>
      <kwd-group>
        <kwd>Austria</kwd>
        <kwd>Biographies</kwd>
        <kwd>Information Extraction</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>Wikidata</kwd>
        <kwd>Wikipedia</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In the last decade, large biographical databases have
become available in a number of languages
        <xref ref-type="bibr" rid="ref12 ref14 ref2 ref21 ref24 ref4">(Reinert and
Ebneth, 2017)</xref>
        . This is the case with many online data
sources and sources which were digitised such as the
Slovenian Biography
        <xref ref-type="bibr" rid="ref8">(Erjavec et al., 2015)</xref>
        , the Deutsche
Biographie
        <xref ref-type="bibr" rid="ref22">(Reinert et al., 2015)</xref>
        , and the O¨ sterreichisches
Biographisches Lexikon
        <xref ref-type="bibr" rid="ref6">(der Wissenschaften, 2012)</xref>
        . As
a result of the large amounts of data available, researchers
have been exploring ways to process this data using
computational methods. In particular, natural language processing
(NLP) and information extraction (IE) methods play an
important role in processing these large amounts of data,
ranging from tasks like tokenisation, part of speech tagging and
sentence splitting, to toponym resolution, semantic
rolelabeling and relation extraction. Due to their size and
availability, Wikipedia1 and Wikidata2 have become popular
online data sources of information for biographies
        <xref ref-type="bibr" rid="ref1 ref4">(Biadsy
et al., 2008; Chisholm et al., 2017)</xref>
        . In addition,
DBpedia3 provides structured information that have been used to
generate biography summaries using natural language
generation methods
        <xref ref-type="bibr" rid="ref18">(Moussallem et al., 2018)</xref>
        . A number of
projects such as the A Prosopographical Information
System (APIS) project at the Austrian Academy of Sciences
(AAS)
        <xref ref-type="bibr" rid="ref12 ref14 ref2 ref21 ref24 ref4">(Schlo¨ gl and Lejtovicz, 2017)</xref>
        and the Dutch
BiographyNet project
        <xref ref-type="bibr" rid="ref9">(Fokkens et al., 2014)</xref>
        have addressed the
problem of retrieval of information from biographical
encyclopedias and dictionaries.
      </p>
    </sec>
    <sec id="sec-2">
      <title>1https://www.wikipedia.org/</title>
    </sec>
    <sec id="sec-3">
      <title>2https://www.wikidata.org/wiki/Wikidata:</title>
      <p>Main_Page provides structured information taken from
Wikipedia (and other sources)</p>
    </sec>
    <sec id="sec-4">
      <title>3https://wiki.dbpedia.org/</title>
      <p>
        The aforementioned APIS project aims to develop new
methods for re-using qualitative (biographical) research
products (encyclopedias) for quantitative research and, in
doing so, facilitate a digital transformation process against
the background of Humanities
        <xref ref-type="bibr" rid="ref12 ref14 ref2 ref21 ref24 ref4">(Gruber and Wandl-Vogt,
2017)</xref>
        . To achieve this, the project has developed a
webbased, customisable virtual research environment that
allows researchers to work with programs especially
designed for processing biographical. texts4 Another goal
of the APIS project is to reveal information encoded in
texts such as people names, institutions, places, and to
detect relationships between them and the person depicted in
the biography, primarily in relation to the O¨ sterreichisches
Biographisches Lexikon (O¨ BL) project, by which APIS is
third-party funded. In this context, the aim is to collect and
make visible the lives and careers of persons with impact
in the area of the former Austrian-Hungarian monarchy, as
well as the first Republic of Austria. In order to find
relevant candidates, O¨ BL is aimed at looking beyond the
primarily usual suspects, in order to find lesser known and
less easy to find knowledge carriers, influencers and impact
holders. Currently, about 18,500 biographies are available
and O¨BL aims to publish the final volumes in 2020. Since
2004, O¨ BL went digital and a database has been established
to support the manifold editorial processes. A rich network
to neighbouring analogue endeavours and close personal
relationships exist, for instance the European
BiographyPortal
        <xref ref-type="bibr" rid="ref12 ref14 ref2 ref21 ref24 ref4">(Gruber and Wandl-Vogt, 2017)</xref>
        .5
In this paper, we present a work-in-progress processing
pipeline which can be used to identify relevant biography
candidates in Wikidata and Wikipedia, and to extract
infor
      </p>
    </sec>
    <sec id="sec-5">
      <title>4https://apis.acdh.oeaw.ac.at/</title>
    </sec>
    <sec id="sec-6">
      <title>5https://www.biographie-portal.eu/</title>
      <p>mation about these candidates. The work presented here is
within the scope of the aforementioned APIS project and
aimed at enriching the O¨ BL, by means of automatically
suggesting candidates for inclusion in the newer additions
of the O¨ BL. We investigate the case of biography
candidates that died between 1951 and 2019 and who had an
impact in the present Republic of Austria. The current work
focuses on processing the English texts, which can easily
be adapted to process other languages as well, and
selecting relevant candidates for inclusion in the O¨ BL. However,
since previous work on adding biographies to the O¨BL was
carried out by a team of historians, with decisions
sometimes being based on factors which are difficult to model,
evaluation of our pipeline is not straight-forward. We
discuss this aspect in Section 4.</p>
      <p>2.</p>
      <sec id="sec-6-1">
        <title>Related Work</title>
        <p>
          The field of information extraction (IE) has long been a
thriving area of research within natural language
processing (NLP) and has steadily maintained a close relationship
with gathering information from the web. A large amount
of the early work on IE was done during the Message
Understanding Conferences (MUC) which took place from
1987 until 1997, and was later continued at the Automatic
Content Extraction (ACE) program
          <xref ref-type="bibr" rid="ref13 ref16 ref7">(Grishman and
Sundheim, 1996; Hirschman, 1998; Doddington et al., 2004)</xref>
          .
These well funded conferences laid the ground work for
extracting entities and information pertaining to the same
from text, audio and image data. Although the earliest
approaches were rule-based
          <xref ref-type="bibr" rid="ref3">(Chinchor et al., 1993)</xref>
          , machine
learning-based contributions quickly followed suit. Most
notably these were made by Freitag (1998) and Soderland
(1999), who both emphasised that their approaches work
on many different types of text, including HTML and
unstructured text, most commonly found on the web.
Modern approaches in IE are usually machine
learningbased approaches, and much focus has been put on
an area of IE called open information extraction (OIE).
These modern methods require little to no human
supervision and are focused on gathering information from web
sources
          <xref ref-type="bibr" rid="ref5">(Del Corro and Gemulla, 2013)</xref>
          . One of the first
OIE systems was TextRunner, presented by Yates et al.
(2007). Another system that gained wide-spread attention
is ClausIE, which made use of syntactic knowledge in
English
          <xref ref-type="bibr" rid="ref5">(Del Corro and Gemulla, 2013)</xref>
          . The most recent
approach that relies heavily on machine learning has been
presented by Stanovsky et al. (2018). This approach, although
supervised, makes use of state-of-the-art machine learning
architectures as well as semantic role labelling.
        </p>
        <p>
          Other areas related to IE are named entity recognition
(NER), named entity linking (NEL) and wikification. NER
refers to the NLP task of detecting entities in text,
including (proper) names, locations, institutions, dates and so on.
On the other hand, NEL refers to the process of linking
named entities to entries in large databases or knowledge
bases, essentially linking information together. Both
Petram et al. (2015) and Brouwer and Nijboer (2017) have
explored entity linking in the context of biographical data,
in order to gain further information on persons. Hachey
et al. (2013) implement and evaluate three existing NEL
approaches, while making use of Wikipedia to augment
their approach. However, this should not be confused with
wikification, which commonly refers to the task of linking
wikipedia pages to concepts, persons and so on mentioned
in texts, i.e. not linking named entities and databases in a
strict sense
          <xref ref-type="bibr" rid="ref15">(Hachey et al., 2013)</xref>
          .
        </p>
        <p>Extracting biographical information from the web is an
area that has been gaining more attention, especially for
the (automatic) creation of biographies and biographical
databases. Increasingly, this area uses IE and closely
related methods. Garcia and Gamallo (2015) have explored
different machine learning methods to extract
biographical relations in Portuguese. Furthermore, Wikipedia has
increasingly become a common source of information for
various applied methods. Approaches have either been used
to extract information from Wikipedia as a source of
information, such as Gotti and Langlais (2017). Russo et al.
(2015) explores methods to extract biographical
information from Wikipedia and DBpedia. Relevant information
that was extracted includes the name, birthplace, birth date,
and so on, of a person.</p>
        <p>Although NLP methods are being used in the context of
creating new biographies and biographical databases, there
are many issues that need to be addressed. More
specifically, Fokkens et al. (2014) point out these issues that were
encountered during the course of the BiographyNet project.
The authors point out that historic methods can often be
hard to transfer to computational or automatic methods,
since they rely on facts that may not be extracted directly,
in addition to being based on interpretations, logic, analysis
and so on. However, they also raise awareness of the
potential biases that historians could face when using NLP
methods. In particular, the authors argue that when using
rulebased methods, the rules and heuristics need to be clearly
indicated, and when machine learning approaches are used,
the training data and features used should be described.
Fokkens et al. (2014) use the example of provenance
modelling to demonstrate where these biases could occur and
how obvious they would be. More obvious cases could
be ambiguous geographical locations, which would factor
heavily in a rule-based approach. Less obvious would be
unbalanced datasets that may be used in machine learning,
leading to persons to be associated mistakenly with certain
topics. In their final conclusion, the authors emphasise that
the awareness on both the historian and NLP sides needs to
be raised to the problems explained in the paper.
Work on extracting biographical information from
Wikipedia using Wikidata on a larger-scale has been
carried out by Plum (2018). Research carried out for that
project can be seen as preliminary work for the methods
we describe here. It features a dataset of around 130, 000
lexicon pages about individual persons, which was selected
by using Wikidata with similar parameters set out here. A
short analysis of common structures containing
information was carried out. Plum (2018) not only demonstrates
how these common structures can be exploited and simple
rules applied in order to extract information about the
date of death and occupation of a person. The author also
points out some of the pit-falls when working with such
a large data-set, including the amount of processing time
and choosing an appropriate data-structure. We make use
of the analysis put forth and apply similar rules to extract
basic information. In addition, we make sure to take into
account the possible obstacles to overcome.</p>
        <p>3.</p>
      </sec>
      <sec id="sec-6-2">
        <title>Methods</title>
        <p>
          In this section we present the processing pipeline used
to carry out the experiments described in Section 4. The
pipeline consists of three main steps, and is depicted in
Figure 1. First, we pre-select a large amount of entities
using Wikidata and in accordance to parameters set out by
the project scope (Section 3.1.). We make the gathered
meta information easily accessible via a local MongoDB6
database. Next, we preprocess the data, using Stanford
CoreNLP
          <xref ref-type="bibr" rid="ref17">(Manning et al., 2014)</xref>
          to perform a variety of
NLP tasks including NER and dependency parsing
(Section 3.2.). We carry out shallow information extraction,
using rules similar to those developed by Plum (2018), and
described in Section 3.3. Finally, we describe an
experimental approach to find relevant candidates using location
matching, as well as a basic method of ranking these
candidates (Section 3.4.).
3.1. Data
The selection of relevant entities was carried out in a
twofold approach. First, we developed a simple Wikidata query
to return a selection of entities. Using this list of entities,
we then retrieved corresponding articles from Wikipedia.
Wikipedia is a large repository of information and offers a
vast amount of articles for almost any conceivable topic, in
various different languages. There are a number of projects
that extract information from Wikipedia and make it
accessible in the form of structured databases. One of these
projects is Wikidata, which we utilise here. As will become
more evident in later sections of this paper, it is clear that
Wikidata not only includes information from Wikipedia,
but also other sources, such as Wikivoyage, Wiktionary and
Wikisource. For this reason, we believe that by
combining information from Wikidata and Wikipedia we are able
to extract more relevant information than we would obtain
from only one of the sources.
        </p>
        <p>According to the parameters of the joint project, we select
entities that are listed as human and that have died between
1951 and 2019. The query returns the date and place of
death and birth, a short description, as well as the Wikidata
link and other identification numbers. Figure 2 shows the
query that was used and Figure 3 shows an example of the
returned results.</p>
        <p>
          A problem that occurs with the links to the individual
articles is that Wikipedia uses two types of links, one using
the name and one using an id number. Therefore, we use
a second Wikidata query to retrieve the ids for each page
in addition to the previously returned links (which use the
name). The id number cannot be returned using a standard
Wikidata query, i.e. there is no relation that can be
specified in order to gather this for all entities at the time of the
query. Instead, we use a separate query which takes each
Wikidata id individually, and generates the corresponding
Wikipedia id by extracting it from the Wikipedia article
itself. The choice was made to use these ids, as it easier to
retrieve the article for each entity from the Wikidump, which
we describe in due course. Using the name could lead to
some errors, due to differences in spelling across different
languages. This problem was also pointed out by
          <xref ref-type="bibr" rid="ref20">(Plum,
2018)</xref>
          , where it was found that using the ID is
unambiguous. Some examples are shown in Table 1.
        </p>
        <p>The Wikidata query returned 401, 695 entities (equivalent
to lexicon pages), and of these 172, 131 had
corresponding articles in English. From the total number returned, we
hope to use cross-lingual methods in future to extract from
articles in languages other than English (see Section 5.). It
is also worth mentioning at this point that some
discrepancies exist between Wikidata and Wikipedia information,
leading to some entities of the overall retrieved not being
used. A account of this will be presented in Section 3.2.
Once the selection of the entities and retrieval of basic
meta information was carried out, the next task involved
extracting the corresponding articles from the Wikidump.
A Wikidump is a snapshot of the whole Wikipedia
encyclopedia in XML format. We use a Python script called
WikiExtractor7 to convert the Wikidump from XML format
to plain text. The script processes the whole Wikidump,
returning each article as plain text, as well as minimal meta
information including the Wikipedia id and name of the
article. It does not retain any structure of the XML file or
provide any further markup or information. This was the
main reason for using the script, as processing a Wikidump
is time-consuming in itself and can involve many
complications. Therefore, we opted to use this already available
script in order to be able to process the vast XML files
which have a complex structure. From the converted plain
text, we extract all relevant articles using the previously
obtained Wikipedia ids. Articles are grouped together,
indicated by a begin and end tag and stored across plain text
files with sizes between 1Mb - 2Mb. We run a Python script
to extract each article by identifying the beginning and end
tags, and store it in individual JSON files, which makes the
following step of pre-processing easier.</p>
        <sec id="sec-6-2-1">
          <title>3.2. Text Preprocessing</title>
          <p>
            In order to facilitate the information extraction task, we
automatically annotate each article with linguistic
information. Due to the large amount of data we are processing, this
task could not be carried out during extraction. Therefore,
we process each article using Stanford CoreNLP
            <xref ref-type="bibr" rid="ref17">(Manning
et al., 2014)</xref>
            accessed via a Python script to carry out
annotation tasks. We run a tokenizer, sentence splitter, part
of speech tagger, lemmatizer, dependency parser and NER.
Previous work carried out in Plum (2018) has shown that
selecting an output format that is well structured and easy to
process is vital to this task. Each annotated article is saved
individually using a preset XML format available within
CoreNLP. It is important to point out that processing such
large amounts of data is extremely time-consuming.
Using very modern system with good system specifications
(6-core CPU, 32 GB Ram) the task took around one week of
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6https://www.mongodb.com/</title>
    </sec>
    <sec id="sec-8">
      <title>7https://github.com/attardi/wikiextractor</title>
      <p>continuous processing time. Using a more multi-threaded
workload could optimise this task, however, this will be
addressed in future work.</p>
      <p>Once processed, we were left with 170, 517 articles,
down 1, 614 from the previous number. Taking the
timeconsuming nature of this processing into account, we had
to set up a time-out for each request to the CoreNLP
annotation pipeline, meaning that extremely long articles were
not processed. Upon further inspection we also discovered
that a small number of texts had not been converted
properly, and contained either no text or corrupted text.
3.3.</p>
      <sec id="sec-8-1">
        <title>Basic Information Extraction</title>
        <p>For the information extraction step, we follow a shallow
rule-based approach in order to take advantage of the
basic information from Wikidata. Exploiting the structure of
Wikipedia articles, we extract the name and date of death
of an entity. Not only do we test if using simple rule-based
methods are viable for this kind of data, we also test a
method to determine candidates of relevance to the project.
As described in the introduction, candidates should have
had an influence in the Republic of Austria. Therefore, we
also extract locations mentioned in the articles and
determine whether they are Austrian, hence possibly hinting at
the fact that the entity of the article has some kind of
connection to Austria.</p>
        <p>
          By extracting the tokens of the heading of each article, we
extract the name. We remove any information that is
contained in brackets, which is sometimes the case in order to
disambiguate certain persons. Furthermore, we use a
simple rule based on observations and preliminary work carried
out by
          <xref ref-type="bibr" rid="ref20">(Plum, 2018)</xref>
          : the second full date that is mentioned
in the first sentence of each article is usually the date of
death. As each article has been tagged in terms of named
entities, including expressions of time, we simply extract
this from the annotations by iterating over the time
expressions. By full date we accept dates in the form
YYYY-MMDD. As a fall-back option, if only one expression is
annotated we select this. Should no dates be detected we do not
use anything. The extraction is carried out using a Python
script, which compares the extracted information with that
contained in the meta information.
3.4.
        </p>
      </sec>
      <sec id="sec-8-2">
        <title>Relevance Ranking</title>
        <p>As the articles have been annotated in terms of named
entities, we are able to extract locations by simply searching for
any LOCATION or CITY tags. We employ this approach
for each text, making a list of each location for each
article. Next, we determine the country of each location. Our
first approach was to use the GeoCoder api8 connected to
GeoNames to retrieve the country of each location.
Unfortunately, GeoNames is restricted to 1, 000 requests for
locations per hour. With such a large dataset, this would not
be a viable approach. Other than buying requests as part of
a premium service, we opt to download the full GeoNames
list of locations, which is freely available. This list
contains locations, as well as their country. Using MongoDB,
we create a database and index the location names to
ensure fast searching. This way we are able to query a local
database to determine the country for each location.
Using a custom Python script we query whether a location
among those found in each article belongs to the
Republic of Austria. If this is the case, we include the article or
entity as a candidate. In the first iteration, we found that
locations were being found in the documents, such as City.
These location names are always part of a longer name, but
are picked up as they are tagged individually. These
locations returned matches in the database, although these were
mostly mistakenly added. Therefore, we added the
criterion that any location has to have a population of more than
zero. We found that locations that have been mistakenly
added or are meant for some other purpose usually have a
population of zero.</p>
        <p>In order to test out how a method of ranking relevant
candidates could work, we count all locations in one article that
are in Austria. In addition, we try to count the false matches
as well. As there are a large amount of false matches we put
these in contrast to the main location count. The idea is that
these count could put the ranking into better perspective, if
for instance the counts are equal, this candidate could
possibly be excluded.</p>
        <p>4.</p>
        <sec id="sec-8-2-1">
          <title>Results</title>
          <p>In this section we present the results of the information
extraction task, as well as the selection of possible
candidates. The evaluation of our results presents a challenge,
due to problems selecting a gold standard. We compare
the extracted names and dates of birth to those returned
by Wikidata, but this assumes Wikidata as a gold standard.
Concerning the location extraction, we rely on the
performance of Stanford CoreNLP and the rules we employ as to
extracting the country. For reasons that will be explained
later on, we do not have a gold standard for evaluation
purposes here. Therefore, it should be clear that this is not an
evaluation of the extraction itself, but rather the process of
selecting candidates.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>8https://geocoder.readthedocs.io/api.html</title>
      <p>4.1.</p>
      <sec id="sec-9-1">
        <title>Wikipedia vs. Wikidata</title>
        <p>As described in the previous section, we extracted the name
and date of death from the Wikipedia articles. We
compared each result with the information obtained from
Wikidata. Of the 170, 517 articles, the name did not match in
18, 267 cases. Upon closer inspection, we found that this is
largely due to differences in spelling, and slight differences
in the name. Table 1 shows a selection of the most
common errors: The first two rows are examples of differences
in shorter names. Rows 3 and 4 show different levels of
preciseness in naming, i.e. abbreviations. The last row shows
an example where Wikidata returned the name in German,
whereas we extracted the name in English.</p>
      </sec>
      <sec id="sec-9-2">
        <title>Wikidata</title>
        <p>Robert Joshua
Francisco Javier Vidarte
Joe C. Davis, Jr.</p>
        <p>Vincent Graber
Karl Aloys von
und zu Liechtenstein</p>
      </sec>
      <sec id="sec-9-3">
        <title>Extracted</title>
        <p>Bob Joshua
Paco Vidarte
Joe C. Davis Jr.</p>
        <p>Vincent J. Graber Sr.</p>
        <p>Prince Karl Aloys
of Liechtenstein
In terms of the date of death, we had 30, 153 cases where
the date of death did not match the Wikidata records.
Using a Python script we counted the different errors that
occurred, and found that we could classify three main errors:
“no date” errors, “minor difference” errors and “birthday”
errors. A breakdown of how many times each error
occurred is shown in Table 2. In the first case, our extraction
rule returned 0000-00-00, indicating no date was extracted.
This error was caused by the Stanford NER algorithm not
detecting a date, or it being missing in the article itself. The
second most common error was the “minor difference
error”. In this case, the difference between the Wikidata date
and our extracted date was minor, i.e. only between one to
five days difference. We suggest that this shows that
Wikidata also gathers information from other sources, or that it
could be caused by timezone differences. The last error to
occur was the “birthday” error. Here, the date we extracted
did not match the date of death extracted from Wikidata,
but rather the corresponding date of birth. This is caused
by the fact that we extract a date in sentences, even if only
one date is found by the NER algorithm.</p>
        <p>Taking these results into account, it is interesting to
see where differences in data lie. Using Wikidata and
Wikipedia as complementary components has many
benefits. On the one hand, Wikidata serves excellently as a
tool to pre-select data according to some criteria. As
processing all Wikipedia, or complete Wikidumps is extremely
time-consuming, this reduces the time dramatically. On the
other hand, Wikidata can serve to some extent as a kind of
gold standard against which to compare the results of any
extraction carried out on Wikipedia articles. Of course, this
is only to a limited extent, as not all relations are available
in Wikidata. This is the case with our extraction of
locations in order to determine candidates.</p>
        <p>Going beyond the use as a gold standard, the two data
sources can also be used to extract information in a more
complementary way, i.e. using Wikidata for basic
information, and building on these known relations to extract
further information from Wikipedia. It may also be of interest
to compare contradicting information, as seen here with the
differences in date of death.
4.2.</p>
      </sec>
      <sec id="sec-9-4">
        <title>Biographical Dictionary Candidates</title>
        <p>Using our location matching script, we were able to obtain
13, 521 possible candidates. A short investigation of
candidates picked at random shows that our technique is
probably not precise enough. For each candidate, we list the
location that caused its inclusion in our candidates list. The ten
most common locations are listed in Table 3. While this list
includes many valid locations, it is clear that many articles
are chosen as candidates due to matches caused by Hall and
Point. Further examples include Sand, Fall and Gray and
so on. These match proper locations in Austria, however,
they also match English nouns and adjectives, and are most
probably part of longer location names. Another common
problem in this regard was the matching of names which
are ambiguous, as they also match locations in Austria, and
therefore contributed to being considered as candidates.</p>
      </sec>
      <sec id="sec-9-5">
        <title>Location</title>
        <p>Hall
Vienna
Point
Salzburg
Bergen
Nassau
Innsbruck
Graz
Ko¨ nigsberg
Inn</p>
        <p>Count
In terms of the ranking of the candidates, this is just as hard
to evaluate as to measure. A brief manual analysis
indicates that the ranking mechanism at this point is too crude.
Quite often candidates rank very highly, even though there
is no relevance to Austria whatsoever. This is mainly due
to the false matches, described previously. An extract of
some of the candidates below shows how highly some
irrelevant candidates rank. At the other end of the scale, the
reverse applies. Candidates that should probably be
considered with priority are ranked very low, due to only a few ore
one locations being matched. However, this is mostly due
to extremely short Wikipedia articles, which do not hold
much information.</p>
        <p>Further evaluation of the locations extracted against a gold
standard is not possible. Wikidata queries rely on a
relation, such as born in, to be present in order to extract the
corresponding location. In our case, we want to go beyond
these relations and find any kind of mention of locations
that are relevant. Ultimately, the candidates we are able to
derive are to be evaluated in an iterative process by
historians from the APIS/OEBL team in order to say how well our
method performs. Other forms of automatic evaluation do
not exist at this point in time, especially considering there
is no gold standard for this work, as it is mainly aimed at
finding completely new candidates.</p>
        <p>5.</p>
        <sec id="sec-9-5-1">
          <title>Conclusion and Future Work</title>
          <p>In this paper we presented an NLP pipeline to identify
biography candidates and to extract information about them
from Wikipedia and Wikidata. We show that shallow
extraction methods work well for obtaining basic
information about biography candidates. However, for determining
possible relevant candidates there is still work to be done.
While our simple method of matching locations works as
a wide net, there are many irrelevant inclusions. We
acknowledge that this metric by itself is too simple, however,
we feel that it could become an aspect of a future metric.
As the goal of this project is detecting relevant candidates,
we are currently working on improving our basic method,
hopefully making use of statistical or machine learning
based approaches in order to determine whether a person
has had some kind of relevance in a certain area. This could
also allow us to rank candidates according to their relevance
for that particular area and time period.</p>
          <p>In collaboration with the AAS we are working on a ranking
system which at present is based on implicit expert
knowledge. We hope to convert this knowledge to the machine,
and as a part of this ongoing effort we are working on an
annotating the dataset of the previous O¨ BL biographies. By
extracting sentences that show some kind of surface
relevance to Austria, we are researching the possibility of
training a machine learning classifier on the word and/or context
embeddings of these sentences, in order to automatically
detect them in text. This would also eliminate the need of
the large-scale pre-processing beforehand.</p>
          <p>In the future, we also aim to refine our information
extraction methods and to test them on different encyclopedic
repositories. In addition, we plan to explore cross-lingual
methods for extracting information from data sources in
other languages such as Czech, German, Hungarian, and
Slovak as the core languages of the former
AustrianHungarian monarchy.</p>
          <p>This research is a pilot endeavour to detect relevant
candidates for a biographical dictionary in online sources. It aims
to contribute to three goals, 1) the further compilation of a
digital, semi-automatic biographical dictionary on the case
study of the O¨ BL, 2) the further development of an
editing system for biographical dictionaries, which might be
used as a research infrastructure, on the use case of O¨ BL,
and 3) triggering interdisciplinary collaboration and further
pilot studies on methods and tools to detect people of
”relevance”.</p>
        </sec>
        <sec id="sec-9-5-2">
          <title>Acknowledgements</title>
          <p>The APIS project is funded by a research grant (project
number O¨ AW0405) of the Austrian Nationalstiftung fu¨ r
Forschung, Technologie und Entwicklung (Programm
“Digital Humanities - Langzeitprojekte zum kulturellen
Erbe”). We are grateful for technical support to access the
APIS and O¨ BL data sets by Katalin Lejtovicz and Matthias
Schlo¨ gl and are thankful for the feedback and evaluation of
the (art)historians A´goston Bernad and Maximilian Kaiser.
6.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Fadi</given-names>
            <surname>Biadsy</surname>
          </string-name>
          , Julia Hirschberg, and
          <string-name>
            <given-names>Elena</given-names>
            <surname>Filatova</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>An Unsupervised Approach to Biography Production Using Wikipedia</article-title>
          .
          <source>In Proceedings of ACL.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Judith</given-names>
            <surname>Brouwer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Harm</given-names>
            <surname>Nijboer</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <string-name>
            <given-names>Golden</given-names>
            <surname>Agents</surname>
          </string-name>
          .
          <article-title>A web of linked biographical data for the Dutch Golden Age</article-title>
          .
          <source>In Proceedings of BD2017.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Nancy</given-names>
            <surname>Chinchor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>David D</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Lynette</given-names>
            <surname>Hirschman</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Evaluating message understanding systems: an analysis of the third message understanding conference (MUC-3)</article-title>
          .
          <source>Computational linguistics</source>
          ,
          <volume>19</volume>
          (
          <issue>3</issue>
          ):
          <fpage>409</fpage>
          -
          <lpage>449</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Chisholm</surname>
          </string-name>
          , Will Radford, and
          <string-name>
            <given-names>Ben</given-names>
            <surname>Hachey</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Learning to generate one-sentence biographies from Wikidata</article-title>
          .
          <source>In Proceedings of ACL.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Luciano Del Corro</surname>
            and
            <given-names>Rainer</given-names>
          </string-name>
          <string-name>
            <surname>Gemulla</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>ClausIE: Clause-based Open Information Extraction</article-title>
          .
          <source>In Proceedings of WWW.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>O¨sterreichische Akademie der Wissenschaften</source>
          .
          <year>2012</year>
          . O¨ sterreichisches
          <source>Biographisches Lexikon</source>
          <year>1815</year>
          -1950, volume
          <volume>63</volume>
          .
          <source>Verlag der O¨ sterreichischen Akademie der Wissenschaften.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>George R Doddington</surname>
          </string-name>
          , Alexis Mitchell,
          <article-title>Mark A Przybocki, Lance A Ramshaw, Stephanie Strassel</article-title>
          , and
          <string-name>
            <surname>Ralph M Weischedel</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>The Automatic Content Extraction (ACE) program-tasks, data, and evaluation</article-title>
          .
          <source>In Proceedings of LREC.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Tomazˇ Erjavec</surname>
          </string-name>
          , Joh Dokler, and Petra Vide Ogrin.
          <year>2015</year>
          .
          <article-title>Slovenian biography</article-title>
          .
          <source>In Proceedings of BD2015.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Antske</given-names>
            <surname>Fokkens</surname>
          </string-name>
          , Serge Ter Braake, Niels Ockeloen, Piek Vossen,
          <string-name>
            <surname>Susan</surname>
            <given-names>Legeˆne</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guus Schreiber</surname>
          </string-name>
          , et al.
          <year>2014</year>
          .
          <article-title>Biographynet: Methodological issues when nlp supports historical research</article-title>
          .
          <source>In LREC</source>
          , pages
          <fpage>3728</fpage>
          -
          <lpage>3735</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Dayne</given-names>
            <surname>Freitag</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Information extraction from HTML: Application of a general machine learning approach</article-title>
          .
          <source>In Proceedings of AAAI/IAAI.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Marcos</given-names>
            <surname>Garcia</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pablo</given-names>
            <surname>Gamallo</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Exploring the effectiveness of linguistic knowledge for biographical relation extraction</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>21</volume>
          (
          <issue>4</issue>
          ):
          <fpage>519</fpage>
          -
          <lpage>551</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Fabrizio</given-names>
            <surname>Gotti</surname>
          </string-name>
          and
          <string-name>
            <given-names>Philippe</given-names>
            <surname>Langlais</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>From french Wikipedia to Erudit: A test case for cross-domain open information extraction</article-title>
          .
          <source>Computational Intelligence</source>
          ,
          <volume>34</volume>
          (
          <issue>2</issue>
          ):
          <fpage>420</fpage>
          -
          <lpage>439</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Ralph</given-names>
            <surname>Grishman</surname>
          </string-name>
          and
          <string-name>
            <given-names>Beth</given-names>
            <surname>Sundheim</surname>
          </string-name>
          .
          <year>1996</year>
          . Message Understanding Conference-
          <article-title>6: A brief history</article-title>
          .
          <source>In Proceedings of COLING.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Christine</given-names>
            <surname>Gruber</surname>
          </string-name>
          and
          <string-name>
            <given-names>Eveline</given-names>
            <surname>Wandl-Vogt</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Mapping historical networks: Building the new Austrian Prosopographical Biographical Information System (APIS). Ein U¨berblick</article-title>
          .
          <source>In Europa baut auf Biographien</source>
          . Aspekte, Bausteine,
          <article-title>Normen und Standards f u¨r eine europ a¨ische Biographik</article-title>
          ., pages
          <fpage>271</fpage>
          -
          <lpage>282</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Ben</given-names>
            <surname>Hachey</surname>
          </string-name>
          , Will Radford, Joel Nothman, Matthew Honnibal, and
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Curran</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Evaluating Entity Linking with Wikipedia</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>194</volume>
          :
          <fpage>130</fpage>
          -
          <lpage>150</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Lynette</given-names>
            <surname>Hirschman</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>The Evolution of evaluation: Lessons from the message understanding conferences</article-title>
          .
          <source>Computer Speech &amp; Language</source>
          ,
          <volume>12</volume>
          (
          <issue>4</issue>
          ):
          <fpage>281</fpage>
          -
          <lpage>305</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Christopher D. Manning</surname>
            , Mihai Surdeanu, John Bauer, Jenny Finkel,
            <given-names>Steven J.</given-names>
          </string-name>
          <string-name>
            <surname>Bethard</surname>
          </string-name>
          , and
          <string-name>
            <surname>David McClosky</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The Stanford CoreNLP Natural Language Processing Toolkit</article-title>
          .
          <source>In Proceedings of ACL.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Diego</given-names>
            <surname>Moussallem</surname>
          </string-name>
          , Thiago Castro Ferreira, Marcos Zampieri, Maria Claudia Cavalcanti, Geraldo Xexe´o, Mariana
          <string-name>
            <surname>Neves</surname>
          </string-name>
          , and
          <string-name>
            <surname>Axel-Cyrille Ngonga Ngomo</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>RDF2PT: Generating Brazilian Portuguese Texts from RDF Data</article-title>
          .
          <source>In Proceedings of LREC.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Lodewijk</given-names>
            <surname>Petram</surname>
          </string-name>
          , Jelle van Lottum,
          <source>Rutger van Koert, and Sebastiaan Derks</source>
          .
          <year>2015</year>
          .
          <article-title>Small Lives, Big Meanings Expanding the Scope of Biographical Data through Entity Linkage and Disambiguation</article-title>
          .
          <source>In Proceedings of BD2015.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Alistair</given-names>
            <surname>Plum</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Rule-based Information Extraction Using Wikipedia and Wikidata</article-title>
          .
          <source>Master's thesis</source>
          , University of Wolverhampton.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Reinert</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bernhard</given-names>
            <surname>Ebneth</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Interfaces: Accessing biographical data and metadata</article-title>
          .
          <source>In Proceedings of BD2017.</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Reinert</surname>
          </string-name>
          , Maximilian Schrott, Bernhard Ebneth, and
          <string-name>
            <given-names>Malte</given-names>
            <surname>Rehbein</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>From Biographies to Data Curation-The Making of www.deutsche-biographie</article-title>
          .de.
          <source>In Proceedings of BD2015.</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Irene</given-names>
            <surname>Russo</surname>
          </string-name>
          , Tommaso Caselli, and
          <string-name>
            <given-names>Monica</given-names>
            <surname>Monachini</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Extracting and Visualising Biographical Events from Wikipedia</article-title>
          .
          <source>In Proceedings of BD2015.</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <source>Matthias Schlo¨ gl and Katalin Lejtovicz</source>
          .
          <year>2017</year>
          .
          <article-title>A Prosopographical Information System (APIS)</article-title>
          .
          <source>In Proceedings of BD2017.</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Soderland</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Learning information extraction rules for semi-structured and free text</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>34</volume>
          (
          <issue>1</issue>
          ):
          <fpage>233</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Gabriel</given-names>
            <surname>Stanovsky</surname>
          </string-name>
          , Julian Michael, Luke Zettlemoyer, and
          <string-name>
            <given-names>Ido</given-names>
            <surname>Dagan</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Supervised Open Information Extraction</article-title>
          .
          <source>In Proceedings of NAACL.</source>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Yates</surname>
          </string-name>
          , Michael Cafarella, Michele Banko,
          <string-name>
            <given-names>Oren</given-names>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Broadhead</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Soderland</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>TextRunner: Open Information Extraction on the Web</article-title>
          .
          <source>In Proceedings of NAACL.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>