<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Semantic Enrichment of Newspapers: A Historical Ecology Use Case</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marieke van Erp</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas van Goethem</string-name>
          <email>tgoethem@science.ru.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katrien Depuydt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jesse de Does</string-name>
          <email>Jesse.Dedoesg@ivdnt.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Instituut voor de Nederlandse Taal</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Radboud University Nijmegen</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vrije Universiteit Amsterdam</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>39</fpage>
      <lpage>44</lpage>
      <abstract>
        <p>Historical ecology research relies on historical accounts of human-animal interactions to study this interaction through space and time. Newspaper archives are a rich source of information, but require careful querying and ltering to collect the relevant information. Traditionally, this is a laborious manual task. In this position paper, we describe our ongoing work on semantically enriching a newspaper collection to create a knowledge base to support historical ecological work.</p>
      </abstract>
      <kwd-group>
        <kwd>text enrichment</kwd>
        <kwd>historical ecology</kwd>
        <kwd>lexicology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Historical research often relies on manual inspection of documents. Historical
ecology investigates, a.o., the occurrence of particular animals in a distinctive
region over time. When using a large newspaper corpus, this would mean having
to sift through thousands of documents to identify whether each document is
relevant or not, before the researcher can even begin to analyse the content at
a more detailed level. In this position paper, we present ongoing work on the
creation of a knowledge base that provides a semantic enrichment layer over a
large newspaper corpus. We highlight the challenges we discovered in our rst
data analysis as well as the solutions we intend to implement to resolve these.</p>
    </sec>
    <sec id="sec-2">
      <title>Background and Related Work</title>
      <p>
        Historically, humans have had an ambivalent relationship with animals,
perceiving animals not only as sources of food, tools or totems, but also as threats
and nuisances. Many birds, small mammals and insects were believed to carry
diseases or to be harmful to crops or livestock. Furthermore, large predatory
species (e.g. wolf) or venomous species (e.g. viper) were feared for injuring or
killing humans [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These perceived threats have led to a `cultural fear' of pest
species, which has been reinforced through storytelling and mythology [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Recently, our relationship to many of these so-called \vermin" species has changed.
Some species are now valued as key species in nature rehabilitation, while others
are reintroduced to our country. It is therefore becoming increasingly relevant to
understand how these historical relationships between man and nature relate to
the present time. A comprehensive historical study on pest and nuisance species
is lacking for the Netherlands. Newspapers reporting on interactions with pest
and nuisance species may be an important source of information for such a study.
      </p>
      <p>Currently, the majority of such newspaper analyses are done manually. They
involve sending a query to the newspaper interface and clicking every article link,
reading the article and recording whether it is relevant to the research question or
not. We propose to automate the classi cation of newspaper articles and storing
the results in a knowledge base that contains structured, semantic information
about and extracted from the articles along with a link to the original articles4
to help researchers focus their time on a deeper analysis of the relevant articles.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Resources</title>
      <p>
        A mix of unstructured and structured resources is used. The newspaper corpus
is the main information source, but structured resources to systematically query
the newspapers and inform the language technology tools are used.
National Library Newspaper Corpus The Dutch National Library has
made available the original texts from 1.3 million newspapers, 1.5 million
magazine pages and 320,000 books from the 15th to the 21st century through
the Delpher portal.5 We focus on newspaper articles published between 1800
and 1940, for two reasons: 1) The OCR quality on these is most likely better
than on the older material and 2) This period also saw the \biological reveil",
a reawakening of interest in biological, in the Netherlands, which also may
be re ected in mentions of animals in newspapers [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Taxonomic Resources and Lexicons A list of pest and nuisance species
compiled in the ATHENA project,6 is used, which provides the latin name
and its common vernacular name. However, due to the local and
temporal variance in animal names we also employ diachronic lexicons that each
contain Dutch language variations across time and dialects78 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
4 Due to copyright restrictions it is not possible to include all article texts in the
knowledge base, but the articles are freely accessible through the Dutch National
Library newspaper portal.
5 http://www.delpher.nl
6 http://www.athena-research.org/
7 http://ivdnt.org/onderzoek-a-onderwijs/projecten/gigant
8 http://ivdnt.org/onderzoek-a-onderwijs/projecten/diamant
      </p>
      <p>Fig. 1. SERPENS Work ow
4</p>
    </sec>
    <sec id="sec-4">
      <title>Design of the Knowledge Base</title>
      <p>We aim to create a knowledge base containing information about animal reports.
We start by broadly querying the National Library Newspaper Corpus through
the taxonomic and diachronic resources (Step (1) in Figure 1). This results in
many articles returned from the newspaper collection that are not necessarily
about animals, such as persons with last name \De Wolf") (2). We therefore
employ document classi cation to lter out irrelevant documents (3). Our initial
analysis on this is presented in Section 5.</p>
      <p>Simply obtaining a set of relevant documents (4) is already useful to
ecological historians, but we would like to dive further into the documents and classify
what type of animal report it is (5). We will investigate what level of speci city
the tools can handle. This results in a knowledge base that contains document
classi cations, links to the Delpher sources, animal mentions, its spelling
variations, document metadata such as publication date and article length, and
factual information extracted from the documents (6). New mentions will be fed
back to enrich the lexicons.The knowledge base enables humanities and biology
researchers to study pest and nuisance species across species, space and time
(7).
5</p>
    </sec>
    <sec id="sec-5">
      <title>European polecats and lynxes</title>
      <p>As a rst use case, we chose to investigate documents mentioning `bunzing'
(European polecat) and `lynx' (Lynx). These two species are chosen as a rst
query on the database, which returns relatively modest result sets (2,515 and
5,530 documents respectively) showing a wide variety of topics in the documents.</p>
      <sec id="sec-5-1">
        <title>Language variation</title>
        <p>As our research covers a relatively long period of time, as well as a corpus that
contains quite some local newspapers, we expect to nd a fair amount of language
variation. Indeed, through the diachronic lexicons as well as the returned hits,
we nd a reasonable set of terms to expand our query with. Table 1 lists the
query terms employed for \bunzing" (European polecat), what type of language
variation they express and in the last column the number of hits in the corpus.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Document categories</title>
        <p>The hypothesis that will be tested is that the negative perception of native
species has subsided with the passing of time, while it has grown for invasive
(non-native) species. Categorization helps to dissect the newspaper corpus,
making it more easy to measure perception and understand its determinants. The
categories have been chosen to be broad enough to be applicable to a broad range
of species over an extended period of time, but speci c enough for a meaningful
analysis.</p>
        <p>
          Initially, we set out to classify whether a document is about the animal or
not. However, upon inspecting the document sets, we discovered that documents
that may not directly be about the animal, may also be informative and useful
to include in the research on species perceptions. For example, an uptick in ads
for `bunzinghonden' (dogs used in the hunt for polecats) or `bunzingklemmen'
(`polecat traps') can be indicative for the species to be a nuisance and thus for
people to want to get rid of them. Furthermore, gurative language use that
has a positive or negative connotation (`eyes like lynx' or `stinks like a polecat')
also says something about the public perception of a species. Upon annotating
about 100 examples with three annotators (consisting of a historical ecologist, a
lexicologist and a computational linguist), we came to the following categories,
which largely correspond to categories described in [
          <xref ref-type="bibr" rid="ref1 ref5">1, 5</xref>
          ].
        </p>
        <p>Natural history General articles about the animal, e.g. it subsists on birds or x
number were stu ed and became part of a museum collection
Nuisance, material damages The article mentions the animal as causing material
damages, e.g. beetles damaging crops or lynxes killing chickens
9 `Ulk' was a German satirical magazine whose cartoons were sometimes republished
in Dutch newspapers.
10 Fret (ferret) is a domesticated polecat
11 We found OCR errors that map `het' ('the') to `fret', further stressing the need for</p>
        <p>OCR correction and automatic document classi cation to yield relevant documents.
12 `egg thief'
Nuisance, immaterial damages The article mentions the animal as a nuisance
without material damages e.g. polecats found to walk over someone's face whilst
they were in bed, or (possibly irrational) fear for a certain animal
Pest control Organised hunt to bring down the number of pest species, e.g. ad for
hunting dogs
Hunt for economic reasons Hunting to use the fur, meat or other parts of the
animal e.g. an article mentioning that the hunting season has started again
Prevention Non-lethal actions against pest species, e.g. advice in the newspaper on
which plants keep away pest species
Accidents Mention of an unintentional encounter with the animal, e.g. roadkill
Figurative Figurative language featuring the animal e.g. eyes like a lynx
Other Articles not pertaining to the animal, e.g. a ship named `Lynx' or a person
whose last name is `Bunzing'</p>
      </sec>
      <sec id="sec-5-3">
        <title>Fact and Fiction</title>
        <p>Another interesting dimension of the dataset is that it does not only cover
`news' but also other types of texts. Currently, the National Library corpus
distinguishes 4 types of documents in its newspaper corpus: article, ad,
announcement and illustration. In our result set, we also encounter
crossword puzzles, feuilletons, poems and cartoons, which are all classi ed as
`article' in the metadata. This is understandable as the newspaper corpus
has been processed largely automatically, but for our purposes it makes
sense to distinguish at least between text with an `imaginative' primary aim
(a.o. ction) and text with an informative primary aim (non- ction), where
we put crossword puzzles in the imaginative category. We are annotating
the articles with these classes, and intend to train a classi er from this to
automatically detect these categories. We expect that the crosswords are the
easiest here (when looking at features such as the occurrence of horizontal,
vertical and numbers), but jokes are more di cult to identify automatically e.g.:
Guest: \Could you perhaps bring me a ferret?"
Waitress: \Why would you want one?"
Guest: \Perhaps it could nd the hare that is hidden in this jugged hare" 13</p>
      </sec>
      <sec id="sec-5-4">
        <title>Document quality</title>
        <p>
          It is well-known that Optical Character Recognition is not perfect, especially not
on older documents [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. As a rst attempt to identify which documents are of
highest and lowest quality, we compare each OCRed text to a historical lexicon
of known words and return the percentage of words recognised (cf. also [
          <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
          ] for
a lexical and a geometrical approach to quality assessment).
6
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Discussion and Future Work</title>
      <p>Semantic Web research revolves around structured data, but for humanities
researchers, text is often the core source of investigation. We argue that some of
13 Arnhemsche Courant, 09-01-1926, http://resolver.kb.nl/resolve?urn=MMKB08:
000106191:mpeg21:a0117
the manual data collection work for humanities researchers can be alleviated
through semantic enrichment of the texts. We propose to use a combination of
language technology and structured resources to create a knowledge base as a
more sophisticated entry point to collections.</p>
      <p>However, working with historical textual collections is not without challenges;
in this contribution. we have identi ed historical language variation, document
classi cation, and document quality as major problems to overcome. By
bringing together the knowledge of historical ecology, (historical) lexicography and
computational linguistics, we believe we are in the best position to address these
issues.</p>
      <p>The knowledge base we are creating will be published through Timbuctoo.14
Besides data access via SPARQL through an API, it provides a
programmingfree manner to access the data for humanities researchers. Currently, ports to
visualisation tools such as Gephi15 are being built. The annotated data and
experiments thusfar can be found at: http://www.github.com/clariah/serpens.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>The research for this paper was made possible by the CLARIAH-CORE project
nanced by NWO: http://www.clariah.nl
14 https://github.com/HuygensING/timbuctoo
15 https://gephi.org/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Lenders</surname>
            ,
            <given-names>H.J.R.</given-names>
          </string-name>
          :
          <article-title>Ten a penny? deadly viper bites in the netherlands in a socioeconomic perspective</article-title>
          .
          <source>Litteratura Serpentium</source>
          <volume>34</volume>
          (
          <year>2014</year>
          )
          <volume>290</volume>
          {
          <fpage>316</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Lenders</surname>
            ,
            <given-names>H.J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>I. A. W.</given-names>
            <surname>Janssen</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          :
          <article-title>The grass snake and the basilisk: From prechristian protective house god to the antichrist</article-title>
          .
          <source>Environment and History</source>
          <volume>20</volume>
          (
          <year>2014</year>
          )
          <volume>319</volume>
          {
          <fpage>346</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. van Berkel,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Voor Heimans en Thijsse: Frederik van Eeden sr. en de natuurbeleving in negentiende-eeuws Nederland</article-title>
          . Volume
          <volume>63</volume>
          .
          <string-name>
            <surname>Koninklijke Nederlandse Akademie van Wetenschappen</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Maks</surname>
            , I., van Erp,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vossen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoekstra</surname>
          </string-name>
          , R., van der Sijs, N.:
          <article-title>Integrating diachronous conceptual lexicons through linked open data</article-title>
          .
          <source>Presented at DHBenelux</source>
          <year>2016</year>
          (
          <volume>9</volume>
          -
          <issue>10</issue>
          <year>June 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dirke</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Where is the big bad wolf? notes and narratives on wolves in swedish newspapers during the eighteenth and nineteenth centuries</article-title>
          . In Masius, P.,
          <string-name>
            <surname>Sprenger</surname>
          </string-name>
          , J., eds.:
          <article-title>A fairy tale in question. Historical interactions between humans and wolves</article-title>
          . The White Horse Press, Cambridge (
          <year>2015</year>
          )
          <volume>101</volume>
          {
          <fpage>118</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Reynaert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Non-interactive ocr post-correction for giga-scale digitization projects</article-title>
          .
          <source>In: CICLing</source>
          , Springer (
          <year>2008</year>
          )
          <volume>617</volume>
          {
          <fpage>630</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Springmann</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fink</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>K.U.</given-names>
          </string-name>
          :
          <article-title>Automatic quality evaluation and (semi-) automatic improvement of mixed models for OCR on historical documents</article-title>
          .
          <source>CoRR abs/1606</source>
          .05157 (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gutierrez-Osuna</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Capitanu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auvil</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grumbach</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Furuta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandell</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Automatic assessment of ocr quality in historical documents</article-title>
          .
          <source>In: Proc. AAAI</source>
          . Volume in press. (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>