<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Challenging knowledge extraction to support the curation of documentary evidence in the humanities</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Enrico Daga</string-name>
          <email>enrico.daga@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrico Motta</string-name>
          <email>enrico.motta@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>The Open University</institution>
          ,
          <addr-line>Milton Keynes</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <abstract>
        <p>The identification and cataloguing of documentary evidence from textual corpora is an important part of empirical research in the humanities. In this position paper, we ponder the applicability of knowledge extraction techniques to support the data acquisition process. Initially, we characterise the task by analysing the endto-end process occurring in the data curation activity. After that, we examine general knowledge extraction tasks and discuss their relation to the problem at hand. Considering the case of the Listening Experience Database (LED), we perform an empirical analysis focusing on two roles: the listener and the place. The results show, among other things, how the entities are often mentioned many paragraphs away from the evidence text or are not in the source at all. We discuss the challenges emerged from the point of view of scientific knowledge acquisition.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Information systems → Information extraction; •
Computing methodologies → Information extraction; • Applied
computing → Arts and humanities.</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        The identification and cataloguing of documentary evidence from
textual corpora is an important part of empirical research in the
humanities. An increasing number of recent initiatives in the
digital humanities have as primary objective the curation of a
database collecting text excerpts augmented with fine-grained
metadata, mentioned entities, and their relations, often in the form of
knoweldge graphs developed adopting the linked data paradigm.
These databases are developed following controlled processes, in
the spirit of digital library management, where the identification
and onboarding of relevant information is substantially entrusted
to research students, librarians, and similar domain experts. The
Listening Experience Database Project (LED)1, for example, is an
initiative aimed at collecting accounts of people’s private experiences
of listening to music [
        <xref ref-type="bibr" rid="ref5">4</xref>
        ]. Since 2012, the LED community explored
a wide variety of sources, collecting over 10.000 unique experiences.
These are catalogued through a sophisticated workflow but more
importantly by means of a rich ontology covering a variety of
aspects related to the experience, for example, the time and place it
occurred, the source where the evidence has been retrieved, and
the entities involved, such as, a performer, a composer, or a creative
work [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Another example is the UK Reading Experience Database
(RED). UK RED includes over 30,000 records of reading experiences
sourced from the English literature. The curatorial efort required
to populate these databases was significant and the size and quality
of these databases is a major achievement of these projects.
      </p>
      <p>In this position paper we ponder the applicability of knowledge
extraction techniques to support the data curation activity. Initially,
we introduce the case study and analyse the data curation activity.
After that, we examine general knowledge extraction tasks and
discuss their relation to the problem at hand. Considering the case
of the Listening Experience Database (LED), we perform an
empirical analysis of a portion of the database, focusing on the role
"listener" and "place". Specifically, we elaborate on the hypothesis
that the related entities can be automatically retrieved from the
source. Finally, we discuss a set of challenges for knowledge
extraction related to supporting the curation of this type of evidence
databases.
2</p>
    </sec>
    <sec id="sec-3">
      <title>DATA CURATION ACTIVITY</title>
      <p>
        In general, the discovery and selection of documentary evidence
is an activity that may not be conducted systematically. However,
in the context of enterprises such as the LED project, there is an
attempt to objectively select, extract, and curate documentary
evidence from texts. From the curator’s perspective, it is not about
searching archives or repositories but exploring specific sources of
value, for example, specific books. In [
        <xref ref-type="bibr" rid="ref9">8</xref>
        ] we developed an approach
for retrieving textual excerpts relevant for a certain theme of
interest in a book by combining language analysis, entity recognition,
and a general purpose knowledge graph (DBpedia) and showed
that many of those pieces of evidence are characterised by implicit
information. In addition, once the text is found, populating all the
metadata is a long and dificult task.
      </p>
      <p>To illustrate the problem, let’s consider two examples from the
LED project:</p>
      <p>E1 "Music is certainly a pleasure that may be reckoned intellectual,
and we shall never again have it in the perfection it is this year, because
Mr. Handel will not compose any more! Oratorios begin next week,
to my great joy, for they are the highest entertainment to me."2 The
2Source: Mary Granville, and Augusta Hall (ed.), Autobiography and Correspondence
of Mary Granville, Mrs Delany: with interesting Reminiscences of King George the
Third and Queen Charlotte, volume 1 (London, 1861), p. 594. https://led.kmi.open.ac.
uk/entity/lexp/1444424772006 accessed: 30 September, 2019.
excerpt refers to Mrs Delany’s report of a (series of) live
performances of Operas and Oratorios by George Frideric Handel,
happened in March, 1737.</p>
      <p>E2 "I then went to Amsterdam to conduct Oedipus at the
Concertgebouw, which was celebrating its fortieth anniversary by a series of
sumptuous musical productions. The fine Concertgebouw orchestra,
always at the same high level, the magnificent male choruses from
the Royal Apollo Society, soloists of the first rank - among them Mme
Hélène Sadoven as Jocasta, Louis van Tulder as Oedipus, and Paul
Huf, an excellent reader - and the way in which my work was
received by the public, have left a particularly precious memory that
I recall with much enjoyment."3 Stravinsky, in the beginning of
1928, celebrates the high level of the Concertgebouw orchestra
and singers performing his Oedipus Rex. All of them are listed as
entities in the LED database.</p>
      <p>In both examples, several of the entities involved are not
mentioned in the excerpt and are derived from the curator’s knowledge
of the source (for example, Mrs Delany is the author of the letter in
E1) and the domain (e.g. the full name of the work is Oedipus Rex
in E2).</p>
      <p>
        Here we focus on the challenge of automatically populating the
record and support an expert in identifying, collecting and inputting
the relevant information. In other words, we aim at automatically
populating (as many as possible) roles of the ontology. For instance,
a listening experience specification can be derived from the
available graph on data.open.ac.uk [
        <xref ref-type="bibr" rid="ref8">7</xref>
        ]. The type ListeningExperience
includes the following properties, among others (we omit
namespaces for readability):
• agent (who is the listener)
• time (when the listening event occurred)
• place (where it occurred)
• subject (what was listened)
• is_reported_in (a link to the source)
• has_environment (e.g. was it a public or a private place,
indoor or outdoor)
A ListeningExperience is related to other relevant items, notably
Performance, WrittenWork, MusicArtist, and Country. The
knowledge extraction system should be able to derive the requirements
from the ontology specification, primarily the data values and roles
involved. For example, it should derive the requirement to find
the agent of the ListeningExperience, its place and time, and
that there may be a specific musical work to be identified and,
eventually, the author of the musical work, filling the roles
associated to the path subject -&gt; ? a Performance -&gt; performance of -&gt; ? a
MusicExpression.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>KNOWLEDGE EXTRACTION</title>
      <p>Knowledge extraction is a branch of artificial intelligence
covering a variety of tasks related to the automatic or semi-automatic
derivation of formal symbolic knowledge from unstructured or
semi-structured sources4.</p>
      <p>
        The area comprehends research in a variety of problems
related to lifting an unstructured or semi-structured source into an
3Igor Stravinksy, Igor Stravinsky: An Autobiography (1936), p. 139. https://led.kmi.
open.ac.uk/entity/lexp/1435674909834 accessed: 30 September, 2019.
4For a general introduction, see [
        <xref ref-type="bibr" rid="ref17">16</xref>
        ].
output described using a knowledge representation formalism.
Entity extraction and classification are two related tasks referring
to the location of mentions of entities in an input text and their
categorization, as in the following example: "We went to the
rehearsal of JoshuaP er son last TuesdayT ime ". Entity Linking,
instead, refers to finding mentions of entities from a database into
a natural language resource or, similarly, to appropriately
disambiguate words by associating a knowledge base identifier. Often,
the three tasks are performed together and labelled Named
Entity Recognition and Classification (NERC) [
        <xref ref-type="bibr" rid="ref13">12</xref>
        ]. Linked Data
and NER together have been extensively employed in a number
of knowledge extraction and data mining tasks (e.g., the work of
H. Paulheim [
        <xref ref-type="bibr" rid="ref22">21</xref>
        ]). Relation extraction refers to the
identification of n − ary relations (for n ≥ 2) within the source, usually
addressed with a combination of NLP and machine learning
techniques [
        <xref ref-type="bibr" rid="ref23">22</xref>
        ]. The relations Composer(Opedipus Rex,Starvinsky)
and Performed(Opedipus Rex,Concertgebouw,1928) are two
examples. Event extraction is a special case of relation
extraction where the focus is on identifying an event, usually an action
being performed by an agent in a certain setting. This task is
extensively studied in domains such as Biomedicine [
        <xref ref-type="bibr" rid="ref6">5</xref>
        ], Finance and
Politics [
        <xref ref-type="bibr" rid="ref16">15</xref>
        ], and Science [
        <xref ref-type="bibr" rid="ref27">26</xref>
        ]. Approaches dedicated to the
detection and extraction of historical and biographical events are
designed in [
        <xref ref-type="bibr" rid="ref26 ref30">25, 29</xref>
        ]. The notion of event is generally considered as
something happening at a specific time and place, which constitutes
an incident of substantial relevance [
        <xref ref-type="bibr" rid="ref15">14</xref>
        ]. Therefore, the objective is
to identify the action triggering the event (e.g. the verb perform) and
then the associated roles. Data-driven approaches usually involve
statistical reasoning or probabilistic methods like Machine Learning
techniques. In contrast, knowledge-based methods are generally
top-down and based on pre-defined templates, for example,
lexicosemantic patterns [
        <xref ref-type="bibr" rid="ref16">15</xref>
        ]. The two approaches can be combined and
machine learning methods used to learn such patterns [
        <xref ref-type="bibr" rid="ref24">23</xref>
        ].
However, the notion of event is still ill-defined in NLP research and
this makes it hard to develop methods which are portable,
efectively, to multiple domains [
        <xref ref-type="bibr" rid="ref15">14</xref>
        ]. Research in open domain event
extraction focuses essentially on social media data [
        <xref ref-type="bibr" rid="ref25">24</xref>
        ] where the
task is the extraction of statements for summarization purposes,
similar to the one of key-phrases extraction [
        <xref ref-type="bibr" rid="ref29">28</xref>
        ]. Ontology-based
information extraction (OBIE) uses formal ontologies to guide the
extraction process [
        <xref ref-type="bibr" rid="ref18 ref28">17, 27</xref>
        ]. Relevant work in the area is surveyed
in [
        <xref ref-type="bibr" rid="ref10 ref20">9, 19</xref>
        ]. In 2013, Gangemi provided an introduction and
comparison of fourteen tools for knowledge extraction over unstructured
corpora, where the task is defined as general purpose machine
reading [
        <xref ref-type="bibr" rid="ref11">10</xref>
        ]. A machine reader transforms a natural language text
into formal knowledge, according to a shared semantics. State of art
methods include FRED [
        <xref ref-type="bibr" rid="ref12">11</xref>
        ] and PIKES [
        <xref ref-type="bibr" rid="ref7">6</xref>
        ]. These approaches are
based on a frame-based semantics that is at the same time
domainand task-independent. Instead, a domain-oriented solution would
identify knowledge components of interest in the text, similarly to
what explored, for example, in the work of Alani [
        <xref ref-type="bibr" rid="ref4">3</xref>
        ]. This task is
also considered as an automatic ontology instantiation [
        <xref ref-type="bibr" rid="ref3">2</xref>
        ] or
semiautomatic creation of metadata [
        <xref ref-type="bibr" rid="ref14">13</xref>
        ]. A suitable approach should be
able to detect the requirements from a domain-specific ontology
and, having as input the text excerpt, the source metadata, and
potentially other knowldge bases, generate suitable hypotheses of
values and entities on any relevant role.
      </p>
      <p>CThhailrlednIgnitnegrnkantoiownlaeldWgeoerxktsrhaocption tCoasputpuprionrgt Scientific Knowledge (Sciknow), November 19th, 2019. Collocated with the tenth International Conference on Knowledge Capture
the curation of documentary evidence in the humanities (K-CAP), Los Angeles, CA, USA.</p>
    </sec>
    <sec id="sec-5">
      <title>EMPIRICAL ANALYSIS</title>
      <p>To discuss the feasibility and dificulty of the task, we relax the
problem and verify to what extent the entities that are part of the
curated metadata could potentially be automatically derived from
the sources. Specifically, we want to answer the questions: ( Q1)
Could a system find the target entities in the excerpt? ( Q2) Could a
system find the target entities in the text surrounding the excerpt?
(Q3) How far from the excerpt the entity is? (Q4) Could it be found
in the metadata of the source?</p>
      <p>
        We consider the case of the LED database and focus on two
relation and roles: the listener and the place of the listening event. The
LED curation workflow reuses entities from DBpedia, MusicBrainz,
and also defines new entities in the Linked Data. Our analysis is
limited to books from archive.org annotated with a listener or place
from DBpedia. We use DBpedia Spotlight [
        <xref ref-type="bibr" rid="ref21">20</xref>
        ] as entity recognition
and linking system.
      </p>
      <p>
        First, we need to find the position of the evidence text back in
the original source. Identifying the position of LED items in the
original book is not an easy task. In fact, the process of reporting
an excerpt from the book involves a number of modifications in
the format that makes it very rare the chance that a precise text
match would work. In addition, the reported text includes often
omissis or rephrasing in order to include co-references derived
from previous paragraphs. To solve the problem, we developed the
algorithm presented in Listing 1. The method is based on using the
longest words as locators. The algorithm selects the occurrences of
the longest words and isolate the surrounding portion of text using
the length of the excerpt as heuristic. The resulting candidates are
then ranked according to their similarity against the excerpt using
the well-known Levenshtein distance [
        <xref ref-type="bibr" rid="ref19">18</xref>
        ]. The candidate with the
lowest score is elected as the original text.
      </p>
      <p>Figure 1 illustrates the features of the corpus. Of the 9059
listening experiences in the database with a textual excerpt reported,
7999 include a place (88.3%) and in 7222 of them the place points to
DBpedia (79.9%). The agent is specified in 8258 of them (91.2%) but
only 2996 refer to a DBpedia entity (33.1%). In all other cases the
listener is created as a novel entity.</p>
      <p>64.8% of the listeners are also the authors of the text - 5874 cases.
This is not surprising as one of the most researched type of resources
were memories, diaries, and collection of letters. In addition, this
answers our Q4 and shows how important it could be to intelligently
derive information from the source metadata. However, less than
half of the agents exist in DBpedia (2130 times, 23.5% of the total).
Finally, only 11.3% of the sources could be retrieved as open texts,
referring to 1026 of the documentary evidence in the database. Of
Listing 1: Detect the location of an excerpt in a source.
e x c e r p t , S o u r c e ;
b e s t [ t , b , e , s ] ; / / t e x t , begin , end , s c o r e
words [ ] = t o k e n i z e ( e x c e r p t )
words [ ] = s o r t B y L e n g t h D e s c ( words [ ] ) / / L o n g e s t on t o p
F o r e a c h word i n words [ ] :
o c c u r r e n c e s [ ] [ b , e ] = f i n d ( word , S o u r c e )
p o s i t i o n [ b , e ] = f i n d ( word , e x c e r p t )
F o r e a c h o c c u r r e n c e [ b , e ] i n o c c u r r e n c e s [ ] [ b , e ] :
b e g i n = o c c u r r e n c e . b − p o s i t i o n . b
end = o c c u r r e n c e . e + l e n ( e x c e r p t ) − p o s i t i o n . e
p o s s i b l e = s u b s t r i n g ( Source , begin , end )
s c o r e = l e v e n s h t e i n ( e x c e r p t , p o s s i b l e )
i f ( s c o r e &lt; b e s t [ s ] )</p>
      <p>b e s t [ t , b , e , s ] = [ p o s s i b l e , begin , end , s c o r e ]
f i
End
End
r e t u r n b e s t
(a) Places.</p>
      <p>(b) Agents.
these, 7.3% includes DBpedia entities as place or agent, 690 excerpts
from 26 books. These are the objects in our analysis.</p>
      <p>Results are summarised in Figure 2. Charts display the distance
of the entity mentions, measured in number of paragraphs5. This
analysis is partial as it only covers DBpedia entities being used as
places or agents (listeners) with relation to books which sources
we could retrieve from the Web. However, the answers to the
remaining questions are quite interesting. (Q1) The DBpedia place
was mentioned in the textual excerpt only in 25.9% of the observed
cases (179). The listener was mentioned in the excerpt only in 13
cases, 13.4% of the observed population (97). (Q2) 10% of the times
the place mention is less then 5 paragraphs from the evidence text.
The agent is mentioned within 5 paragraphs from the evidence
in 4% of the observed cases. (Q3) 83.2% of the times the DBpedia
place was explicitely mentioned at least once in the source (574). In
79 cases (11.4%) the place hasn’t been found either in the excerpt
or anywhere else in the source. A similar result is observable for
agents. Finally, there is good chance the entity is somewhere away
from the evidence text.
5</p>
    </sec>
    <sec id="sec-6">
      <title>CHALLENGES</title>
      <p>There are several aspects that make the task of automatically
supporting the acquisition of knowledge about documentary evidence
particularly interesting from the point of view of scientific
knowledge acquisition.
5Text segmentation is itself a dificult task. In our analysis, we measured distances
in number of characters, considered one word to be 5 characters (the approximated
average length in english) and one paragraph to amount to 200 words.</p>
      <p>An important characteristic is the amount of implicit
information necessary to characterise the documentary evidence that is
not derivable from the reference text. As a result, a typycal
knowledge extraction approach may fail at performing an inference that
is normally the result of user’s expertise. A domain-independent
machine reader could produce a formal representation of the text
with entities and roles linked together. Theoretically, processing a
text through a machine reading system would reduce the problem
to one of ontology alignment. However, as we have seen, the needed
entities may not be mentioned in the text excerpt at a reasonable
proximity. In addition, having to deal with an ontology alignment
problem does not necessarily reduces the distance to the goal.</p>
      <p>Crucially, metadata about the sources should be used to derive
information such as the time span of the documentary material or
information about the author(s). Determining who is the person
reporting the event could contribute to populate the agent (for
firstperson reports) but also on deriving more contextual information,
for example, related to the historical period or the interests of the
author. Linking an author to a knowledge graph (such as DBpedia)
could provide insight on the validity of the hypotheses for assigning
certain roles, for example, by deriving that Stravinsky is the author
of Oedipus Rex (E2). Therefore, a general solution should be able
to reason upon contextual knowledge. Intuitively, the system
should be capable of fitting within the constraints of the domain
specific ontology and exploit it to tailor the approach. The ontology
specification would provide information about the main types and
relations of interest, and those can be used to derive contextual
information from existing commons sense knowledge bases (e.g.
ConceptNet6).</p>
      <p>Although it may seem that these databases have a limited domain
of interest, there are few chances that the variety of types and
entities useful could be found in a single, encyclopedic, knowledge
base. In the case of the LED project, part of the Linked Open Data,
the documentary evidence links to a variety of external resources
(e.g. MusicBrainz7 and Geonames8). The system should be able to
work across distributed and heterogeneus datasets in search for
relevant resources. These may include common-sense knowledge
and linguistic resources, textual corpora, gazetteeres, thesauri, and
specialised digital libraries. Ultimately, the system should be able
to recognise entities and their roles despite the fact that they can
be linked to any reference database.</p>
      <p>Ultimately, cultural studies like the ones performed in the LED
and RED projects often coin novel concepts, such as Listening
Experience, whose structure and features cannot be found in
preexisting databases. In fact, the definition of a concept of interest
is itself a scientific output for which the database constitutes the
empirical proof of relevance to scholarship in the related field. It is
an open question to what extent learning from one of such databases
could help in supporting a new, coming one.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Adamou</surname>
          </string-name>
          , Mathieu d'Aquin,
          <string-name>
            <given-names>Helen</given-names>
            <surname>Barlow</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Simon</given-names>
            <surname>Brown</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>LED: curated and crowdsourced linked data on music listening experiences</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>Proceedings of the ISWC 2014 Posters &amp; Demonstrations Track</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Harith</given-names>
            <surname>Alani</surname>
          </string-name>
          , Sanghee Kim, David E Millard, Mark J Weal, Wendy Hall, Paul H Lewis, and
          <string-name>
            <given-names>Nigel</given-names>
            <surname>Shadbolt</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Web based knowledge extraction and consolidation for automatic ontology instantiation</article-title>
          . (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Harith</given-names>
            <surname>Alani</surname>
          </string-name>
          , Sanghee Kim, David E Millard, Mark J Weal, Wendy Hall, Paul H Lewis, and
          <string-name>
            <surname>Nigel R Shadbolt</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Automatic ontology-based knowledge extraction from web documents</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          <volume>18</volume>
          ,
          <issue>1</issue>
          (
          <year>2003</year>
          ),
          <fpage>14</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Helen</given-names>
            <surname>Barlow</surname>
          </string-name>
          and
          <string-name>
            <given-names>David</given-names>
            <surname>Rowland</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Listening to music: people, practices and experiences</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Jari</given-names>
            <surname>Björne</surname>
          </string-name>
          , Filip Ginter, Sampo Pyysalo,
          <string-name>
            <surname>Jun'ichi Tsujii</surname>
            , and
            <given-names>Tapio</given-names>
          </string-name>
          <string-name>
            <surname>Salakoski</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Complex event extraction at PubMed scale</article-title>
          .
          <source>Bioinformatics</source>
          <volume>26</volume>
          ,
          <issue>12</issue>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Francesco</given-names>
            <surname>Corcoglioniti</surname>
          </string-name>
          , Marco Rospocher, and Alessio Palmero Aprosio.
          <year>2016</year>
          .
          <article-title>Frame-based ontology population with PIKES</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>28</volume>
          ,
          <issue>12</issue>
          (
          <year>2016</year>
          ),
          <fpage>3261</fpage>
          -
          <lpage>3275</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Daga</surname>
          </string-name>
          , Mathieu d'Aquin,
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Adamou</surname>
          </string-name>
          , and Stuart Brown.
          <year>2016</year>
          . The Open University Linked Data - data.open.ac.uk.
          <source>Semantic Web</source>
          <volume>7</volume>
          ,
          <issue>2</issue>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Daga</surname>
          </string-name>
          and
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Motta</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Capturing themed evidence, a hybrid approach</article-title>
          .
          <source>In 18th Int</source>
          .
          <article-title>Conference on Knowledge Capture (K-CAP)</article-title>
          . ACM, To appear.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Dejing</given-names>
            <surname>Dou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Hao</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Haishan</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Semantic data mining: A survey of ontology-based approaches</article-title>
          .
          <source>In Proceedings of the 2015 IEEE 9th international conference on semantic computing (IEEE ICSC</source>
          <year>2015</year>
          ). IEEE,
          <fpage>244</fpage>
          -
          <lpage>251</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Aldo</given-names>
            <surname>Gangemi</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A comparison of knowledge extraction tools for the semantic web</article-title>
          .
          <source>In Extended semantic web conference</source>
          . Springer,
          <fpage>351</fpage>
          -
          <lpage>366</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Aldo</surname>
            <given-names>Gangemi</given-names>
          </string-name>
          , Valentina Presutti, Diego Reforgiato Recupero, Andrea Giovanni Nuzzolese, Francesco Draicchio, and
          <string-name>
            <given-names>Misael</given-names>
            <surname>Mongiovì</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Semantic web machine reading with FRED</article-title>
          .
          <source>Semantic Web</source>
          <volume>8</volume>
          ,
          <issue>6</issue>
          (
          <year>2017</year>
          ),
          <fpage>873</fpage>
          -
          <lpage>893</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Archana</surname>
            <given-names>Goyal</given-names>
          </string-name>
          , Vishal Gupta, and
          <string-name>
            <given-names>Manish</given-names>
            <surname>Kumar</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Recent named entity recognition and classification techniques: a systematic review</article-title>
          .
          <source>Computer Science Review</source>
          <volume>29</volume>
          (
          <year>2018</year>
          ),
          <fpage>21</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Siegfried</surname>
            <given-names>Handschuh</given-names>
          </string-name>
          , Stefen Staab, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Ciravegna</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>S-CREAM-semiautomatic creation of metadata</article-title>
          .
          <source>In International Conference on Knowledge Engineering and Knowledge Management</source>
          . Springer,
          <fpage>358</fpage>
          -
          <lpage>372</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Frederik</surname>
            <given-names>Hogenboom</given-names>
          </string-name>
          , Flavius Frasincar, Uzay Kaymak, Franciska De Jong, and
          <string-name>
            <given-names>Emiel</given-names>
            <surname>Caron</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A survey of event extraction methods from text for decision support systems</article-title>
          .
          <source>Decision Support Systems</source>
          <volume>85</volume>
          (
          <year>2016</year>
          ),
          <fpage>12</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Wouter</surname>
            <given-names>IJntema</given-names>
          </string-name>
          , Jordy Sangers, Frederik Hogenboom, and
          <string-name>
            <given-names>Flavius</given-names>
            <surname>Frasincar</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A lexico-semantic pattern language for learning ontology instances from text</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>15</volume>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Jing</given-names>
            <surname>Jiang</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Information extraction from text</article-title>
          .
          <source>In Mining text data</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Vangelis</surname>
            <given-names>Karkaletsis</given-names>
          </string-name>
          , Pavlina Fragkou, Georgios Petasis, and
          <string-name>
            <given-names>Elias</given-names>
            <surname>Iosif</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Ontology based information extraction from text. In Knowledge-driven multimedia information extraction and ontology evolution</article-title>
          . Springer,
          <fpage>89</fpage>
          -
          <lpage>109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Vladimir</surname>
            <given-names>I</given-names>
          </string-name>
          <string-name>
            <surname>Levenshtein</surname>
          </string-name>
          .
          <year>1966</year>
          .
          <article-title>Binary codes capable of correcting deletions, insertions, and reversals</article-title>
          .
          <source>In Soviet physics doklady</source>
          , Vol.
          <volume>10</volume>
          .
          <fpage>707</fpage>
          -
          <lpage>710</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Jose L Martinez-Rodriguez</surname>
            ,
            <given-names>Aidan</given-names>
          </string-name>
          <string-name>
            <surname>Hogan</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ivan</surname>
          </string-name>
          Lopez-Arevalo.
          <year>2018</year>
          .
          <article-title>Information extraction meets the semantic web: a survey</article-title>
          .
          <source>Semantic Web</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Pablo</surname>
            <given-names>N Mendes</given-names>
          </string-name>
          , Max Jakob, Andrés García-Silva, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>DBpedia spotlight: shedding light on the web of documents</article-title>
          .
          <source>In Proceedings of the 7th international conference on semantic systems. ACM</source>
          , 1-
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Heiko</given-names>
            <surname>Paulheim</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Exploiting Linked Open Data as Background Knowledge in Data Mining</article-title>
          .
          <source>DMoLD</source>
          <volume>1082</volume>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Sachin</surname>
            <given-names>Pawar</given-names>
          </string-name>
          ,
          <article-title>Girish K Palshikar,</article-title>
          and
          <string-name>
            <given-names>Pushpak</given-names>
            <surname>Bhattacharyya</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Relation extraction: A survey</article-title>
          .
          <source>arXiv preprint arXiv:1712.05191</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Jakub</surname>
            <given-names>Piskorski</given-names>
          </string-name>
          , Hristo Tanev, and Pinar Oezden Wennerberg.
          <year>2007</year>
          .
          <article-title>Extracting violent events from on-line news for ontology population</article-title>
          .
          <source>In International Conference on Business Information Systems</source>
          . Springer,
          <fpage>287</fpage>
          -
          <lpage>300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Alan</surname>
            <given-names>Ritter</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Oren</given-names>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sam</given-names>
            <surname>Clark</surname>
          </string-name>
          , et al.
          <year>2012</year>
          .
          <article-title>Open domain event extraction from twitter</article-title>
          .
          <source>In Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM</source>
          ,
          <volume>1104</volume>
          -
          <fpage>1112</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Roxane</surname>
            <given-names>Segers</given-names>
          </string-name>
          , Marieke Van Erp,
          <string-name>
            <surname>Lourens Van Der Meij</surname>
          </string-name>
          , Lora Aroyo, Jacco van Ossenbruggen,
          <string-name>
            <surname>Guus Schreiber</surname>
            , Bob Wielinga, Johan Oomen, and
            <given-names>Geertje</given-names>
          </string-name>
          <string-name>
            <surname>Jacobs</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Hacking history via event extraction</article-title>
          .
          <source>In Proceedings of the sixth international conference on Knowledge capture. ACM.</source>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Maria</given-names>
            <surname>Vargas-Vera</surname>
          </string-name>
          and
          <string-name>
            <given-names>David</given-names>
            <surname>Celjuska</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Event recognition on news stories and semi-automatic population of an ontology</article-title>
          .
          <source>In IEEE/WIC/ACM International Conference on Web Intelligence (WI'04)</source>
          . IEEE,
          <fpage>615</fpage>
          -
          <lpage>618</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Daya</surname>
            <given-names>C Wimalasuriya</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Dejing</given-names>
            <surname>Dou</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Ontology-based information extraction: An introduction and a survey of current approaches</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Ian</surname>
            <given-names>H Witten</given-names>
          </string-name>
          , Gordon W Paynter, Eibe Frank, Carl Gutwin, and Craig G NevillManning.
          <year>2005</year>
          .
          <article-title>Kea: Practical automated keyphrase extraction</article-title>
          .
          <source>In Design and Usability of Digital Libraries: Case Studies in the Asia Pacific . IGI Global</source>
          ,
          <volume>129</volume>
          -
          <fpage>152</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Kalliopi</surname>
            <given-names>Zervanou</given-names>
          </string-name>
          , Ioannis Korkontzelos, Antal Van Den Bosch, and
          <string-name>
            <given-names>Sophia</given-names>
            <surname>Ananiadou</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Enrichment and structuring of archival description metadata</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <source>In Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage</source>
          ,
          <source>Social Sciences, and Humanities</source>
          . ACL.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>