<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thomas Ploeger</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maxine Kruijt</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lora Aroyo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frank de Bakker</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iina Hellsten</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antske Fokkens</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jesper Hoeksema</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Serge ter Braake</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Activists have a signi cant role in shaping social views and opinions. Social scientists study the events activists are involved in order to nd out how activists shape our views. Unfortunately, individual sources may present incomplete, incorrect, or biased event descriptions. We present a method where we automatically extract event mentions from di erent news sources that could complement, contradict, or verify each other. The method makes use of o -the-shelf NLP tools. It is therefore easy to setup and can also be applied to extract events that are not related to activism.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Computer Science Department VU University Amsterdam 2 Organization Sciences Department VU University Amsterdam 3 Language and Communication Department VU University Amsterdam 4 History Department VU University Amsterdam</title>
      <sec id="sec-1-1">
        <title>Introduction</title>
        <p>
          The goal of an activist is to e ect change in societal norms and standards [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
Activists can thus both make an impact on the present and play a signi cant
role in shaping the future. Considering the events activists are engaged in allows
us to see current controversial issues, and gives social scientists the means to
identify the (series of) methods through which activists are trying to achieve
change in society.
        </p>
        <p>The MONA5 project is an interdisciplinary social/computer science e ort
which aims at producing a visual analytics suite for e ciently making sense of
large amounts of activist events. Speci cally, we intend to enable the discovery of
event activity patterns that are `hidden' in human-readable text, as well as
provide detailed analyses of these patterns. The project currently focuses on activist
organizations that have recently been protesting against petroleum exploration
in the Arctic.</p>
        <p>Social scientist are interested in nding out which activist organizations are
trying to in uence the oil giants, and speci cally which events they are
organizing to do so. This could be addressed by aggregating events that took place in
this context, enabling a quantitative (e.g. \What is the common type of event
organized?") as well as a qualitative (e.g. \Why are these types of events
organized?") analysis.</p>
        <p>
          In an earlier paper [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], we described initial work in the MONA project.
This work primarily concerned the evolutionary explorations we performed in
the activist use case to make our event modeling requirements concrete. These
explorations led to the decision to use the Simple Event Model (SEM) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which
models events as \who did what to whom, where and when". In addition, we
considered how visualizations of event data could aid an end-user in answering
speci c types of questions about aggregated activist events.
        </p>
        <p>This paper describes our approach for event extraction from human-readable
text so we can aggregate them and `feed' them to a visualization suite. Our
approach repurposes o -the-shelf natural language processing software and services
(primarily named entity recognizers and disambiguators) to automatically
extract events from news articles with a minimal amount of domain-speci c tuning.
As such, the method described in this paper goes beyond the domain of activism
and can be used to extract events related to other topics as well.</p>
        <p>
          The output of our system are representations in the Grounded Annotation
Framework (GAF) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], which links representations in SEM to the text and
linguistic analyses they are derived from. A more detailed description of GAF will
be given in Section 3.2.
        </p>
        <p>We use news articles because they are available from a huge variety of sources
and in increasingly large numbers. Being able to tap into such a large and
diverse source of event descriptions is extremely valuable in event-based research,
because individual event descriptions may be incomplete, incorrect, out of
context, or biased. These problems could be alleviated by using multiple sources
and increasing the number of descriptions considered: Events extracted from
multiple articles could complement each other in terms of completeness, serve as
veri cation of correctness, place events in a larger context, and present multiple
perspectives.</p>
        <p>We consider both quantitative measurements and the usefulness of the
extracted events in our evaluation. We quantitatively evaluate performance by
calculating the traditional information retrieval metrics of precision, recall, and
F1 for the recognition of events and their properties. Through examples, we give
a tentative impression of the usefulness of the aggregated event data.</p>
        <p>The rest of this paper is structured as follows. In Section 2, we give an
overview of previous work in event extraction and how it relates to this work.
The representation frameworks we use are explained in Section 3. In Section 5,
we outline our methodology. We show how we model events, how events are
typically described in text and how we use existing NLP software and services
to extract them. Section 5 contains an overview of the results. We present both
a quantitative evaluation as well as a detailed error analysis of the performance
of our event extraction method. We go beyond performance numbers in Section
6 by discussing the usability and value of our contribution leading us to the
direction future work should take.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Related work</title>
        <p>In this section, we demonstrate the heterogeneous nature of the eld of event
extraction by giving a non-exhaustive overview of contemporary approaches from
several domains. The diversity in event representations and extraction methods
makes it inappropriate to make direct comparisons (e.g. in terms of performance)
between our work and that of others, but we can still show how work in other
domains relates to our own work.</p>
        <p>
          In molecular biology, gene and protein interactions are described in
humanreadable language in scienti c papers. Researchers have been working on
methods for extracting and aggregating these events to help understand the large
numbers of interactions that are published. For example, Bjorne [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
demonstrated a modular event extraction pipeline that uses domain-speci c modules
(such as a biomedical named entity recognizer) as well as general purpose NLP
modules to extracted a prede ned set of interaction events from a corpus of
PubMed papers.
        </p>
        <p>
          The European border security agency Frontex uses an event extraction
system [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] to extract events related to border security from online news articles.
Online news articles are used because they are published quickly, have
information that might not be available from other sources, and facilitate cross-checking
of information. This makes them valuable resources in the real-time monitoring
of illegal migration and cross-border crime. The system developed for Frontex
uses a combination of traditional NLP tools and pattern matching algorithms to
extract a limited set of border security events such as illegal migration,
smuggling, and human tra cking.
        </p>
        <p>
          Van Oorschot et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] extract game events (e.g. goals, fouls) from tweets
about football matches to automatically generate match summaries. Events were
detected by considering increases in tweet volume over time. The events in those
tweets were classi ed using a machine learning approach, using the presence
of certain words, hyperlinks, and user mentions as features. There is a limited
set of events that can occur during a football match, so there is a pre-de ned,
exhaustive list of events to extract. These events have two attributes: The time
at which they occurred and the football team that was responsible.
        </p>
        <p>The recurring theme in event extraction across di erent domains is the desire
to extract events from human-readable text (as opposed to structured data) to
aggregate them, enabling quantitative and qualitative analysis. Our research has
the same intentions, but the domain-speci c nature of event representations and
extraction methods in the current event extraction literature limits the reuse of
methods across domains and (to our knowledge) there has been no research into
extracting events for the purpose of studying activists.</p>
        <p>Speci cally, the existing work on event extraction is typically able to take
advantage of an exhaustive lists of well-de ned events created a priori. In our
case, we cannot make any assumptions about which types of events are relevant
to the end user because we intend to facilitate discovery of new event patterns,
which necessitates a minimally constrained de nition of `event'.</p>
        <p>
          Ritter et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] present an open-domain approach to extract events from
twitter. They use supervised and semi-supervised machine learning training a
model on 1,000 annotated tweets. Due to the di erence in structure and
language use, this corpus is not suitable for extracting events from newspaper text.
Moreover, tweets will generally address only one event whereas newspaper
articles can also be stories that involve sequences of events. This makes our task
rather di erent from the one addressed in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
        <p>The goal of our research was to create an approach that can identify events
in newspaper text while exclusively making use of o -the-shelf NLP tools. We
do not make use of a prede ned list of potentially interesting events like most
of the approaches mentioned above. Our approach di ers from Ritter et al.'s
work, because there is no need to annotate events in text for training. Our
approach, which will be described in the following section, can be applied for
event extraction in any domain.
3</p>
      </sec>
      <sec id="sec-1-3">
        <title>Event Representation</title>
        <p>
          In this section, we describe the representations we use as output of our system.
We rst outline the Simple Event Model in Section 3.1. This is followed by an
explanation of the Grounded Annotation Framework (GAF) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] which forms the
overall output of our extraction system in Section 3.2.
3.1
        </p>
        <p>The Simple Event Model
We use the Simple Event Model (SEM) to represent events. SEM uses a graph
model de ned using the Resource Description Framework Schema language (RDFS)
and the Web Ontology Language (OWL). SEM is designed around the
following de nition of event. \Events [..] encompass everything that happens, even
ctional events. Whether there is a specic place or time or whether these are
known is optional. It does not matter whether there are specic actors involved.
Neither does it matter whether there is consensus about the characteristics of
the event." This de nition leads to a more formal speci cation in the form of an
event ontology which models events as having actors, places and times(tamps).
Each of these classes may have a type, which may be speci ed by a foreign type
system. A unique feature of SEM is that it allows specifying multiple views on
a certain event, which hold according to a certain authority. A basic example of
an instantiated SEM-event can be seen in Figure 1.
3.2</p>
        <p>
          The Grounded Annotation Framework
In addition to SEM, we use the Grounded Annotation Framework (GAF). The
basic idea behind this framework is that it links semantic representations to
mentions of these representations in text and semantic relations to the syntactic
relations they are derived from. This provides a straight-forward way to mark the
provenance of information using the PROV-O [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. When presenting multiple
        </p>
        <p>dbpedia:
NonProfitOrg
rdf:type
freebase:
Protest
rdf:type
dbpedia:
Company
rdf:type
sem:Actor
geonames:p</p>
        <p>sem:Place
rdf:type
sem:hasTimeStamp
rdf:type
sem:hasActor</p>
        <p>sem:hasActor sem:hasPlace
sem:Event
dbpedia:Royal_
Dutch_Shell
dbpedia:Greenpeace
geonames:
2643743</p>
        <p>Tuesday
sem:eventType
sem:actorType rdf:type rdf:type
sem:actorType
sem:placeType rdf:type
sem:eventType
sem:actorType
sem:actorType
sem:placeType
views next to each other, it is important to know where these views come from.
Furthermore, Natural Language Processing techniques do not yield perfect
results. It is thus essential that social scientists can easily verify whether extracted
information was indeed expressed in the original source. Finally, insight into the
derivation process can be valuable for system designers as they aim to improve
their results.
4</p>
      </sec>
      <sec id="sec-1-4">
        <title>Method</title>
        <p>As establised in the previous section, we consider everything that happens an
event. An event may have actors involved, a certain location, and occurs at a
point in time. We use a rapidly prototyped event extraction tool which integrates
several generic, o -the-shelf natural language processing software packages and
Web services in a pipeline to extract this information. This section describes this
pipeline which is illustrated in Figure 2.</p>
        <p>Preprocessing &amp; Metadata extraction The pipeline takes a news article's
URL as input, with which we download the article's raw HTML. We use the
Nokogiri6 XML-parser to nd time and meta tags in the HTML. These tags
typically contain the article's publication date, which we need later for date
normalization. Next, we use AlchemyAPI's7 author extraction service on the
raw HTML to identify the article's author, which enables us to attribute the
extracted events. We then run the HTML through AlchemyAPI's text extraction
service to strip any irrelevant content from the HTML, giving us just the text
of the article.
6 http://nokogiri.org/</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>7 http://www.alchemyapi.com/</title>
      <p>1
2
3</p>
      <sec id="sec-2-1">
        <title>Tuesday Greenpeace protested against Shell in London</title>
      </sec>
      <sec id="sec-2-2">
        <title>Tuesday Greenpeace protested against Shell in</title>
        <p>DATE ORG ORG</p>
      </sec>
      <sec id="sec-2-3">
        <title>TIME ACTOR ACTOR</title>
      </sec>
      <sec id="sec-2-4">
        <title>London</title>
        <p>LOC</p>
      </sec>
      <sec id="sec-2-5">
        <title>PLACE</title>
        <p>Dependency Parsing</p>
      </sec>
      <sec id="sec-2-6">
        <title>Tuesday Greenpeace protested against Shell in</title>
      </sec>
      <sec id="sec-2-7">
        <title>London</title>
        <p>NN</p>
        <p>
          NSUBJ
Processing The article's text is split into sentences and words using Stanford's
sentence splitter and word tokenizer8. We consider each verb of the sentence to
be an event, because verbs convey actions, occurrences, and states of being. This
is a very greedy approach, but this is necessarily so: We do not wish to make
any a priori assumptions about which types of events are relevant to the end
user. We use Stanford's part-of-speech tagger [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] to spot the verbs.
        </p>
        <p>
          Actors and places are discovered using Stanford's named entity recognizer [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
The type (e.g. person, organization, location) of the named entity determines
whether it is an Actor or a Place. Dates and times are also identi ed by the
named entity recognizer.
        </p>
        <p>
          The mere existence of named entities, a timestamp, and a verb in the same
sentence does not immediately mean that they together form one event. One
sentence may describe multiple events or a place might be mentioned without it
being the direct location of the event. Therefore we only consider named entities
and timestamps grammatically dependent on a speci c event to be part of that
event. For this we use Stanford's dependency parser [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
        </p>
        <p>
          Normalization &amp; Disambiguation Using Stanford's SUTime [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], We
normalize any relative timestamps (e.g. \Last Tuesday") to the publication date to
transform them into full dates (e.g. \23-06-2013"). We complement Stanford's
named entity recognizer with TextRazor's9 API to disambiguate found named
entities to a single canonical entity in an external data source such as DBpedia.
Storage &amp; Export The output of the preprocessing, metadata extraction,
processing, normalization, and disambiguation steps is stored in a Neo4j10 graph
database. For each article, we create a document node with metadata properties,
such as the URL, author, and publication date. The document node has sentence
nodes as its children, which in turn have word nodes as their children. The word
nodes have the properties that were identi ed earlier in the pipeline, such as
their part-of-speech tags, named entity tags, etc. The grammatical dependencies
between words are expressed as typed edges between word nodes. We traverse the
resulting graph to identify verbs with dependent named entities and timestamps.
We export the event as a SEM event together with provenance in GAF.
Implementation details All of the software packages and services above are
integrated using several custom Ruby scripts. We have also used several
existing Ruby gems for various supporting tasks: A Ruby wrapper11 for Stanford's
NLP tools, HTTParty12 for Web API wrappers, Chronic13 for date parsing, and
Neography14 for interacting with Neo4j.
8 nlp.stanford.edu/software/tokenizer.shtml12 http://github.com/jnunemaker/httparty
9 http://www.textrazor.com/ 13 http://github.com/mojombo/chronic
10 http://www.neo4j.org/ 14 http://github.com/maxdemarzi/neography
11
http://github.com/louismullie/stanford
        </p>
        <p>core-nlp</p>
        <sec id="sec-2-7-1">
          <title>Evaluation</title>
          <p>Before we present the results of our method of event extraction in Section 5.2,
we describe the corpus we used for evaluation and the creation of a gold
standard in Section 5.1. In Section 5.3, we describe the major issues impacting the
performance of our method.
5.1</p>
          <p>Experimental Setup
We extracted events from a corpus of 45 documents concerning arctic oil
exploration activism. 15 of these documents are blog posts, the other 30 are news
articles. The majority of articles are from The New York Times15 (70%) and the
Guardian16 (15%), the rest from similar news websites.</p>
          <p>Three domain experts manually annotated every article (each annotator
individually annotated 1/3 of the corpus) to create a gold standard for evaluation.
The experts were asked to annotate the articles with events, actors, places, and
times and then link the actors, places, and times to the appropriate events, in
such a way that the resulting events would be useful for them if aggregated and
visualized. No further explicit instructions were given to the annotators. The
Brat rapid annotation tool17 was used by the experts for annotation.</p>
          <p>Table 1 illustrates the inter-rater agreement of the annotators on a subset
of the corpus that was annotated by each annotator. For each type of
annotation we show the percentage of annotations that were annotated by only 1 of the
annotators, by 2 of the annotators, or by all 3 annotators. For each class the
majority of annotations are shared by at least 2 annotators. Events have the largest
amount of single-annotator annotations, showing that inter-rater consensus is
lowest for this concept.
The second and third columns of Table 2 show the amounts of events, actors,
places, and times in the gold standard and the amounts extracted from the
corpus. The next 3 columns show the true positives, false positives, and false
15 http://www.nytimes.com/
16 http://www.guardian.co.uk/
17 http://brat.nlplab.org/
negatives. The nal 3 columns show the resulting precision, recall, and F1 per
class.</p>
          <p>For each of the 1299 events correctly recognized, we checked if they were
associated with the correct actors, places, and times. Table 3 shows the mean
precision, recall, and F1 scores for the linking of events to the appropriate actors,
places, and times.
We carried out an error analysis for each class and identi ed several issues that
bring down performance of our system. This section describes these errors and
indicates how we may improve our system in future work.</p>
          <p>Actors masquerading as places (and vice versa) In the sentence \Shell
is working with wary United States regulators.", our annotators are interested
in the United States as an actor, not a location. Still, it is recognized as a
location by the named entity recognizer. This is a contributor to the large number
of false negatives (and false positives) for actors and places. The grammatical
dependency between the verb and a named entity could give us some clues to the
role an entity plays in an event. In the example, the kind of preposition (\with")
makes it clear that United States indicates an actor, not a place.
Ambiguous times The named entity recognizer only identi es expressions that
contain speci c time indications as times. Relative timestamps such as \last
Tuesday" or \next winter" are resolvable by the extraction pipeline, but more
ambiguous times such as \after" or \previously" and conditional times such
as \if" and \when" are not detected. This contributes to the false negatives
for timestamps and could be solved by hand-coding a list of such temporal
expressions into the extraction process.</p>
          <p>Unnamed actors &amp; places The pipeline only recognizes named entities as
actors and places, so any common nouns or pronouns that indicate actors are not
recognized by the pipeline. This issue could be solved by relaxing the restriction
that only named entities are considered for actors and places. Similar to the
actors masquerading as places, looking at the grammatical dependencies could
indicate whether we are dealing with an actor or a place. This may however
increase the number of false positives because of the ambiguous nature of some
grammatical dependencies (e.g. \about"). We propose two tactics to address this
issue: coreference resolution and linking noun phrases to ontologies.</p>
          <p>
            Consider the following 2 sentences: \The Kulluk Oil Rig was used for test
drilling last summer. The Coast Guard ew over the rig for a visual inspection."
A coreference resolver in the pipeline could indicate that \the rig" in the second
sentence is a coreferent of a named entity and may thus be considered a location.
Sometimes, actors or places do not refer to a speci c person or location (e.g.
\scientists", \an area") in which case they will not corefer to a named entity. If
we link noun phrases to an ontology such as WordNet [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ], we can identify whether
they refer to a potential agent or location by inspecting their hyponyms. Because
nouns can also refer to events (e.g. \strike"), this may also increase recall on event
detection.
          </p>
          <p>Gold Standard Annotations The percentages of inter-rater agreement (as
shown earlier in Table 3), especially for events, indicate that the gold standard
could bene t from a more rigorous annotation task description. We realize that if
the task is loosely de ned, human annotators may have di erent interpretations
of what an `event' is in natural language.</p>
          <p>For this reason, it is interesting to compare the tool output to the three
annotators individually. Table 4 shows the pipeline's F1-scores per class per
individual annotator. The scores for annotator 1 and 3 are very close for all four
classes. Annotator 2 di ers signi cantly for places and times. This demonstrates
the variance that annotators with di erent interpretations of the annotation task
introduce to performance scores of the tool.</p>
        </sec>
        <sec id="sec-2-7-2">
          <title>Conclusion</title>
          <p>In this paper we reported on the development and performance of our extraction
method for activist events: A pipeline of existing NLP software and services with
minimal domain-speci c tuning. The greatest value of this contribution is the
fact that it will enable further work in the MONA project. The goal of the
project is to produce a visual analytics suite for e ciently making sense of large
amounts of activist events. Through these visual analytics, we intend to enable
the discovery and detailed analysis of patterns in event data. The extraction
pipeline described in this paper (and any future revisions of it) will be able to
feed our visual analytics suite with event data.</p>
          <p>Work is already underway on the development of the visual analytics suite
and details will be available in a forthcoming paper. The e ectiveness of the
visual analytics will be dependent on the quality of the event data our extraction
pipeline produces. We already have candidate solutions for issues that negatively
impact the pipeline's performance. In future work we will implement these
solutions and report on their e ectiveness. In the meantime, we can already get a
tentative impression of the value the extracted event data has, for both discovery
and more detailed analysis.</p>
          <p>Aggregating and counting event types that a certain actor is involved in
enables the discovery of the primary role of actors. Similarly, by aggregating and
counting the places of events we can discover the geographical areas an actor has
been active in. Filtering the events by time can give us insight into changes in
active areas over time. Because we have extracted events from multiple sources,
events can complement each other in terms of completeness, serve as veri cation
of correctness, place events in a larger context, and present multiple perspectives.
In future work, we intend to de ne measurements for these concepts (e.g. when
are events complementary, when do they verify each other) in order to quantify
them.</p>
        </sec>
        <sec id="sec-2-7-3">
          <title>Acknowledgements</title>
          <p>This research is partially funded by the Royal Netherlands Academy of Arts and
Sciences in the context of the Network Institute research collaboration between
Computer Science and other departments at VU University Amsterdam, and
partially by the EU FP7-ICT-2011-8 project NewsReader (316404). We would
like to thank Chun Fei Lung, Willem van Hage, and Marieke van Erp for their
contributions and advice.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Atkinson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piskorski</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goot</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yangarber</surname>
          </string-name>
          , R.:
          <article-title>Multilingual real-time event extraction for border security intelligence gathering</article-title>
          . In: Wiil, U.K. (ed.)
          <source>Counterterrorism and Open Source Intelligence, Lecture Notes in Social Networks</source>
          , vol.
          <volume>2</volume>
          , pp.
          <volume>355</volume>
          {
          <fpage>390</fpage>
          . Springer Vienna (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Bjorne, J.,
          <string-name>
            <surname>Van Landeghem</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pyysalo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohta</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ginter</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van de Peer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakoski</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Pubmed-scale event extraction for post-translational modi cations, epigenetics and protein structural relations</article-title>
          .
          <source>In: Proceedings of BioNLP 2012</source>
          . pp.
          <volume>82</volume>
          {
          <issue>90</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <issue>3</issue>
          .
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>A.X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Sutime: A library for recognizing and normalizing time expressions</article-title>
          . In: et al., N.C. (ed.)
          <source>Proceedings of LREC 2012. ELRA</source>
          , Istanbul, Turkey (may
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Fellbaum</surname>
          </string-name>
          , C. (ed.):
          <article-title>WordNet: An Electronic Lexical Database</article-title>
          . MIT Press, Cambridge, MA (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Finkel</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grenager</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Incorporating non-local information into information extraction systems by gibbs sampling</article-title>
          .
          <source>In: Proceedings of the 43rd ACL</source>
          . pp.
          <volume>363</volume>
          {
          <fpage>370</fpage>
          . ACL '05,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          , Stroudsburg, PA, USA (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Fokkens</surname>
            , A., van Erp,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vossen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tonelli</surname>
            , S., van Hage,
            <given-names>W.R.</given-names>
          </string-name>
          , Sera ni, L.,
          <string-name>
            <surname>Sprugnoli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoeksema</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>GAF: A grounded annotation framework for events</article-title>
          .
          <source>In: Proceedings of the rst Workshop</source>
          on Events: De nition, Dectection, Coreference and Representation. Atlanta, USA (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. van Hage,
          <string-name>
            <given-names>W.R.</given-names>
            ,
            <surname>Malaise</surname>
          </string-name>
          , V., van Erp,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Schreiber</surname>
          </string-name>
          , G.:
          <article-title>Linked Open Piracy</article-title>
          .
          <source>In: Proceedings of the sixth international conference on Knowledge capture</source>
          . pp.
          <volume>167</volume>
          {
          <fpage>168</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>June 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. den Hond, F.,
          <string-name>
            <surname>de Bakker</surname>
            ,
            <given-names>F.G.A.</given-names>
          </string-name>
          :
          <article-title>Ideologically motivated activism: How activist groups in uence corporate social change activities</article-title>
          .
          <source>Academy of Management Review</source>
          <volume>32</volume>
          (
          <issue>3</issue>
          ),
          <volume>901</volume>
          {
          <fpage>924</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Accurate unlexicalized parsing</article-title>
          .
          <source>In: Proceedings of the 41st ACL</source>
          . pp.
          <volume>423</volume>
          {
          <fpage>430</fpage>
          . ACL '03,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          , Stroudsburg, PA, USA (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Moreau</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Missier</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belhajjame</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>B'Far</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cheney</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coppens</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cresswell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gil</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klyne</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lebo</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCusker</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miles</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Myers</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sahoo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tilmes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>PROV-DM: The PROV Data Model</article-title>
          .
          <source>Tech. rep., W3C</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. van Oorschot, G., van Erp,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Dijkshoorn</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Automatic extraction of soccer game events from twitter</article-title>
          .
          <source>In: Proceedings of the Workhop on Detection, Representation, and Exploitation of Events in the Semantic Web (DeRiVE</source>
          <year>2012</year>
          ). pp.
          <volume>21</volume>
          {
          <issue>30</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Ploeger</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Armenta</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aroyo</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>de Bakker</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellsten</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Making sense of the arab revolution and occupy: Visual analytics to understand events</article-title>
          .
          <source>In: Proceedings of the Workhop on Detection, Representation, and Exploitation of Events in the Semantic Web (DeRiVE</source>
          <year>2012</year>
          ). pp.
          <volume>61</volume>
          {
          <issue>70</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ritter</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Open domain event extraction from twitter</article-title>
          .
          <source>In: Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          . pp.
          <volume>1104</volume>
          {
          <fpage>1112</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Feature-rich part-of-speech tagging with a cyclic dependency network</article-title>
          .
          <source>In: Proceedings of the NAACL and HLT 2003</source>
          . pp.
          <volume>173</volume>
          {
          <fpage>180</fpage>
          . NAACL '03,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          , Stroudsburg, PA, USA (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>