<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of CLEF 2019 Lab ProtestNews: Extracting Protests from News in a Cross-context Setting</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ali Hurriyetoglu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erdem Yoruk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Deniz Yuret</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cagr Yoltar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Burak Gurel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>F rat Durusan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Osman Mutlu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arda Akdemir</string-name>
          <email>aakdemirg@ku.edu.tr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Koc University</institution>
          ,
          <addr-line>Istanbul 34450</addr-line>
          ,
          <country country="TR">Turkey</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present an overview of the CLEF-2019 Lab ProtestNews on Extracting Protests from News in the context of generalizable natural language processing. The lab consists of document, sentence, and token level information classi cation and extraction tasks that were referred as task 1, task 2, and task 3 respectively in the scope of this lab. The tasks required the participants to identify protest relevant information from English local news at one or more aforementioned levels in a cross-context setting, which is cross-country in the scope of this lab. The training and development data were collected from India and test data was collected from India and China. The lab attracted 58 teams to participate in the lab. 12 and 9 of these teams submitted results and working notes respectively. We have observed neural networks yield the best results and the performance drops signi cantly for majority of the submissions in the cross-country setting, which is China.</p>
      </abstract>
      <kwd-group>
        <kwd>natural language processing</kwd>
        <kwd>information retrieval</kwd>
        <kwd>ma- chine learning</kwd>
        <kwd>text classi cation</kwd>
        <kwd>information extraction</kwd>
        <kwd>event ex- traction</kwd>
        <kwd>computational social science</kwd>
        <kwd>generalizability</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        We describe a realization of our task set proposal [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] in the scope of CLEF-2019
Lab ProtestNews.1;2 The task set aims at facilitating development of
generalizable natural language processing (NLP) tools that are robust in a cross-context
setting, which is cross-country in this lab. Since the performance of NLP tools
signi cantly drop in a context di erent from the one they are created and
validated [
        <xref ref-type="bibr" rid="ref1 ref2 ref6">1, 2, 6</xref>
        ], measuring and improving state-of-the-art NLP tool development
methodology is the primary aim of our e orts.
      </p>
      <p>
        Comparative social and political science studies facilitate protest
information to analyze cross-country similarities, di erences, and e ect of these actions.
Therefore our lab focuses on classifying and extracting protest event
information in English local news articles from India and China. We believe our e orts
will contribute to enhance the methodologies applied to collect data for these
studies. This need was motivated based on the recent results that shows NLP
tools, those of text classi cation and information extraction, have not been
satisfactory against the requirements of longer time coverage and working on data
from multiple countries [
        <xref ref-type="bibr" rid="ref3 ref7">7, 3</xref>
        ].
      </p>
      <p>This rst iteration of our lab attracted 58 teams from all around the world.
12 of these teams submitted their results to one or more tasks on the CodaLab
page of the lab.3 9 teams described their approach in terms of a working note.</p>
      <p>We introduce the task set we tackle, the corpus we have been creating, and
the evaluation methodology in Sections 2, 3, and 4 respectively. We report the
results in Section 5 and conclude our report in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Task Set</title>
      <p>The lab consists of the tasks document classi cation, event sentence detection
and event extraction, which are referred as task 1, task 2, and task 3 respectively,
as demonstrated in Figure 1. The document classi cation task, which is task 1,
requires predicting whether a news article report at least one protest event that
has happened or is happening. It is a binary classi cation task that require to
predict whether a news article label should be 1 (positive) or 0 (negative). The
sentences that contain any event trigger should be identi ed in task 2, which
is event sentence detection task. Sentence labels are 0 and 1 as well. This task
could be handled either as classi cation or extraction task as we provide order
of the sentences in their respective articles. Finally, the event triggers and event
information, which are place, facility, time, organizer, participant, and target,
should be extracted in task 3. This order of tasks provides a controlled setting
that enables error analysis and optimization possibility during annotation and
tool development e orts. Moreover, this design enable analyzing steps of the
analysis that contributes to explainability of the automated tool predictions.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data</title>
      <p>We provide the number of instances for each task in Table 1 in terms of training,
development, test 1, and test 2 data. The training and development data was
collected from online local English news from India. Test 1 and test 2 data refer
to data from India and China respectively.</p>
      <sec id="sec-3-1">
        <title>3 https://competitions.codalab.org/competitions/22349</title>
        <p>Protests</p>
        <p>Protest sentence(s)</p>
        <p>Event
Participant
Target
Place
Time
...</p>
        <p>A sample from task 1 contains the news article's text, its URL and its label
that is assigned by the annotation team. For task 2: sentences, their labels, their
order in the article's text they belong, and the URL of their article are available
in samples. The release format of the data for task 2 enable participants to
treat this task either as classi cation of individual sentences or extracting event
relevant sentences from a document.</p>
        <p>The data for task 3 consists of snippets that contain one or more event
trigger that refer to the same event. Multiple sentences may occur in a snippet
in case these sentences refer to the same event.4 The tokens in these snippets
are annotated using IOB, inside, outside, beginning, scheme. The examples of
data is provided in Figure 2.</p>
        <p>There is not any overlap of news articles across tasks. This separation was
required in order to avoid any misuse of data from one task to infer the labels
for another task without any e ort.
3.1</p>
        <p>Distribution
We distributed the data set in a way that does not violate copyright of the news
sources. This involves only sharing information that is needed to reproduce the
corpus from the source for task 1 and task 2 and only relevant snippets for task
3. We released a Docker image that contains the toolbox5 required to reproduce
the news articles on the computer of a participant. The toolbox generates a log
of the process that reproduce the data set and we have requested these log les
from the participants. The toolbox is a pipeline that scrapes HTMLs, converts
HTMLs to text and nally performs speci c lling operation for each of task
1 and task 2. To the best of our knowledge, the toolbox succeeded in enabling
participants create the data set on their computers. Only one participant from
4 Snippets we share contain information about only a single event.
5 https://github.com/emerging-welfare/ProtestNews-2019
{"text": "... Police suspect that the panchayat members, including the Salwa
Judum leader, were abducted and killed by Maoist rebels, who had left the
bodies near the village. Meanwhile, security forces and Naxalites had an
encounter near village Belgaon. ...", "label":1}</p>
        <p>A sample from data set for task 1.
{"sentence": "Police suspect that the panchayat members, including the
Salwa Judum leader, were abducted and killed by Maoist rebels, who had</p>
        <p>left the bodies near the village.", "sentence_number":3, "label": 1},
{"sentence": "Meanwhile, security forces and Naxalites had an encounter
near village Belgaon.", "sentence_number":4, "label": 1}</p>
        <p>Corresponding sentence samples for the article above.
including O
the O
Salwa B-target
Judum I-target
leader I-target
, O
were O
abducted B-trigger
and I-trigger
killed I-trigger
by O
Maoist B-participant
rebels I-participant
Annotations in IOB scheme.
Iran was not able to download the news articles due to restrictions to access
online content that are speci c to his geolocation.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Setting</title>
      <p>
        We use macro averaged F1 due to class imbalance present in our data for
evaluating the task 1 and task 2. The event extraction task, which is task 3, was
evaluated on the average F1 score of all information types that was based on the
ratio of the full match between the prediction and the annotations in the test
sets, using a python implementation6 of CoNLL 2003 shared task [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] evaluation
script.
      </p>
      <p>We performed two levels of evaluation that were on data from the source
country (Test 1) and from target country (Test 2). The participants were
informed only about labels of the training and development data from the source
country. They did not see labels of any test set. The number of allowed
submissions and the period the participants can submit their predictions for Test 1
and Test 2 was determined in a way that restrict the possibility of over- tting
on test data. We limited the number of submission and the submission period
in order to make sure the participants do not over t to the test data based.</p>
      <p>We applied three cycles of evaluation. Participants could submit unlimited
number of results without being able to see their score. Their last submitted
results' scores were announced at the end of each evaluation cycle. First and
second cycles aimed at providing feedback to the participants. The third and
nal cycle was the deadline for submitting results.</p>
      <p>Finally, we provided a baseline submission for task 1 and task 2 in order
to guide the participants. This baseline was based on predictions of the the
best scoring machine learning model among Support Vector Machines, Naive</p>
      <sec id="sec-4-1">
        <title>6 https://github.com/sighsmile/conlleval</title>
        <p>Bayes, Rocchio classi er, Ridge Classi er, Perceptron, Passive Aggressive
Classi er, Random Forest, K Nearest Neighbors, and Elastic Net on development
set. The best scoring model was a linear support vector machines classi er that
was trained using stochastic gradient descent.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>We facilitated CodaLab platform for managing the submissions and maintain a
leaderboard.7 The leaderboard for task 1 and task 2 is presented in Table 2. The
column names have the following format Test &lt;task number&gt;-&lt;test number&gt;,
e.g. Test 1-1 stands for the Test 1 of task 1. The results for ProtestLab Baseline
is the aforementioned baseline that was submitted by us.
ASafaya (Sakarya University) submitted the best results for task 2 and for
average of task 1 and task 2 using Bidirectional Gated Recurrrent Unit
(GRU) based model. Although this model perform the best on average, the
performance of this model drops across the context signi cantly.</p>
      <p>PrettyCrocodile (National Research University HSE) submitted the
second best average results that were predicted using Embeddings from
Language Models (ELMo). The performance of the model is comparable in the
cross-context setting for task 2.
7 https://competitions.codalab.org/competitions/22349#results
8 We have not received details of the submissions from CIC-NLP, iAmirSoltani, and
Sayeed Salam. The details of other approaches can be found in the respective working
notes that were published in proceedings of CLEF 2019 Lab ProtestNews.</p>
      <p>LevelUp Research (University of North Carolina at Charlotte) has
applied multi-task learning based on LSTM units using word embeddings from
a pre-trained FastText model. This method yielded the best results for task
1.</p>
      <p>Provos RUG (University of Groningen) has implemented a feature based
stacked ensemble model based on FastText embeddings and a set of di erent
basic Logistic Regression classi ers that enabled their predictions to rank
fourth among the participating teams.</p>
      <p>GS (University of Bremen) has stacked the word embeddings such as GloVe
and FastText together with the contextualized embeddings generated from
Flair language models (LM). This approach was ranked fourth in general
and third for task 2.</p>
      <p>Be-LISI (Universit de Carthage) combined the logistic regression with
linguistic processing and expansion of the text with related terms using word
embedding similarity. This approach marked a signi cant di erence in terms
of overall performance, which is the drop from .64 to .54, in comparison to
higher ranked submissions.</p>
      <p>SSNCSE1 (Sri Sivasubramaniya College of Engineering) reported results
of their bi-directional LSTM that applies Bahdanau, Normed-Bahdanau,
Luong, and Scaled-Luong attentions. The submission that uses Bahdanau
attention yielded the results reported in Table 2.</p>
      <p>SEDA lab (University of Exeter) applied support vector machines and
XGBoost classi ers that are combined with various word embedding approaches.
Results of this submission showed promising performance in terms of
precision on both document and sentence classi cation tasks.</p>
      <p>We analyze task 3 results separate from task 1 and task 2 as it di ers from
them. The F1 scores for task 3 are presented in Table 2.
GS (University of Bremen) submitted the best results for task 3 using a
BiLSTM-CRF model incorporating pooled contextualized air embeddings,
and their model was the best in generalizing.</p>
      <p>DeepNEAT (FloodTags &amp; Radboud University) compares the submitted
ELMO+BiLSTM model to a traditional CRF and shows that the former is
better and more generalizable.</p>
      <p>Provos RUG (University of Groningen) divides the task 3 into two as event
trigger detection task and event argument detection task using
BiLSTMCRF model with word embeddings, POS embeddings, and character-level
embeddings for both subtasks. He further extends the features for latter
subtask with learned embeddings for dependency relations and event
triggers.</p>
      <p>PrettyCrocodile (National Research University HSE) makes use of ELMO
embeddings with di erent architectures, achieving her best score for task 3
using a BiLSTM.</p>
      <p>LevelUp Research (University of North Carolina at Charlotte) implemented
a multi-task neural model that require a time-ordered sequence of word
vectors representative of a document or sentence. The LSTM layer has been
replaced by a layer of bidirectional gated recurrent units (GRU).
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>The results show how text classi cation and information extraction tool
performances drops between two contexts. The scores on data from the target country
are signi cantly lower than on data from the source country. Only the
PrettyCrocodile team performed comparatively well across contexts for task 2.
Although it is not the best scoring system for neither task 1 nor task 2,
PrettyCrocodile team's approach show some promise toward tackling the
generalizability of NLP tools.</p>
      <p>The generalization of automated tools is an issue that has recently attracted
much attention.9 However, as we have determined in our lab, generalizability
is still a challenge for state-of-the-art methodology. Consequently, we will
continue our e orts by repeating this practice and extending the data and will be
adding data from new countries and languages to our setting. The next
iteration will run in the scope of the Workshop on Challenges and Opportunities in
Automated Coding of COntentious Political Events (Cope 2019) at European
Symposium Series on Societal Challenges in Computational Social Science (Euro
CSS 2019).10;11;12</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work is funded by the European Research Council (ERC) Starting Grant
714868 awarded to Dr. Erdem Yoruk for his project Emerging Welfare. We are
grateful to our steering committee members for the CLEF 2019 lab Sophia
Ananiadou, Antal van den Bosch, Kemal O azer, Arzucan O zgur, Aline
Villavicencio, and Hristo Tanev. Finally, we thank to Theresa Gessler and Peter Makarov
9 https://sites.google.com/view/icml2019-generalization/cfp
10 https://competitions.codalab.org/competitions/22842
11
https://emw.ku.edu.tr/?event=challenges-and-opportunities-in-automated-codingof-contentious-political-events&amp;event date=2019-09-02
12 http://symposium.computationalsocialscience.eu/2019/
for their contribution in organizing the CLEF lab by reviewing the annotation
manuals and sharing their work with us respectively.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Akdemir</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Hurriyetoglu,
          <string-name>
            <surname>A.</surname>
          </string-name>
          , Yoruk, E., Gurel,
          <string-name>
            <surname>B.</surname>
          </string-name>
          , Yoltar,
          <string-name>
            <surname>c.</surname>
          </string-name>
          , Yuret, D.:
          <article-title>Towards Generalizable Place Name Recognition Systems: Analysis and Enhancement of NER Systems on English News from India</article-title>
          .
          <source>In: Proceedings of the 12th Workshop on Geographic Information Retrieval</source>
          . pp.
          <volume>8</volume>
          :
          <issue>1</issue>
          {8:
          <fpage>10</fpage>
          . GIR'18,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2018</year>
          ). https://doi.org/10.1145/3281354.3281363, http://doi.acm.
          <source>org/10</source>
          .1145/3281354.3281363
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ettinger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daume</surname>
            <given-names>III</given-names>
          </string-name>
          , H.,
          <string-name>
            <surname>Bender</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          :
          <source>Towards Linguistically Generalizable NLP Systems: A Workshop</source>
          and Shared Task.
          <source>In: Proceedings of the First Workshop on Building Linguistically Generalizable NLP Systems</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>10</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2017</year>
          ), http://aclweb.org/anthology/W17-5401
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hammond</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weidmann</surname>
            ,
            <given-names>N.B.:</given-names>
          </string-name>
          <article-title>Using machine-coded event data for the micro-level study of political violence</article-title>
          .
          <source>Research &amp; Politics</source>
          <volume>1</volume>
          (
          <issue>2</issue>
          ),
          <volume>2053168014539924</volume>
          (
          <year>2014</year>
          ). https://doi.org/10.1177/2053168014539924, https://doi.org/10.1177/2053168014539924
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Hurriyetoglu,
          <string-name>
            <surname>A.</surname>
          </string-name>
          , Yoruk, E., Yuret,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Yoltar</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          , Gurel,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Durusan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Mutlu</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.:</surname>
          </string-name>
          <article-title>A task set proposal for automatic protest information collection across multiple countries</article-title>
          . In: Azzopardi,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Fuhr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Mayr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Hau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Hiemstra</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.) Advances in Information Retrieval. pp.
          <volume>316</volume>
          {
          <fpage>323</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Sang</surname>
            ,
            <given-names>E.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Meulder</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Introduction to the conll-2003 shared task: Languageindependent named entity recognition</article-title>
          .
          <source>arXiv preprint cs/0306050</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Soboro</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuhr</surname>
          </string-name>
          ,
          <source>N.: Report on GLARE 2018: 1st Workshop on Generalization in Information Retrieval: Can We Predict Performance in New Domains? SIGIR Forum</source>
          <volume>52</volume>
          (
          <issue>2</issue>
          ),
          <volume>132</volume>
          {
          <fpage>137</fpage>
          (
          <year>2018</year>
          ), http://sigir.org/wpcontent/uploads/2019/01/p132.pdf
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kennedy</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lazer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramakrishnan</surname>
          </string-name>
          , N.:
          <article-title>Growing pains for global monitoring of societal events</article-title>
          .
          <source>Science</source>
          <volume>353</volume>
          (
          <issue>6307</issue>
          ),
          <volume>1502</volume>
          {
          <fpage>1503</fpage>
          (
          <year>2016</year>
          ). https://doi.org/10.1126/science.aaf6758, http://science.sciencemag.org/content/353/6307/1502
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>