<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploiting Social Media to Address Fundamental Human Rights Issues</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>The Human Rights</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Big Data</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Technology (HRBDT) project</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>based mainly at the Human Rights Centre</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>of the</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>, the Harvard FXB Center for Health and Human Rights</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>. A core activity of one of the four workstreams is to explore and apply the potential of natural language processing to the auto-</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Human Rights, Big Data and Technology</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Massimo Poesio</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Minority Rights Group</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Essex with partners that include the World Health Organisation</institution>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>and the Geneva Academy for International Humanitarian Law and Human Rights</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This invited talk provided an overview of some of our work in relation to extracting meaningful knowledge from social media feeds to help in addressing human rights issues highlighting the potential that the rise of 'big data' offers in this respect looking at both sides of the coin regarding big data and human rights: how big data can help human rights work, but also the potential dangers that can originate from the ability to analyse massive amounts of data very quickly. The primary focus of our work is on applying natural language processing methods to turn large-scale unstructured and partially structured data streams into actionable knowledge.</p>
      </abstract>
      <kwd-group>
        <kwd>Human Rights</kwd>
        <kwd>Social Media</kwd>
        <kwd>NLP</kwd>
        <kwd>Arabic NLP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Vast amounts of social media data are being generated
every second. This represents a paradigm shift in publishing
from largely carefully edited data to user-generated content
which, as a result, has rapidly changed the way people
exchange and consume information as well as how they
communicate. Managing such data streams comes with many
challenges as has been discussed extensively in the research
literature. Nevertheless, it also offers new opportunities.
One such opportunity is the potential to more easily detect
and document human rights violations. In fact, these
developments have already resulted in changes to how human
rights organisations work. The ‘investigator on the ground’
will not be completely replaced but there are many new
modes of identifying evidence of human rights violations.
Social media such as Facebook, YouTube and Twitter are
ideal platforms to push content to the world. Obviously,
there is a big challenge in validating any such postings.
Progress in natural language processing (NLP) means that
off-the-shelf tools can now be used to quickly assemble a
processing pipeline that takes social media data and turns
it into structured knowledge. We are primarily interested
in this type of processing pipeline but that needs to be seen
as part of a bigger picture. Two research projects we are
involved in illustrate the point.
3.</p>
    </sec>
    <sec id="sec-2">
      <title>Knowledge Transfer Partnership</title>
      <p>The second part of our keynote talk focussed on a
practical application of NLP techniques to support human rights
work in a collaboration between the University of Essex
and Minority Rights Group International (MRG)7. This
project is funded by InnovateUK8 through a Knowledge
Transfer Partnership (KTP) project. The aim of this project
is to provide support to civilian-led reporting of human
rights violations, in the context of MRG’s involvement in
the Ceasefire Centre, and in particular in the Ceasefire Iraq
project9. This project complements the objectives of the
more general HRBDT project, exploring the contributions
of big data – and in particular, social media – to the
identification of human rights violations.</p>
      <p>Specific objectives of the collaboration with MRG are, first
of all, to develop a portal that will make it possible to collate
reports of human rights violations sent by civilians using a
variety of formats, from SMS to emails to social media.
The portal10, currently undergoing beta-testing and soon to
go live, will allow personnel by MRG and associated
organizations to view and analyse reports of human right
violations sent by civilians.</p>
      <p>Second, the project aims to develop tools to filter and
analyse this type of information. The analysis techniques
developed so far, and at the moment tested with tweets, include
methods for detecting human rights violations reports using
machine learning-based text categorization to classify text
(e.g., tweets) according to a classification scheme which, in
our case, includes categories such as human right violation
reports (for tweets such as “The army of Assad in
Damascus committed a terrible massacre claiming the lives of
dozens of children in their school”), reporting of general
violence (as in “Four people injured as a result of a brawl
in Darb Alarbaeen”), or reporting of an accident (e.g., “At
7http://minorityrights.org
8https://www.gov.uk/government/
organisations/innovate-uk</p>
      <p>
        9http://minorityrights.org/what-we-do/
ceasefire-project/
10http://iraq.ceasefire.org/
least 24 dead in the sinking of boat for illegal immigrants
off the coast of Istanbul”). As part of the project, a dataset
of over 15,000 Arabic tweets was collected and annotated
according to these categories, and used to train a
classifier to recognize such categories in text
        <xref ref-type="bibr" rid="ref1">(Alhelbawy et al.,
2016)</xref>
        . The objective is to apply classifiers of this type to
filter the data collected through the portal and/or to gather
additional evidence not directly sent to the portal.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>The emergence of ‘big data’ in the form of social media is
affecting all parts of life. This development offers a lot of
new challenges but also opportunities such as the
application of natural language processing techniques to detect and
document human rights violations. NLP tools have matured
to a level that they can easily be applied, are scalable and
robust. This stream of work offers the additional benefit
that it applies state-of-the-art technology to practical
applications that will have a measurable impact on the quality of
life of many people.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>The Human Rights, Big Data and Technology project is
funded by Economic and Social Research Council grant
ES/M010236/1. We also acknowledge support from
InnovateUK through a Knowledge Transfer Partnership (KTP)
project between MRG and the University of Essex,
partnership number 9488.</p>
      <p>5.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Alhelbawy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poesio</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kruschwitz</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Towards a corpus of violence acts in arabic social media</article-title>
          .
          <source>In Proceedings of LREC.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>