<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Q u e r y C o n d i t i o n
F r o m</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Comparative Analysis of GDELT Data Using the News Site Contrast System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Masaharu Yoshioka</string-name>
          <email>yoshioka@ist.hokudai.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Noriko Kando</string-name>
          <email>kando@nii.ac.jp</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Copyright c by the paper's authors. Copying permitted for</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hokkaido University</institution>
          ,
          <addr-line>N14 W9, Kita-ku, Sapporo-shi, Hokkaido, 060-0814</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Institute of Informatics</institution>
          ,
          <addr-line>2-1-2, Hitotsubashi, Chiyoda-ku, Tokyo, 101-8430</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>private and academic purposes. This volume is published and</institution>
          ,
          <addr-line>copyrighted by its editors., In: M. Martinez, U. Kruschwitz, G. Kazai, D. Corney, F. Hopfgartner, R. Campos</addr-line>
          ,
          <institution>and D. Albakour (eds.): Proceedings of the, NewsIR'16 Workshop at ECIR</institution>
          ,
          <addr-line>Padua, Italy, 20-March-2016, published at http://ceur-ws.org</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <volume>2</volume>
      <issue>0</issue>
      <abstract>
        <p>The News Site Contrast (NSContrast) system analyzes news articles retrieved from multiple news sites based on the concept of contrast set mining. It can extract terms that characterize different topics of interest across news sites, countries, and regions. In this study, we used NSContrast to analyze Global Database of Events, Language, and Tone (GDELT) data by comparing news articles from different regions (e.g., USA, Asia, and the Middle East). We also present examples of analyses performed using this system.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>named entity, polarity, news site, and country) to
analyze news articles. The interface supports the
construction of structured queries that use one
or more facets, where the facet information can
be represented using various styles (e.g., time
sequence graph, table, or bar chart). (Figure 2)
2.2 Data conversion
To apply NSContrast to the analysis of GDELT data,
it was necessary to convert the GDELT data into news
article data. There are two databases in GDELT:
GDELT Event and GDELT Global Knowledge Graph
(GKG). GDELT GKG is a database based on a raw
output format of the original news articles for
constructing the GDELT Event database. Because the
GDELT Event database does not have detailed
original news article sources, GDELT GKG was used for
NSContrast.</p>
      <p>GDELT GKG was constructed by extracting the
following metadata information from the original
news articles: DATE, THEMES, LOCATIONS,
PERSONS, ORGANIZATIONS, TONE (as a real value;
0 means neutral), CAMEOEVENTIDS (references
to the GDELT Event database), SOURCES, and
SOURCEURLS. When there are two or more
articles that share all name sets (THEMES, LOCATIONS,
PERSONS, and ORGANIZATIONS), those news
articles are aggregated as one datum and SOURCES
and SOURCEURLS have multiple entries. Example
of SOURCES and SOURCEURLS information for one
datum in January 19, 2016 are shown below.
SOURCES punchng.com; punchng.com;
onlinenigeria.com; onlinenigeria.com
SOURCEURLS http://www.punchng.com/25909-2/,
http://www.punchng.com/i-amresolved-to-better-lagos-ambode/,
http://news2.onlinenigeria.com/news/general/
453949-i-am-resolved-to-betterlagos-%E2%80%93w-ambode.html,
http://news2.onlinenigeria.com/news/general/
453949-i-am-resolved-to-better-lagos-ambode.html
Two types of multiple SOURCEURLS are shown
above. In one, almost the same content has a different
URL for the same news site (the rst two URLs and
the last two URLs above) and the other is a
different URL with different news sites (the rst and third
URLs).</p>
      <p>Most of the former cases are simply URL variations
of the same content; e.g., the rst URL is redirected to
the second URL and the third URL is a variation of the
fourth URL (the URL encoding of \%E2%80%93w"
is \-" for UTF-8). It is better to select one of them
for deduplication. The latter cases are meaningful for
representing the importance of the contents, because
different news sites have selected the same content for
their sites.</p>
      <p>By using these metadata, the following information
was constructed for NSContrast.</p>
      <p>Date Date of the article.</p>
      <p>Person, Organization, Location Lists of people,
organizations, and locations extracted from the
article using the GDELT GKG.</p>
      <p>Polarity We classi ed articles into three types
(positive, negative, and neutral) to simplify the
analysis of the polarity information. The tone extracted
by the GDELT GKG was used for classi cation
(tone &gt; 1: positive; tone &lt; 1: negative; other:
neutral).</p>
      <p>Site Site information extracted from GDELT GKG.</p>
      <p>To count the number of articles from different
news sites, we duplicate one datum for each site.
However, if there are two or more entries for the
same news site information, one of these entries
is used for deduplication. In the above example,
one datum is duplicated for \punchng.com" and
\onlinenigeria.com."
URL URL for the original news article. When
there is one URL for a site, the
corresponding URL is used for each site. However, when
there are two or more URLs for a given news
site, the shortest URL is selected for each news
site (e.g., http://www.punchng.com/25909-2/ for
punchng.com).</p>
      <p>SiteCountry We constructed a database of news
sites to identify their countries of origin. We used
http://www.world-newspapers.com/ to extract
these relationships. For \BBC monitoring," we
used \United Kingdom" as the site country for the
news site. In addition, if news sites used country
code top-level domains (e.g., .jp for Japan), we
used this domain information to estimate the site
country. Finally we used a geolocation service 3 to
estimate the site country by using the IP address
of the top domain. However, the country was left
blank if we could not obtain appropriate location
information from the geolocation service.</p>
      <p>SiteRegion Countries were grouped into the
following eight regions: USA, Asia, Europe, Middle
East, Africa, Oceania, North America (excluding
USA), and South America. News articles that
lacked site country information were categorized
as Unclassi ed.</p>
      <p>We could use all of these information types other
than the URL to perform multifaceted analyses.</p>
    </sec>
    <sec id="sec-2">
      <title>3 NSContrast with GDELT</title>
      <p>We set up our system based on the GDELT GKG
from July 20, 2015 to January 19, 2016. Using the
data conversion process described above, we extracted
31,584,327 articles from 70,781 news sites.</p>
      <p>3https://freegeoip.net/,
http://ipinfo.io
http://ip-api.com/,
and</p>
      <p>First, we present information related to the
country and region estimation. Because our manually
constructed news site list is small, only 2201 sites
(8,555,263 articles) were identi ed by using this
information. Table 1 shows the number of articles
(sites) by the top-level domain of URLs (Top 6).
Because 81.2% (47,259/70,781) of news sites and (71.9%
(22,716,591/31,584,327) of articles have .com as their
top-level domain, only 10,139 sites (5,671,259 articles)
were identi ed by their top-level domain.</p>
      <p>Table 1: Number of articles (sites) for top-level
domains (Top 6)
.com 22,716,591 (47,259) .au 2,623,813 (1048)
.uk 1,682,960 (2705) .org 1,029,232 (8049)
.net 645,996 (3015) .ca 326,184 (1008)</p>
      <p>Finally, by using the geolocation service 57,459
sites (16,816,980 articles) were identi ed. As a
result, most of the sites (98.6%: 69,799/70,781) and
articles (98.3%: 31,043,442/31,584,327) were classi ed
into countries and regions.</p>
      <p>Table 2 shows the number of articles for each region.
From this table, news articles from the USA were
dominant in the database (61.6%: 19,443,005/31,581,063).
In contrast, there were only 903,811 articles from
North America excluding the USA. With such
unbalanced numbers of articles, making a category North
America including the USA is almost equivalent to
USA alone. Therefore, we divided North America into
the USA and North America (excluding USA).</p>
      <p>Table 2: Number of articles for each region
USA 19,443,005 Europe 3,696,359
Oceania 2,962,792 Asia 2,891,865
North America (excluding USA) 903,811
Africa 726,626 Middle East 373,411
South 45,573 Unclassi ed 537,621
America</p>
      <p>Our multifaceted analysis interface was used to
compare the results with different query conditions.
Figure 2 shows a time-sequence graph of polarity in
different countries: all countries (upper left), China
(upper right), the USA (lower left), and Europe
(lower right). These graphs were constructed by
adding new query conditions when selecting the data.
For example, the graph for China uses news articles
that included \Asian Infrastructure Investment Bank"
(AIIB) as the organization, an article date July 20,
2015, and the SiteCountry = \China."</p>
      <p>This gure shows that there were many positive
articles about AIIB in China. Europe was slightly
positive than the USA. This information re ects the
attitudes to AIIB in these countries (or regions).</p>
    </sec>
    <sec id="sec-3">
      <title>4 Conclusion</title>
      <p>In this study, we have analyzed the characteristics of
GDELT data and propose a data conversion process to
utilize this information for NSContrast. In this
conversion process, we conducted deduplication of news
article URLs and added source country and region
information to analyze the characteristic differences
between them. Because of the large coverage of news
sites, the system can conduct comparative analyses of
various countries and regions by using large numbers
of news articles from different news sites. However, for
future work, it may be better to check the
appropriateness of the estimated country by using a geolocation
service.</p>
      <p>Acknowledgement
This work was partially supported by JSPS KAKENHI
Grant Number 25280035.</p>
      <p>References
[Kle02] Jon Kleinberg. Bursty and hierarchical
structure in streams. In Proceedings of the 8th ACM
SIGKDD Intl. Conf. on Knowledge Discovery
and Data Mining, pages 91{101, New York,
NY, USA, 2002. ACM Press.
[LS13] Kalev Leetaru and Philip A. Schrodt.</p>
      <p>Gdelt:global data on events, location, and
tone, 1979-2012. In ISA Annual Convention
2013, volume 2, page 4, 2013.
[YK12] Masaharu Yoshioka and Noriko Kando.
Multifaceted analysis of news articles by using
semantic annotated information. In Proceedings
of the fth workshop on Exploiting semantic
annotations in information retrieval, ESAIR
'12, pages 19{20, New York, NY, USA, 2012.
ACM.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>