<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Building a Knowledge Base using Microblogs: the Case of Cultural MicroBlog Contextualization Collection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thi-Bich-Ngoc Hoang</string-name>
          <email>thi-bich-ngoc.hoang@irit.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Josiane Mothe</string-name>
          <email>josiane.mothe@irit.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>(1) IRIT, UMR5505 CNRS, Universite de Toulouse</institution>
          ,
          <addr-line>Toulouse</addr-line>
          ,
          <country country="FR">France (</country>
          <institution>2) University of Economics, the University of Danang</institution>
          ,
          <addr-line>Vietnam (3) ESPE, UT2J</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Cultural MicroBlog Contextualization (CMC) Workshop provides a collection of tweets on cultural events related to festivals. Given the size of a tweet, the information obtained by a single post is often very partial. We develop the idea that using a set of tweets about an event could enable having a more complete view of that event by combining all information posted. In this paper, we propose a model to represent the collection of microblogs into a knowledge base. Considering the set of tweets on festival events from CMC, we de ne a domain ontology and show how to populate this ontology based not only on the tweet collection but on external data too. We detail how the knowledge base could be used to provide a complete view of an event. This paper presents the preliminary results.</p>
      </abstract>
      <kwd-group>
        <kwd>Information retrieval</kwd>
        <kwd>Tweet analysis</kwd>
        <kwd>Cultural MicroBlog Contextualization</kwd>
        <kwd>Knowledge base from tweets</kwd>
        <kwd>Microblog</kwd>
        <kwd>Information extraction from tweets</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The Cultural MicroBlog Contextualization (CMC) Workshop aims at discussing
applications and tools based on a collection of tweets on cultural events related
to festivals [
        <xref ref-type="bibr" rid="ref30 ref31">30, 31</xref>
        ]
      </p>
      <p>A tweet is composed of 140-characters and in the case of CMC, it corresponds
to a twitter's post that contains the \festival" term. Such a collection can be used
to analyze what has been said about a given festival, what happens during the
festival, ... However, when considering a single tweet, the information is often
very partial and it is more likely that a human rather needs to read a set of
tweets to get a clear picture of an event.</p>
      <p>For example, the three following tweets, all related to Cannes 2015, provide
di erent and complementary pieces of information:</p>
      <p>Ouverture de la route des Golden Globes avec Carol de Todd Haynes, Le ls de
Saul et Mustang! A suivre! #Cannes2015 pic.twitter.com/YKd43HORmk
Vincent Lindon &amp; Gaspar Noe, guests of honour at #VentanaSur Festival de
Cannes Film Week from 30/11/15 to 6/12/15! pic.twitter.com/slPVK t24
Irina Shayk, somptueuse, lors du tapis rouge du 19 mai 2015 a Cannes,
pinterest.com/pin/4530340437. . .</p>
      <p>The rst tweet is about the lm Carol directed by Todd Haynes to be
presented at the Cannes 2015 festival. While the second tweet provides the date of
a related event in Buenos Aires (VentanaSur ) along with two actors who were
there; it is an add for the Buenos Aires festival. Finally the third tweet gives the
information about a speci c date at festival de Cannes 2015 where the model
Irina Shayk showed up.</p>
      <p>When considering these three individual tweets, it is obvious that some users
will lack of context to understand them individually. However, some pieces of
information from various tweets could help understanding a given tweet. For
example, given the second tweet, if the user does not know the VentanaSur
festival, he may mismatch festival de Cannes and VentanaSur festival. When
considering both the second and the third tweets, he will nd that festival of
Cannes is in May and not at the end of the year, which was not obvious when
considering the second tweet only. Each tweet taken individually provides partial
information; but the sum of them could give a better picture of the information
or of an event. If all pieces of information from the tweet set could be used to
enrich a knowledge base, it would then be possible to understand better each
tweet individually by contextualizing it using additional knowledge.</p>
      <p>Moreover, some parts of the knowledge could rely on existing resources such
as geographical hierarchies or domain knowledge rather than on tweets only.
For example, understanding the second tweet would be easier if the user knew
the entity types \Vincent Lindon" and \Gaspar Noe" belong to (V. Lindon is a
player and G. Noe a director) and that \VentanaSur" is a \Festival".</p>
      <p>In this paper, we propose a model to represent a collection of microblogs
by a domain ontology. By combining the festival tweet collection (the CMC
collection of CLEF 2016) with other Internet resources, we aim at bringing a
complete picture of the collection content that can make a complete view of
festival events referenced in this collection.</p>
      <p>To populate the skeleton of the ontology, we use Wikipedia (or rather
DBPedia1) as well as websites which provide o cial pieces of information about
geography, list of festivals and related details. This information is quite stable
in time. Next, the tweets related to each festival are selected using information
retrieval methods. They are analyzed to recognize and extract named entities
(NE) such as locations, artists, festival names, time. This extracted information
can be used to populate instances of the corresponding classes in the ontology.</p>
      <p>The knowledge base we design could be used in applications where the
users (1) would choose a speci c festival name and have a picture of that
festival through the tweets (2) would choose a location and would get a list
1 BDpedia structures the information from Wikipedia pages; it can be queried using</p>
      <p>SPARQL to extract structured information
of corresponding festivals, etc. For example, from the three tweets mentioned
above, our ontology could help a user inferring from Ouverture de la route
des Golden Globes avec Carol de Todd Haynes #Cannes2015 and Irina
Shayk, somptueuse, lors du tapis rouge du 19 mai 2015 a Cannes that
the lm Carol was presented in May 2015 at the Cannes festival.</p>
      <p>Currently, there are various ways to represent knowledge, but we believe that
ontology (e.g. OWL-based) is an appropriate and e cient solution because of
the following reasons. Firstly, it makes our system an easily accessible
knowledge base. The ontology-based knowledge represents data in a common language
platform which can be shared and retrieved by Resource Description Framework
(RDF) query language. Moreover, it allows inferring new knowledge from
existing data that make users understand more about incomplete data in tweets.
Finally, it could provide complete and updated information about festivals by
combining Internet resources and the tweet collection.</p>
      <p>This paper does not cover the entire project, rather it focuses on the domain
representation and on the ontology population. We also mention some ways this
knowledge base could be used in some applications.</p>
      <p>The rest of the paper is organized as follows: Section 2 presents the related
work. Section 3 details the model we suggest to represent the festival domain.
Section 4 explains how the knowledge base is populated. Finally, section 5
concludes this paper, discusses about applications and future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>Due to the rising popularity of social media, many studies propose ways to
extract information from this resource. Prior works related to ours are grouped
into three categories: ontology-based information extraction, event detection,
and location estimation in microblogs.
2.1</p>
      <sec id="sec-2-1">
        <title>Ontology-based information extraction</title>
        <p>
          In recent years, a number of papers have addressed the ontology-based
information extraction. Narayan et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] suggest an approach to populate an ontology
with the events retrieved from Twitter. Data is parsed and mined for various
features such as name, date, time, location, type and URL that are later used
to populate the ontology. The authors use the existing ontology from [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] to
identify time and use Alexandria Digital Library Gazetteer (1999) to recognize
Location and Name. Using these methods, they are not able to detect NE when
it is not explicitly mentioned in a tweet content.
        </p>
        <p>
          Kontopoulos et al. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] present a method for sentiment analysis of tweets based
on an ontology. They rst identify the topic discussed in tweets and then give
each tweet the sentiment score for each distinct aspect relevant to the topic.
Another study is from [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], the authors propose an ontology-based information
extraction for recognizing and semantically disambiguating NE in tweets. They
solve the problem of entity disambiguation by using syntactical context and
Linked Data as Freebase.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Event detection</title>
        <p>
          In the area of event detection, Weng et al. [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] build signals for individual words
and lter out trivial words based on their corresponding auto correlations signal.
They extract events by clustering signals and using modularity-based graph
partitioning. Similarly, Zhao et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] propose a text-based clustering and temporal
segmentation combined with information ow-based graph analysis. Besides, by
aggregating information across multiple messages, Benson et al. [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] present a
graphical model to detect entertainment events while Sakaki et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] use a
probabilistic spatio-temporal model to detect earthquake and use Kalman and
particle ltering to estimate location. Using a di erent approach, Quack et al.
[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] detect local events by analyzing community photo collections while Lee et al.
[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and Watanabe et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] analyze the geographical distribution of geo-tagged
microblogs to detect events.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Location extraction</title>
        <p>
          A location is either explicitly mentioned or should be inferred from content. NE
recognition (NER) systems have addressed the problem of retrieving location
speci ed in documents; however they do not perform very well on informal texts
[
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. The literature proposes some methods to improve this limitation. Liu et al.
[
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] combine a K-Nearest Neighbors classi er with a linear Conditional Random
Fields model under a semi-supervised learning framework to tackle the lack
of information in microblogs, while Krishnan et al. [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] propose a two-stage
approach to handle non-local dependencies in NER. By aggregating information
garnered from the World Wide Web to build local and global contexts from
tweets, Li et al. [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] target the error-prone and short nature challenges. Another
location estimation approach is to rely on analyzing geo-location by content
analysis either with terms in gazetteer [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], with probabilistic model [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], or users'
networking [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>In the next sections, we present the knowledge base we promote as well as
the way we populate it. We also present some preliminary results based on the
CMC CLEF 2016 festival tweet set.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Knowledge base model: the geographical-festival ontology</title>
      <p>Events have several dimensions, the main ones are:
{ Location information which indicates where the event takes place;
{ Temporal information that indicates when the event takes place;
{ Entity-related information which indicates what the even is about.</p>
      <p>In the case of festival-related events, we can have a more speci c
representation.
Music class consists of Classical, Rock, Jazz, Pop... We use a set of categories
to contribute to the Festival part of our ontology including a number of
classes such as Music, Art, Film, Parades, which are types of festivals. This
hierarchy of categories is proposed by DBPedia. It might not be complete but
it is appropriate to start with and it can be completed later on, considering
tweets contents.
{ Lastly, the Tweet class contains tweets which relate either to a speci c
festival or a location. Tweets that cannot be related to either a festival or a
location are not stored and considered as useless. One tweet might be about
entities from the Performance part of the ontology such as Time, Artist,
Show ...or contain fresh information of a festival or a location such as tra c,
weather, stories and feedback of attendees. We do not store this type of
information in various classes but keep the tweets that can be associated to either
a location, or a festival (or both) to be able to retrieve fresh information on
atmosphere, twitters' comments.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Populating the domain ontology</title>
      <p>In this section, we rst provide the general principles of the knowledge base
population then we detail the various steps of the ontology population.
4.1</p>
      <sec id="sec-4-1">
        <title>Principles</title>
        <p>The domain ontology is populated considering complementary resources. We use
both a ow of tweets that match the information need festival and which can
be seen as our main resource for fresh (and possibly subjective) information,
and external resources such as DBPedia or tourism websites that contain more
stable information even if they can be frequently up-dated (speci cally
considering festivals to come). Figure 2 depicts the overall principle of the ontology
population: Web and DBPedia resources are used to rst populate the skeleton
of the ontology. DBPedia provides general information about existing locations,
festival categories and even most of well-known festivals in the world; o cial
festival and tourism websites provide more speci c information about some
festivals (for example for the Jazz festival in Marciac, the o cial festival website
can be analyzed) and some hubs such as the Syndicats d'initiative websites can
also provide some additional links to other festivals.</p>
        <p>Then the ontology and the tweet collection are used in a process that
combines the information: from the ontology, we know festivals and locations that
help analyzing the tweets which in turns can be used to extract new information
to populate the ontology. For instance, from the ontology, it is possible to know
that in Cannes, there is a event named Cannes lm festival. Then, Cannes lm
festival is used to detect all tweets related to this event. These tweets, in turn,
are used to extract time, artists, and shows to populate the ontology.</p>
        <p>The ontology population using DBPedia and o cial websites resources can
be seen as resources for background ontology population while tweet collection
is a resource for providing complementary views about the events.</p>
        <p>To begin with, we chose Protege 2 to build the ontology that implements
the knowledge base. We created the ontology structure as described in Figure
1 including classes such as Continent, Country, Town, Tweet, Festival.... The
Location and Festival parts are to be created by data extracted from resources
such as DBPedia and o cial websites. Then tweets related to each festival can
be identi ed and populate the Tweet class; the relationships with Location and
Festival are established in the knowledge base. In addition, information from
those tweets such as Time, Artist and Show are extracted to populate the
Performance part of the ontology when possible. The process will be nalized by
applying inference mechanism to get new information from existing data.</p>
        <p>In the next sections, we explain in details the populating process accompanied
by preliminary results. We run the main steps of our approach on 500 tweets
about Cannes and Lyon extracted from the CMC CLEF 2016 collection. This
collection contains 38,686,650 tweets about festivals in the world collected from
May to October 2015.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Location population</title>
        <p>
          The location part of the ontology is populated using Ngo et al. results [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. They
extract the geographic data from Wikipedia which provides the list of locations
for each countries. For example, for France it includes communes (overseas
departments included) with a population over 20,000. The data is structured using
3 levels: \commune", \departement", and \region". We use the country and town
(\commune") of their data to populate the ontology. There are 3; 885 instances
of locations for France. Concretely, we only keep a few in our rst prototype
since Protege is limited in the number of instances it can handle without using
a database.
2 http://protege.stanford.edu/ Protege is an open-source platform for building
knowledge-based ontologies.
        </p>
        <p>An alternative solution for geographic data could have been to use other
geographic resources such as GeoName 3 or GEOnet Names Server 4, but Wikipedia
provides accurate and reliable information on this topic and was enough for our
Proof-of-Concept application.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Festival population</title>
        <p>The Festival part of the ontology is populated using the list of festivals
provided by DBPedia 5. Although the information from these resources changes,
the update rate is not necessarily very high to keep the ontology accurate. This
structured information can be extracted using SPARQL on locally stored
DBpedia or through endpoint framework 6. In our work, for the rst implementation,
we query information from DBPedia through the second way.</p>
        <p>In addition, other information related to a festival could also be retrieved
from DBPedia such as the festival location and o cial website. In turn, it would
then be possible to collect the corresponding Twitter account, hashtags (from
twitter page) and keywords about the festivals and consider them as additional
properties to detect festivals in tweets as presented in the section 4.4. We keep
the automation of this process for later and handle this task manually for a few
festivals for now for Proof-of-Concept.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Relationship between tweets, festivals and locations</title>
        <p>We associate tweets related to speci c festivals or locations. We compare the list
of festivals and properties resulting from the Festival population (section 4.3)
with the tweet contents in order to identify all tweets related to each festival.
The priority is set for festival names, twitter accounts, hashtags and keywords
respectively.</p>
        <p>When considering the sub-collection of 500 tweets, we detected 137 festivals
from 137 tweets including 70 festivals detected by names, 61 festivals detected
by hasgtags and 6 festivals detected by Twitter account.</p>
        <p>To recognize locations in tweets, we combine Stanford NER with other
techniques such as inferring from festival location and mining the Twitter user's
pro le.</p>
        <p>
          We use Stanford NER to recognize locations that are explicitly mentioned in
tweet contents. Because numerous twitters specify locations in their text right
after a hashtag (#) Stanford NER do not extract it. For this reason, we remove
all hashtags in texts before using Stanford NER. In the case locations are not
speci ed in a tweet, we infer the location from the festival that this tweet relate
to. Finally, if a tweet does not contain any text about location or festival, we
3 http://www.geonames.org/
4 http://geonames.nga.mil/gns/html/
5 http://dbpedia.org/page/Lists_of_festivals: The root page provides festivals
by categories of all countries in the world
6 http://dbpedia.org/snorql/
mine the Twitter user's pro le to extract the home residence. We consider this
hometown as the location that his tweets are about due to a conclusion from
[
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]: 50% users post most of their tweets in their home residence.
        </p>
        <p>We set a priority for the three location extraction techniques: Stanford NER,
inference mechanism and pro le mining. In case a location in a tweet is
recognized by more than one method, we chose the most suitable one (detected by
the highest priority technique).</p>
        <p>Using the 500 tweets, we detected 487 locations from 409 tweets including:
1) 313 locations identi ed by Stanford NER in 225 tweets, 2) 137 locations for
137 tweets based on the festivals 3) 245 locations recognized by Twitter users'
pro le. We are currently working on more sophisticated techniques to extract
location from a tweet.
4.5</p>
      </sec>
      <sec id="sec-4-5">
        <title>Performance population</title>
        <p>From tweets that can be related to festivals or locations (see section 4.4), we
use Stanford NER to extract entities such as time, artists, shows... In the 500
tweet collection, we detected 131 artists from 103 tweets, 99 time points from
99 tweets. These instances and relationships and the corresponding tweets are
stored in the ontology.
4.6</p>
      </sec>
      <sec id="sec-4-6">
        <title>Inferring new knowledge</title>
        <p>The inference mechanism is used to infer the relationships between instances in
the case they are not directly set up from previous steps. Back to an example
mentioned in the introduction part, a user can extract that festival of Cannes
is in May even if the time is not mentioned in the rst and second tweets. It
is inferred from the third tweet. In our approach, we inferred 137 locations for
137 tweets based on the festivals that these tweets related to, 30 relationships
between Festival and Artist, 19 relationships between Artist and Time classes,
55 relationships between Festivals and Time.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Discussion</title>
      <p>In this paper, we introduced an approach for building a knowledge base using
Twitter and other external resources for the case of Cultural MicroBlog
Contextualization (CMC) collection.</p>
      <p>The model considers festivals organized in a speci c location and related
information such as time, artists or shows. By combining the festival tweet
collection with DBPedia and o cial websites resources, we help building a more
complete picture of festivals occuring in the data collection.</p>
      <p>For this purpose, we de ne a festival ontology. As a background task, the
population of the location and festival parts is based on resources such as
DBPedia and o cial websites. In addition, tweets related to speci c festivals or
locations are retrieved and analyzed to extract related data.</p>
      <p>We believe that by employing ontology technology, we provide an easily
accessible knowledge base system. Comparing to storing data in traditional databases,
our approach has several pros. Firstly, data is presented in a common language
platform which can be much easily retrieved by SPARQL. A RDF data model is
also easier to updated without adverse e ects to the application, thus it requires
less maintenance. Secondly, the inference mechanism of ontology language
allows inferring new knowledge from existing data easily (in the proof-of-concept
we program the inference, but ontology allows such a process). Lastly, by
combining several resources such as DBPedia, websites and Twitter, our system could
bring a complete and fresh knowledge about festivals by cities in the world
including o cial information from websites and the latest stories from Twitters.</p>
      <p>
        To recognize NE in the CMC collection, we combine Standford NER with
inferring technique and mining user's pro les. Applying Standford NE extraction
on microblogs might not be optimal; some methods have been developed on the
speci c case of tweets such as [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] that could be tested. We will leave this task
for future work.
      </p>
      <p>We would also want to extract short summaries of festivals from BDpedia or
o cial websites to propose the users a basic idea of the festivals he selects.</p>
      <p>We suppose that the knowledge base model we built have a broad range of
applications in several domains such as tourism, transportation, marketing and
advertisement.</p>
      <p>In the eld of tourism, using our knowledge base to build a graphical
recommender system with highly informative summaries about events, famous people,
related activities aggregated from tweets would be valuable. Tourists do not have
to spend time to search and process information for their need. Moreover, latest
news, opinions and feedback are more likely to appear in tweets rather than in
o cial websites.</p>
      <p>In the transportation domain, a system based on our knowledge base that
would suggest a suitable route or transportation mean to avoid crowds, tra c
jams or other problems could be welcomed by travels.</p>
      <p>Besides, festivals could be perfect places for companies to market their brand.
They can communicate with thousands of participants and engage participants
through targeted campaigns. Knowing the type of festivals, type of participants
as well as the artists, shows, dates, companies could propose and implement
e ective advertisement campaigns for their products.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bontcheva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derczynski</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Funk</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greenwood</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maynard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aswani</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <article-title>TwitIE: An Open-Source Information Extraction Pipeline for Microblog Text</article-title>
          . In
          <string-name>
            <surname>RANLP</surname>
          </string-name>
          , (pp.
          <fpage>83</fpage>
          -
          <lpage>90</lpage>
          ) (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ritter</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <article-title>Named entity recognition in tweets: an experimental study</article-title>
          .
          <source>In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Nebhi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>Ontology-based information extraction from twitter (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Iwanaga</surname>
            ,
            <given-names>I. S. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>T. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kawamura</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakagawa</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tahara</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohsuga</surname>
            ,
            <given-names>A. Building</given-names>
          </string-name>
          <article-title>an earthquake evacuation ontology from twitter</article-title>
          .
          <source>In Granular Computing (GrC)</source>
          , 2011 IEEE International Conference on (pp.
          <fpage>306</fpage>
          -
          <lpage>311</lpage>
          ), IEEE (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Narayan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prodanovic</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elahi</surname>
            ,
            <given-names>M. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bogart</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <article-title>Population and Enrichment of Event Ontology using Twitter. Information Management SPIM (</article-title>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kontopoulos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berberidis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dergiades</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bassiliades</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <article-title>Ontology-based sentiment analysis of twitter posts</article-title>
          .
          <source>Expert systems with applications</source>
          , (pp.
          <fpage>4065</fpage>
          -
          <lpage>4074</lpage>
          ) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Cheng,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Caverlee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>You are where you tweet: a content-based approach to geo-locating twitter users</article-title>
          .
          <source>In Proceedings of the 19th ACM international conference on Information and knowledge management</source>
          (pp.
          <fpage>759</fpage>
          -
          <lpage>768</lpage>
          ) (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Chandra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muhaya</surname>
            ,
            <given-names>F. B.</given-names>
          </string-name>
          <article-title>Estimating twitter user location using social interactions{a content based approach</article-title>
          . In Privacy, Security,
          <source>Risk and Trust (PASSAT)</source>
          and
          <source>2011 IEEE Third Inernational Conference on Social Computing (SocialCom)</source>
          ,
          <year>2011</year>
          IEEE Third International Conference on (pp.
          <fpage>838</fpage>
          -
          <lpage>843</lpage>
          ) (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Fink</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piatko</surname>
          </string-name>
          , C. D., May eld, J.,
          <string-name>
            <surname>Finin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martineau</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Geolocating Blogs from Their Textual Content</article-title>
          .
          <source>In AAAI Spring Symposium: Social Semantic Web: Where Web 2.0 Meets Web 3.0</source>
          (pp.
          <fpage>25</fpage>
          -
          <lpage>26</lpage>
          ) (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Abel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Celik</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Houben</surname>
            ,
            <given-names>G. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siehndel</surname>
            ,
            <given-names>P. Leveraging</given-names>
          </string-name>
          <article-title>the semantics of tweets for adaptive faceted search on twitter</article-title>
          .
          <source>In The Semantic Web{ISWC</source>
          <year>2011</year>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          ). Springer Berlin Heidelberg (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cui</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Analyzing social media via event facets</article-title>
          .
          <source>In Proceedings of the 20th ACM international conference on Multimedia</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Sakaki</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Okazaki</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsuo</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Earthquake shakes Twitter users: real-time event detection by social sensors</article-title>
          .
          <source>In: The 19th international conference on World wide web, ACM</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Quack</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leibe</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Gool</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>World-scale mining of objects and events from community photo collections</article-title>
          . In:
          <article-title>The 2008 international conference on Contentbased image and video retrieval</article-title>
          ,
          <source>ACM</source>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Watanabe</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ochi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Okabe</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Onai</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Jasmine: a real-time local-event detection system based on geolocation information propagated to microblogs</article-title>
          .
          <source>In: The 20th ACM international conference on Information and knowledge management</source>
          (pp.
          <fpage>2541</fpage>
          -
          <lpage>2544</lpage>
          ) (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sumiya</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>Measuring geographical regularities of crowd behaviors for Twitter-based geo-social event detection</article-title>
          .
          <source>In Proceedings of the 2nd ACM SIGSPATIAL international workshop on location based social networks</source>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          ) (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>Temporal and information ow based event detection from social text streams</article-title>
          .
          <source>In AAAI (Vol. 7</source>
          , pp.
          <fpage>1501</fpage>
          -
          <lpage>1506</lpage>
          ) (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Aiello</surname>
            ,
            <given-names>L. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petkos</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corney</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papadopoulos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skraba</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , ...
          <string-name>
            <surname>Jaimes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Sensing trending topics in Twitter</article-title>
          . Multimedia, IEEE Transactions on, (pp.
          <fpage>1268</fpage>
          -
          <lpage>1282</lpage>
          ) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srihari</surname>
            ,
            <given-names>R. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <article-title>Location normalization for information extraction</article-title>
          .
          <source>In Proceedings of the 19th international conference on Computational linguistics-Volume</source>
          <volume>1</volume>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          ) (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <article-title>Location-based event search in social texts</article-title>
          .
          <source>In Computing, Networking and Communications (ICNC)</source>
          , 2015 International Conference on (pp.
          <fpage>668</fpage>
          -
          <lpage>672</lpage>
          ) (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Datta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>B. S.</given-names>
          </string-name>
          <string-name>
            <surname>Twiner</surname>
          </string-name>
          <article-title>: named entity recognition in targeted twitter stream</article-title>
          .
          <source>In: The 35th international ACM SIGIR conference on Research and development in information retrieval</source>
          (pp.
          <fpage>721</fpage>
          -
          <lpage>730</lpage>
          ) (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Ngo</surname>
            ,
            <given-names>Q. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winiwarter</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <article-title>Using Wikipedia for extracting hierarchy and building geo-ontology</article-title>
          .
          <source>International Journal of Web Information Systems</source>
          , (pp.
          <fpage>401</fpage>
          -
          <lpage>412</lpage>
          ) (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Wimalasuriya</surname>
            ,
            <given-names>D. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Ontology-based information extraction: An introduction and a survey of current approaches</article-title>
          .
          <source>Journal of Information Science</source>
          .(
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Nagarajan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Purohit</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          <article-title>A Qualitative Examination of Topical Tweet and Retweet Practices</article-title>
          . ICWSM, pp. (
          <volume>295</volume>
          -
          <fpage>298</fpage>
          ) (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Benson</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haghighi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barzilay</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Event discovery in social media feeds</article-title>
          .
          <source>In: The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume</source>
          <volume>1</volume>
          (pp.
          <fpage>389</fpage>
          -
          <lpage>398</lpage>
          ) (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Hobbs</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>An ontology of time for the semantic web</article-title>
          .
          <source>ACM Transactions on Asian Language Information Processing (TALIP)</source>
          , (pp.
          <fpage>66</fpage>
          -
          <lpage>85</lpage>
          ) (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Krishnan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C. D.</given-names>
          </string-name>
          <article-title>An e ective two-stage model for exploiting non-local dependencies in named entity recognition</article-title>
          .
          <source>In: The 21st International Conference on Computational Linguistics</source>
          and
          <article-title>the 44th annual meeting of the Association for Computational Linguistics</article-title>
          (pp.
          <fpage>1121</fpage>
          -
          <lpage>1128</lpage>
          ) (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Weng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>B. S. Event</given-names>
          </string-name>
          <article-title>Detection in Twitter</article-title>
          . ICWSM,
          <volume>11</volume>
          ,
          <fpage>401</fpage>
          -
          <lpage>408</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hwang</surname>
            ,
            <given-names>B. Y.</given-names>
          </string-name>
          <article-title>A Study of the Correlation between the Spatial Attributes on Twitter</article-title>
          .
          <source>In Data Engineering Workshops (ICDEW)</source>
          ,
          <year>2012</year>
          IEEE 28th International Conference on (pp.
          <fpage>337</fpage>
          -
          <lpage>340</lpage>
          ) (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Recognizing named entities in tweets</article-title>
          .
          <source>In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume</source>
          <volume>1</volume>
          (pp.
          <fpage>359</fpage>
          -
          <lpage>367</lpage>
          ) (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>SanJuan</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Moriceau</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tannier</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bellot</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Overview of the INEX 2012 tweet contextualization track. Initiative for XML Retrieval INEX</article-title>
          ,
          <volume>148</volume>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murtagh</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanjuan</surname>
          </string-name>
          , E..
          <source>Overview of the CLEF 2016 Cultural Microblog Contextualization Workshop</source>
          , Experimental IR Meets Multilinguality, Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          .
          <source>Proceedings of the Seventh International Conference of the CLEF Association (CLEF 2016), Lecture Notes in Computer Science (LNCS) 9822</source>
          , Springer, Heidelberg, Germany,
          <year>2016</year>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Ermakova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>J. IRIT</given-names>
          </string-name>
          at INEX 2012:
          <article-title>Tweet Contextualization</article-title>
          . In CLEF (Online Working Notes/Labs/Workshop) (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>