<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A tweets classier based on cosine similarity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Carolina Fcil-Arias</string-name>
          <email>focil.carolina@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Zoeaeiga</string-name>
          <email>zujorge@live.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grigori Sidorov</string-name>
          <email>sidorov@cic.ipn.mx</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ildar Batyrshin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Gelbukh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CIC, Instituto PolitØcnico Nacional (IPN)</institution>
          ,
          <addr-line>Mexico City</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The 2017 Microblog Cultural Contextualization task consists in three challenges: (1) Content Analysis, (2) Microblog search, and (3) TimeLine illustration. This paper describes the use of cosine similarity, which is characterized by the comparison of similarity between two vectors of an inner product space. This research used two approaches: (1) word2vec and (2) Bag-of-Words (BoW) for extracting all relevant tweets to each event related to the four festivals: Charrues, Transmusicales, Avignon and Edinburgh.</p>
      </abstract>
      <kwd-group>
        <kwd>cosine similarity</kwd>
        <kwd>natural language processing</kwd>
        <kwd>Bag-of-Words</kwd>
        <kwd>word2vec</kwd>
        <kwd>opinion mining</kwd>
        <kwd>information retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Opinion mining is dened as "the task of classifying texts into categories
depending on whether they express positive or negative sentiment, or whether they
enclose no emotion at all" [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        The growth of social media provides a domain of great interest for several
studies related on opinion mining (also known as sentiment analysis) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] such as
opinion analysis related to topics or problems of political preferences [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], opinion
about a specic product [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], news [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and others. According to [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Twitter is
the most popular microblogging network in the world; this microblogging service
has more than 300 million users, which represents a competitive advantage for
many organizations.
      </p>
      <p>In this project we propose the usage of cosine similarity with two features:
Bag-of-Words and word2vec using a dataset of workshop Microblog Cultural
Contextualization 2017 to determine the tweets relevance according to each
event from four European festivals.</p>
      <p>The remainder of this paper is structured as follows. In section 2, related
work on timeline illustration, cosine similarity and opinion mining are presented.
Section 3 is focused on showing the materials and methods. Section 4 describes
the experimental results where two features: Bag-of-Words and word2vec are
used with cosine similarity. Finally, section 5 gives a summary of this work.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <sec id="sec-2-1">
        <title>TimeLine illustration</title>
        <p>
          The CLEF 2017 TimeLine illustration shared task [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] was dedicated to
retrieve all relevant tweets based on festival events. A variety of analysis, such
as descriptive (duplicate and web addresses removal, abbreviations, etc.) [
          <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
          ],
correspondence and interactive (FactoMineR package) [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] were used.
        </p>
        <p>
          When looking at the State-of-the-Art on opinion mining for TimeLine
illustration, we distinguish that Dogra et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] used several approaches, such
as (i) content-based retrieval, where a query represents a topic for matching in
the documents. A common method used in this task is BM25 (Best matching),
which is a probabilistic information retrieval function for documents features
such as term and document frequencies, and document length [
          <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
          ], (ii)
diversication is used when a retrieval function does not take into account the
relations among returned documents; they may have relevant and redundant
information [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], and (iii) re-ranking for improving the retrieval results through a
baseline system [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          In related research, Murtagh [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] used correspondence analysis to map the
dual spaces of days and hashtags to a latent semantic space with social narratives
information, and related work on Pierre Bourdieu news.
        </p>
        <p>
          According to Hoan [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], the model usage according to base knowledge, makes
use of an ontology to identify the relation among tweets, festivals and locations.
They used a combination of Stanford NER, festival location and user prole. The
tweets comparison related to each festival consists of festivals and properties lists
with the tweet content.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Cosine similariy method In this paper, we select a cosine similarity approach for unsupervised learning, and now, we will present some works related to this method with similar objectives.</title>
        <p>
          Shi and Macy [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] compared a standardized Co-incident Radio (SCR) with
Jaccard index and cosine similarity. They chose SCR to map sport league studies,
music artists, and congress members that can obtain more followers on Twitter.
        </p>
        <p>
          AL-Smadi et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] provided a semantic text similarity approach using text
overlap, word alignment and semantic features to identify the paraphrasing in
Arabic news tweets. This semantic text similarity is based on Support Vector
Regression, which is used for regression analysis. The method achieved good
accuracy according to State-of-the-Art.
        </p>
        <p>
          Tajbakhsh and Bagherzadeh [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] used cosine similarity to nd coincidences
among tweets. Also, the results were compared with several semantic-based
algorithms such as Shortest path, Wu &amp; Palmer, Lin, JiangConrath, Resnik, Lesk,
LeacockChodorow, and Hirst-STOnge.
        </p>
        <p>Other related works that use cosine similarity are presented in [1921].
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Sentiment analysis approaches</title>
        <p>
          In this section, we overview several researches based on tweets classication.
Looking at particular systems, Dela Rosa, et al. [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] presented a study related on
clustering and classifying tweets into dierent categories such as highly relevant,
somewhat relevant, not relevant and spam using hash-tags as indicators. The
aim of this approach is to nd the most representative tweets for some story
such as video’s Lady Gaga, Obama’s help to Japan, disturbing video of Charlie
Sheen, and others. The results of this technique performed well based on
Stateof-the-Art.
        </p>
        <p>
          NÆdia et al. [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] besought a combination of classier ensembles (Random
Forest, Suppport Vector Machines, Multinomial Naive Bayes and Logistic
Regression) and lexicon for predicting whether a tweet is positive or negative
concerning a query term. This ensemble method used two feature representations
(Bag-of-Words and hashing) and the results provided an improvement in the
State-of-the-Art
        </p>
        <p>
          Another approach was presented in [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], where a combination of classier
ensembles (Multinomial Naive Bayes, SVM, Random Forest, and Logistic
Regression) and lexicons were trained for identication of tweets polarity.
        </p>
        <p>
          In related research, Paltoglou and Thelwall [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] proposed an unsupervised
approach based on lexicon-classier using Bag-of-Words as features. This method
estimates the polarity and subjectivity of informal texts on the web, such as
tweets, social network and online discussions. The results show that the
proposed algorithm presents a trustworthy solution for analyzing feelings of informal
communication on internet.
        </p>
        <p>
          With regard to sentiment analysis of tweets posted on Twitter during a
disaster, Venkata et al. [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] provided an approach with Bag-of-Words, polarity clues,
emoticons, internet acronyms sentistrength and puntuaction as features for
identifying and categorizing the sentiments.
        </p>
        <p>
          As we can see, many researchers have focused on the information that twitter
can express, due to the opinions of consumers concerning brands and products
[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], political and social events [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], health care [
          <xref ref-type="bibr" rid="ref28 ref29">28, 29</xref>
          ], higher education [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ],
microblogging and social network services [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ], foresight [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ] and others.
        </p>
        <p>In this work, we selected an approach for unsupervised learning called cosine
similarity method using two types features: word2vec and Bag-of-Words. Each
experiment was evaluated individually.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Materials and Methods</title>
      <sec id="sec-3-1">
        <title>Dataset</title>
        <p>
          The dataset for this study came from the workshop Microblog Cultural
Contextualization 2017 [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ]. The data consists of festivals and topics collected
during July and December 2015. A brief summary of the dataset is shown
below.
1. Dataset "clef microblogs festival" : This dataset has the tweets related to four
festivals, where there are two French Musical festivals, one French theater
festival and one Great Britain theater festival. Altogether, this data contains
17.1 GB of information, which is represented by several variables such as:
id
username
date
content tweet
tweet link
microblogging name
2. Dataset "clef mc2 task3 topics" : This dataset is a XML le which contains
the four festivals. From this set, there are 664 events, which are divided by
61 events that correspond to the Charrues festival; 138 to the Edinburgh
festival, 365 to the Avignon festival, and 100 to the Transmusicales festival. An
example of structure in the XML document is given below.
        </p>
        <p>&lt;topics&gt;
...
&lt;topic&gt;
&lt;id&gt;5&lt;/id&gt;
&lt;title&gt;&lt;/title&gt;
&lt;artist&gt;Klangstof&lt;/artist&gt;
&lt;festival&gt;transmusicales&lt;/festival&gt;
&lt;startdate&gt;04/12/16-17:45&lt;/startdate&gt;
&lt;enddate&gt;04/12/16-18:30&lt;/enddate&gt;
&lt;venue&gt;UBU&lt;/venue&gt;
&lt;/topic&gt;
...</p>
        <p>&lt;/topics&gt;
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Preprocessing steps</title>
        <p>Tweets gathered from Twitter without a preprocessing stage have noise,
confusing words, URLs, stop words and so on. Therefore, this study used the following
preprocessing:
1. Consider tweets, which are written in latin alphabet such as English, Spanish,</p>
        <p>
          German, Italian, French, and others.
2. Eliminate all retweets.
3. Remove special characters, accents and lingistic inections.
4. Convert the dates mm/dd/yy into an only format dd/mm/yy.
5. Serialize and reduce the dataset into 10,000 tweets due to the large amount
of tweets.
6. Tokenize all words via NLTK toolkit [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ].
7. Replace several white spaces into a white space.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Features</title>
        <p>
          This study uses the following features after preprocessing the tweets.
1. Bag-of-Words. This is a basic representation of words as features. The aim
is to compute the word frequencies in a document to nd the similarities
and dierences among the documents [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ]. Figure 1 shows Bag-of-Words
approach applied in this study.
2. Word2Vec. This is a useful technique used to create word embedding [
          <xref ref-type="bibr" rid="ref36">36</xref>
          ].
        </p>
        <p>
          This means, representing a word as a vector [37]. We use Word2Vec to create
an embedding vector for each word in the tweet, compare all vectors and take
the maximum value to obtain sentence representation.
Cosine similarity is a measure that calculates the angle cosine between two
vectors [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. This technique indicates the degree of similarity between documents
which are represented by vectors; when the two vectors are equal then the
similarity is high and we obtain a value of 1.
        </p>
        <p>In this context, each topic and tweet are represented as vectors, where each
vector has the word frequencies (See Figure 1), and then, the cosine formula is
applied as follows.
where the topic document is represented by A, and the tweet document is
represented by B.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental results</title>
      <p>In these experiments, we used the format .res for representing the results as
shown in Table 1.</p>
      <p>id_event
1</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and future work</title>
      <p>The motivation of this work is focused on retrieval of the most relevant tweets
for each of the four festivals. We concentrated on two types of features:
Bag-ofWords and word2vec using the cosine similarity approach. Our results show that
this technique is capable of detecting the more popular opinion gathered from
Twitter based on similarity between tweets and topics.</p>
      <p>There is much future work to be done in this study such as the analysis of
hashtags, URL, emoticons and tweets characteristics. Furthermore, we will
conduct experiments using Bag-of-Words and word2vec features with other measures
of structural similarity such as Jaccard index and Phi coecient.
37. Enrquez, F., Troyano, J.A., Lpez-Solaz, T.: An approach to the use of word
embeddings in an opinion classication task. Expert Systems with Applications
66 (2016) 1 6</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Tsirakis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poulopoulos</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsantilas</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varlamis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Large scale opinion mining for social, news and blog data</article-title>
          .
          <source>Journal of Systems and Software</source>
          <volume>127</volume>
          (
          <year>2017</year>
          )
          <fpage>237</fpage>
          <lpage>248</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A review of natural language processing techniques for opinion mining systems</article-title>
          .
          <source>Information Fusion</source>
          <volume>36</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Golbeck</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>A method for computing political preference among twitter followers</article-title>
          .
          <source>Social Networks</source>
          <volume>36</volume>
          (
          <year>2014</year>
          ) 177184 Special Issue on Political Networks.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Fernandes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Douza</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Analysis of product twitter data though opinion mining</article-title>
          .
          <source>In: 2016 IEEE Annual India Conference (INDICON)</source>
          .
          <article-title>(</article-title>
          <year>2016</year>
          )
          <fpage>15</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jumadi</surname>
            , Maylawati,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subaeki</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ridwan</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Opinion mining on twitter microblogging using support vector machine: Public opinion about state islamic university of bandung</article-title>
          .
          <source>In: 2016 4th International Conference on Cyber and IT Service Management</source>
          . (
          <year>2016</year>
          )
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <source>Cortez</source>
          , p.,
          <string-name>
            <surname>Areal</surname>
            ,
            <given-names>N.:</given-names>
          </string-name>
          <article-title>The impact of microblogging data for stock market prediction: Using twitter to predict returns, volatility, trading volume and survey sentiment indices</article-title>
          .
          <source>Expert System with Application</source>
          <volume>73</volume>
          (
          <fpage>125</fpage>
          -144)
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawless</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , N.:
          <article-title>Experimental ir meets multilinguality, multimodality, and interaction</article-title>
          ,
          <source>8th International Conference of the CLEF Association, CLEF</source>
          <year>2017</year>
          , (
          <year>2017</year>
          )
          <fpage>1114</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Murtagh</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Semantic mapping : Towards contextual and trend analysis of behaviours and practices</article-title>
          , Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum (</article-title>
          <year>2016</year>
          )
          <fpage>12071225</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Chaham</surname>
            ,
            <given-names>Y.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scohy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Tweet data mining: the cultural microblog contextualization data set</article-title>
          , Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum (</article-title>
          <year>2016</year>
          )
          <fpage>12461259</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Dogra</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amer</surname>
            ,
            <given-names>N.O.</given-names>
          </string-name>
          :
          <article-title>Lig at clef 2016 cultural microblog contextualization : Timeline illustration based on microblogs pre-processing of the ocial tweet corpus content-based matching</article-title>
          , Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum (</article-title>
          <year>2016</year>
          )
          <fpage>12011206</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Svore</surname>
            ,
            <given-names>K.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burges</surname>
            ,
            <given-names>C.J.:</given-names>
          </string-name>
          <article-title>A machine learning approach for improved bm25 retrieval, Proceeding of the 18th ACM conference on Information and knowledge management (</article-title>
          <year>2009</year>
          )
          <fpage>18111814</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The probabilistic relevance framework: Bm25 and beyond</article-title>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          <volume>3</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Zheng</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
          </string-name>
          , H.:
          <article-title>A comparative study of search result diversication methods</article-title>
          ,
          <source>Proc. of DDR</source>
          (
          <year>2011</year>
          )
          <fpage>5562</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Qi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Object retrieval with image graph traversal-based re-ranking</article-title>
          .
          <source>Signal Processing: Image Communication</source>
          <volume>41</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Thi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngoc</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mothe</surname>
          </string-name>
          , J.:
          <article-title>Building a knowledge base using microblogs : the case of festivals and location-based events</article-title>
          , Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum (</article-title>
          <year>2016</year>
          )
          <fpage>12261237</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Measuring structural similarity in large online networks</article-title>
          .
          <source>Social Science Research</source>
          <volume>59</volume>
          (
          <year>2016</year>
          )
          <article-title>97106 Special issue on Big Data in the Social Sciences</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>AL-Smadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jaradat</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>AL-Ayyoub</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jararweh</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Paraphrase identication and semantic text similarity analysis in arabic news tweets using lexical, syntactic, and semantic features</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>53</volume>
          (
          <year>2017</year>
          )
          <fpage>640</fpage>
          <lpage>652</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Tajbakhsh</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bagherzadeh</surname>
          </string-name>
          , J.:
          <article-title>Microblogging hash tag recommendation system based on semantic tf-idf: Twitter use case</article-title>
          ,
          <source>4th International Conference on Future Internet of Things and Cloud Workshops</source>
          (
          <year>2016</year>
          )
          <fpage>252257</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Pandey</surname>
          </string-name>
          , N.:
          <article-title>Density based clustering for cricket world cup tweets using cosine similarity and time parameter</article-title>
          ,
          <source>2015 Annual IEEE India Conference (INDICON)</source>
          (
          <year>2015</year>
          )
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Kaur</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelowitz</surname>
            ,
            <given-names>C.M.:</given-names>
          </string-name>
          <article-title>A tweet grouping methodology utilizing inter and intra cosine similarity</article-title>
          ,
          <source>2015 IEEE 28th Canadian Conference on Electrical and Computer</source>
          Engineering (CCECE) (
          <year>2015</year>
          )
          <fpage>756759</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Prasetyo</surname>
            ,
            <given-names>V.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winarko</surname>
          </string-name>
          , E.:
          <article-title>Rating of indonesian sinetron based on public opinion in twitter using cosine similarity</article-title>
          ,
          <source>2016 2nd International Conference on Science and Technology-Computer (ICST)</source>
          (
          <year>2016</year>
          )
          <fpage>200205</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>K.D.</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gershman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frederkin</surname>
          </string-name>
          , R.:
          <article-title>Topical clustering of tweets</article-title>
          ,
          <source>Proceedings of the ACM SIGIR: SWSM</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>Da</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.F.F.</given-names>
            ,
            <surname>Coletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.F.S.</given-names>
            ,
            <surname>Hruschka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.R.</given-names>
            ,
            <surname>Hruschka</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.J.:</surname>
          </string-name>
          <article-title>Using unsupervised information to improve semi-supervised tweet sentiment classication</article-title>
          .
          <source>Information Sciences</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>Da</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.F.F.</given-names>
            ,
            <surname>Hruschka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.R.</given-names>
            ,
            <surname>Hruschka</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.J.:</surname>
          </string-name>
          <article-title>Tweet sentiment analysis with classier ensembles. Decision Support Systems (</article-title>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Paltoglou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thelwall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Twitter, myspace, digg: Unsupervised sentiment analysis in social media</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          <volume>3</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Neppalli</surname>
            ,
            <given-names>V.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caragea</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Squicciarini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tapia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stehle</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Sentiment analysis during hurricane sandy in emergency response</article-title>
          .
          <source>International Journal of Disaster Risk Reduction</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Adedoyin-Olowe</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaber</surname>
            ,
            <given-names>M.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dancausa</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stahl</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomes</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          :
          <article-title>A rule dynamics approach to event detection in twitter with its application to sports and politics</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>55</volume>
          (
          <year>2016</year>
          )
          <fpage>351</fpage>
          <lpage>360</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>B.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Redmond</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nason</surname>
            ,
            <given-names>G.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Healy</surname>
            ,
            <given-names>G.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horgan</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heernan</surname>
            ,
            <given-names>E.J.:</given-names>
          </string-name>
          <article-title>The use of twitter by radiology journals: An analysis of twitter activity and impact´ factor</article-title>
          .
          <source>Journal of the American College of Radiology</source>
          <volume>13</volume>
          (
          <year>2016</year>
          )
          <fpage>1391</fpage>
          <lpage>1396</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Cardona-Grau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sorokin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leinwand</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welliver</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Introducing the twitter impact factor: An objective measure of urology's academic impact on twitter</article-title>
          .
          <source>European Urology Focus</source>
          <volume>2</volume>
          (
          <year>2016</year>
          )
          <fpage>412</fpage>
          <lpage>417</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hew</surname>
            ,
            <given-names>K.F.</given-names>
          </string-name>
          :
          <article-title>Using twitter for education: Benecial or simply a waste of time?</article-title>
          <source>Computers &amp; Education</source>
          <volume>106</volume>
          (
          <year>2017</year>
          )
          <fpage>97</fpage>
          <lpage>118</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <given-names>Chandra</given-names>
            <surname>Pandey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Singh Rajpoot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Saraswat</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Twitter sentiment analysis using hybrid cuckoo search method</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>53</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Kayser</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bierwisch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Using twitter for foresight: An opportunity?</article-title>
          <source>Futures</source>
          <volume>84</volume>
          (
          <year>2016</year>
          )
          <fpage>50</fpage>
          <lpage>63</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Ermakova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nie</surname>
          </string-name>
          , J.y.,
          <string-name>
            <surname>Sanjuan</surname>
          </string-name>
          , E.:
          <article-title>Cultural micro-blog contextualization 2016 workshop overview : data and pilot tasks</article-title>
          , Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum (</article-title>
          <year>2016</year>
          )
          <fpage>17971200</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loper</surname>
          </string-name>
          , E.:
          <article-title>Natural Language Processing with Python</article-title>
          .
          <article-title>(</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>H.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Bag-of-concepts: Comprehending document representation through clustering words in distributed representation</article-title>
          .
          <source>Neurocomputing</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A topic-enhanced word embedding for twitter sentiment classication</article-title>
          .
          <source>Information Sciences</source>
          <volume>369</volume>
          (
          <year>2016</year>
          )
          <fpage>188</fpage>
          <lpage>198</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>