<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Opinion Analysis Applied to Politics: A case study based on Twitter</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gilberto Nunes</string-name>
          <email>gilberto.nunes@ifpi.edu.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Denivaldo Lopes and Zair Abdelouahab</string-name>
          <email>denivaldo.lopes@ufma.br</email>
          <email>denivaldo.lopes@ufma.br, zair@dee.ufma.br</email>
          <email>zair@dee.ufma.br</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Federal Institute of Education, Science and Technology, of Piau ́ı - IFPI /</institution>
          ,
          <addr-line>Picos, Piau ́ı</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Federal University of Maranha ̃o - UFMA /, Sa ̃o Lu ́ıs</institution>
          ,
          <addr-line>Maranha ̃o</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <fpage>35</fpage>
      <lpage>42</lpage>
      <abstract>
        <p>1Forecasting Elections: Voter Intentions versus Expectations - Brookings Institution - Link for ebook: http://www.brookings.edu/research/papers/2012/11/01voter-expectations-wolfers.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Nowadays, social networks such as
Facebook and Twitter are openly available
for everyone around the world over the
Internet. These websites provide some
functionality without costs, such as:
creation/edition of communities and social
networks; it provides support to a large
variety of multimedia contents (e.g. audio
and video) and support to interactive
communications (e.g. chats and post).
Twitter’s users post comments about a range
of subjects, such as, products, famous
persons and politics. The dissemination of
the information in these social networks
should be considered due to their global
coverage. An important functionality of
Twitter is the support to georeferenced
posts making the localization of posts
possible. In this paper, we propose an
approach to make the Sentiment Analysis or
Opinion Mining. Our approach is based on
Mining Web and of Opinion, Geographic
Information System (GIS) and Machine
Learning in order to recover relevant
information from tweets. The information
recovered follows our approach is
essential to provide support to the verification
of population trends, e.g. in politics
domain. We propose a prototype that makes
the analysis of population trends, in
special, Brazil’s politic context and the
impeachment process in course.</p>
      <p>Keywords — Web Mining; Opinion Mining;
Machine Learning; Opinion Analysis; Twitter;
Geographic Information System.
1 Introduction
Prasetyo and Hauff (2015), Jungherr (2013) and
Lampos (2012) propose approaches to determine
the voter intention polls based on information
recovered from Twitter.</p>
      <p>In this paper, we propose another approach
based on opinion analysis applied to politics in
order to colect information from Twitter and
determine the public opinion about the current
impeachment process in Brazil that is submitted the
elected president in October 2014.</p>
      <p>During the process of impeachment, as well as
the electoral process, opinion surveys are applied
such as presented by Rothschild1. He says that
this opinion survey is generally based on data
obtained from printed forms filled by the population.
Our approach is based on opinion analysis to
analyze messages obtained from Twitter to determine
the Brazilian population’s opinion about the
impeachment of the Brazil’s president. According to
Currie (1998), impeachment is considered a
process that can result in the removal of a person from
public office after this person has violated the
Constitution of her country.</p>
      <p>In this paper, we show an approach based on
knowledge discovery of textual sources, data from
social networks (e.g. Twitter), Mining Web and of
Opinion, Geographic Information System and
Machine Learning. Applying our proposed approach,
opinion trends about impeachment can be
identified in the Brazilian population.</p>
      <p>This paper is presented as follows. Section 3
presents some fundamental concepts to this
research work. Section 4 presents our approach for
performing opinion analysis from data obtained in
Twitter, with Web Mining support, about
impeachment process development in Brazil. Section 5
presents some results about the impeachment
process development in Brazil. Section 6 shows a
case study according to the impeachment process.
Section 7 presents some conclusions and future
directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>
        Most the related works use sentiment analysis
and opinion mining for evaluate voting
intentions, taking into consideration only the content
the posts. For instance, we use for illustration
purpose the following approaches: Prasetyo and
Hauff
        <xref ref-type="bibr" rid="ref17 ref4">(Dwi Prasetyo and Hauff, 2015)</xref>
        , Jungherr
        <xref ref-type="bibr" rid="ref9">(Jungherr, 2013)</xref>
        and Lampos
        <xref ref-type="bibr" rid="ref11">(Lampos, 2012)</xref>
        , as
will be described below.
      </p>
      <p>
        Prasetyo and Hauff
        <xref ref-type="bibr" rid="ref17 ref4">(Dwi Prasetyo and Hauff,
2015)</xref>
        , propose the use Twitter-based election
forecasting, sentiment analysis and machine
learning techniques to determine voting intention. For
Indonesia’s presidential elections 2014.
      </p>
      <p>
        Jungherr
        <xref ref-type="bibr" rid="ref9">(Jungherr, 2013)</xref>
        , shows a work
using four metrics to determine voting intention,
likewise: the total number hashtags mentioning a
given political party; the dynamics between
mentions positive or negative a given political party;
the total number hashtags mentioning one the
candidates; and the total number users who used
hashtags mentioning a given party or candidate. For
Germany presidential elections 2009.
      </p>
      <p>
        Lampos
        <xref ref-type="bibr" rid="ref11">(Lampos, 2012)</xref>
        , shows a study
techniques and patterns for extracting positive or
negative sentiment from tweets, which build on each
other, through a supervised approach for turning
sentiment into voting intention percentages. For
United Kingdom presidential elections 2010.
      </p>
      <p>Differently from approaches mentioned above,
our work uses georeferenced data, addition to the
textual content. Thus we can easily perform a
spatial analysis, as shown the proposed case study
(vide section 6). In the next section, we described
the technological used in our case study.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Overview</title>
      <p>In this section, we present the subjects Web
Mining, Opinion Mining, Geographic Information
System (GIS), Machine Learning and Twitter.
3.1</p>
      <p>
        Web Mining
Web Mining is a process extracting data or
information from web sources, as described by Zhang
(2011). Second author
        <xref ref-type="bibr" rid="ref18">(Zhang, 2011)</xref>
        , Web
Mining aims to find useful knowledge from Web and
on the basis data mining, text mining, and
multimedia to combine the traditional data mining
techniques with Web. This mining type can be
subdivided in: Web Content Mining, Web Usage
Mining, Web Structure Mining. Mining Web Content
refers to the extraction Web page content, the text
contained on those pages is a good example
contents to be extracted. Web Usage Mining is the
automated recognition user utilization patterns based
on the Web site. Web Structure Mining is based
on interconnection between data or information in
documents or sites by Web. Figure 1 illustrates the
subdivision.
Opinion mining can be defined as a computational
technique that takes care opinion in textual sources
        <xref ref-type="bibr" rid="ref13">(Pang and Lee, 2008)</xref>
        . It aims to extract
information based on sentiment analysis (e.g. positive,
negative and neutral) expressed by one or more
writers and their texts
        <xref ref-type="bibr" rid="ref13">(Pang and Lee, 2008)</xref>
        .
Opinion mining has a process that analyzes a large
volume textual documents that contains a range
subjects, such as, entertainment, politics, education
and marketing. The social networks like Twitter
have supported their users to express and share
opinions and points view. Thus, social networks
can be seen as a large documents volume in
textual source and digital format.
According to Nuhcan (2014), Geographic
Information System (GIS) can be understood as a
computational information system like any other, but
the differential is the database that stores
georeferenced data, i.e. the database includes
latitude/longitude information linked to the data.
Initially, GIS applications were restricted to desktop
computers, but nowadays they are present the Web
(Servers to maps) and the Smartphones (map
applications).
      </p>
      <sec id="sec-3-1">
        <title>Proposed Approach for Opinion</title>
      </sec>
      <sec id="sec-3-2">
        <title>Mining Applied to Politics</title>
        <p>3.4</p>
        <sec id="sec-3-2-1">
          <title>Machine Learning</title>
          <p>
            Machine Learning is a subarea of Artificial
Intelligence where the focus is to develop
computational methods order to provide intelligent
behavior to computers
            <xref ref-type="bibr" rid="ref1">(Arel et al., 2010)</xref>
            . Examples
of Machine Learning are Support Vector Machine
(SVM) (1998), Random Forest (2001) and Naive
Bayes (2006).
3.5
          </p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Twitter</title>
          <p>
            Twitter is a social networking service that enables
the users to send and receive messages
denominated tweets that have 140 character of maximum
size for each post. Twitter has a large number
of content such as profiles, general information,
tweets, emotions, hastags and other
            <xref ref-type="bibr" rid="ref17">(Tiara et al.,
2015)</xref>
            . This social network provides basically two
API2 to support the recovery of data: Search API
and Streaming API. In our approach, we apply the
Twitter API in order to recover the tweets and the
georeferenced location where they were posted.
3.6
          </p>
        </sec>
        <sec id="sec-3-2-3">
          <title>Metrics Evaluation</title>
          <p>
            Once the Twitter data have been collected and
processed, it needs a mechanism to determine the
validity of the classification applied
            <xref ref-type="bibr" rid="ref15 ref9">(Sokolova and
Lapalme, 2009)</xref>
            . Table 1 introduces the confusion
matrix that is used to assist the calculation of the
evaluation.
          </p>
          <p>
            Our proposed approach for opinion mining
applied to politics is based on Knowledge
Discovery in Databases (KDD)
            <xref ref-type="bibr" rid="ref5">(Fayyad et al., 1996)</xref>
            . To
reach the proposed objectives this article, it was
developed an approach that consists of five steps.
This approach has been implemented by a
Software Prototype3 that assists its execution. The
prototype composed to two modules, one for
recovery (Works interconnected to Search API) and
another for the analysis (Works interconnected to
API WEKA) of data. In the first stage, occurs the
acquisition of data (tweets). In the second stage,
preprocessing data, to remove noisy structures. In
the third stage, the feature extraction of the data
using TF-IDF
            <xref ref-type="bibr" rid="ref14">(Robertson, 2004)</xref>
            . In the fourth
step, we have the Text Mining, by means of the
applied of the algorithms to Machine Learning
presented previously. The fifth and last step,
contemplates the evaluation the results obtained through
the analysis of the Confusion Matrix
            <xref ref-type="bibr" rid="ref15 ref9">(Sokolova
and Lapalme, 2009)</xref>
            . Finally, we found the
positive and negative opinions. Figura 2 introduces
the proposed approach which is based on KDD
            <xref ref-type="bibr" rid="ref5">(Fayyad et al., 1996)</xref>
            .
          </p>
          <p>It is worth mentioning the importance of using
the WEKA4 tool and its API during execution of
the steps present in this approach, with the
exception of the data stage acquisition.
4.1</p>
        </sec>
        <sec id="sec-3-2-4">
          <title>Acquisition of datas</title>
          <p>The data (tweets) were recovered using the Search
Twitter API. During the recovery process of tweets
it is necessary Web Mining, specifically the Web
Content Mining, in which recovered the texts
contained in the posts by users the Twitter. A total of
1,218 georeferenced tweets were collected, based
on posts related to impeachment of the President
Dilma, during the March of 2016 which contained
the following hashtags:
#FicaDilma,
#NaoVaiTerGolpe,
#FicaLula,
#ForaDilma,
#ForaLula,
#SouMaisDilma,</p>
          <p>#FicaPT,
#NaoAoGolpe,
#ForaPT, #ForaPTralhas,</p>
          <p>
            #ForaDilmaLulaPT,
3Software Prototype - It is the result of applying a
software process, as defined
            <xref ref-type="bibr" rid="ref16">(Sommerville, 2006)</xref>
            .
          </p>
        </sec>
        <sec id="sec-3-2-5">
          <title>4Machine Learning Group at the University of</title>
          <p>Waikato - Version 3.7.12 and documentation, link:
http://www.cs.waikato.ac.nz/ml/weka/documentation.html.</p>
          <p>The georeferencing of tweets corresponds to the
26 Brazilian state capitals and the Brazilian
Federal District. This step it is performed by recovery
module the prototype made during the search.
4.2</p>
          <p>
            Preprocessing
Before feature extraction of the tweets, it is
important to remove the unwanted structures, such
like: hyperlinks irrelevant words, special
characters, and other references. After removing those,
it is necessary the stemming and normalization
applied on tweets. It is important to emphasize
that the preprocessing occurs in copies of tweets
collected (corpus
            <xref ref-type="bibr" rid="ref10 ref9">(Khairnar and Kinikar, 2013)</xref>
            ).
Such device, it seeks to maintain the original
tweets intact for avoid any inconsistencies. As the
previous step, this step it is also performed by the
recovery module contained in the prototype.
4.3
          </p>
          <p>
            Feature Extraction
After the preprocessing stage, tweets were
submitted to the feature extraction process, through
the TF-IDF
            <xref ref-type="bibr" rid="ref14">(Robertson, 2004)</xref>
            method. Once you
have applied the method of TF-IDF
            <xref ref-type="bibr" rid="ref14">(Robertson,
2004)</xref>
            , the tweets are represented by the matrix of
numeric values (bag-of-words model) as the
mathematical definition of the TF-IDF
            <xref ref-type="bibr" rid="ref14">(Robertson,
2004)</xref>
            model. The Text Mining has a model the
representation using often as feature set, known
as “bag-of-words model”, with the help of the
WEKA4 tool and using your StringToWordVector
method one created the model used this paper. In
this model, documents are represented as a word
vector. Thus, all documents are represented as a
giant document/term matrix. In this paper, TF/IDF
            <xref ref-type="bibr" rid="ref14">(Robertson, 2004)</xref>
            was used as the cell value to
dampen the importance of those terms if it appears
in many documents. This step it is performed by
the analysis module assisted by the prototype.
4.4
          </p>
          <p>Text Mining
Once generated the numeric matrix values, these
values are used as inputs to the classification
algorithms presented previously. These algorithms
are seeking patterns of data interpretable within
the matrix of values for determinate the classes of
the tweets in positive or negative for Dilma’s
impeachment. This step it is also performed by the
analysis module.
4.5</p>
          <p>Evaluation
Lastly, we have the evaluation of the
classification of data the confusion matrix and its metrics.</p>
          <p>
            Providing the obtaining of information, which will
provide the acquisition of knowledge at the end of
the process of KDD
            <xref ref-type="bibr" rid="ref5">(Fayyad et al., 1996)</xref>
            .
• 20% of the samples for training and 80%
of test samples.
5
          </p>
          <p>Results
The Results Section of this research is divided
into three subsections. Subsection 5.1 is
responsible for describing the database that contains the
samples used for training and testing. Subsection
5.2 includes the training models and test.
Subsection 5.3 shows the results for the classification of
tweets.
5.1</p>
          <p>Data Base
This research, the database has 500 positive
samples and 500 negative of tweets to posts related to
the impeachment of the president Dilma. Totaling
1,000 samples in the database. Is worth
emphasizing that the samples were divided only into
positives and negatives, because the neutral samples
have no representativity, as seen during the
experiments. Samples were collected an automatic
manner by Search API, but the labeling process
was performed manually. During manual labeling
it was aimed the selection of samples which had
good representativity for the classification process,
that is the most variable possible. Recalling that
the tweets used this subsection are different from
those used in subsection regarding the Case Study.
These are geo-referenced to the capital and federal
district that make up Brazil and a period of posts
different from the month of March 2016. Thus,
we seek to avoid potential problems in the tweets
classification.
5.2</p>
          <p>Training and Test Models
The generation of training models and test took
place with the help of the WEKA4 tool. Through
this, we used the implementations of algorithms
(SVM, Naive Bayes and Random Forest)
classification, necessary for the creation of models.
Scenarios were generated, respecting the training
models and test as:
• 80% of the samples for training and 20%
of test samples;
• 60% of the samples for training and 40%
of test samples;
• 40% of the samples for training and 60%
of test samples;</p>
          <p>The algorithm that showed the best model was
used in the case study this paper. The results
for the proposed scenario and the best designs for
each algorithm can be viewed in subsection (5.3)
next.
5.3</p>
          <p>Training and Cross-Validation results</p>
          <p>According to Table 5.3, can be checked that the
greatest amount of accuracy was found for the
proportion of 60% - 40%, using the SVM, with a hit
rate 96.9%. While the lowest value was recorded
by the accuracy Naive Bayes with a hit rate of
89.3% for the proportion of 20% - 80%.</p>
          <p>According to the analysis results for Sensitivity
in Table 5.3, we can conclude that the SVM has
the highest rate in relation to the number of true
positive feedback. With a Sensitivity rate of 98.1
% for the proportion of 80% - 20%.</p>
          <p>Analyzing the data in Table 5.3 concerning
Specificity, one can infer that the Random Forest
presents the best result for true negative reviews,
with a Specificity rate of 97.7% for the proportion
of 80% - 20%.</p>
          <p>According to the analysis results for Precision,
Recall and F1-Score in Table 5.3, we can conclude
that the SVM has the highest rate for the metrics
used in cross-validation with 10 folds. With a
Precision rate of 98.5 %, Recall rate of 97.8% and
F1-Score rate of 98.4%.
6</p>
          <p>Case study
In this case study were analyzed a total of 1,218
tweets georeferenced, highlighting that the tweets
not georeferenced were discarded. These posts are
referring to the period of March 2016, linked to
the process of impeachment the president of the
country. This period was selected based on two
large manifestation schedules for the month. The
first manifestation favorable5 to impeachment,
occurred on day 13 and the second contrary6 on day
31.</p>
          <p>
            Figure 3 presents the results to tweets collected
and analyzed in the form of map for the regions of
Brazil, using the approach proposed. Reminding
that for plotting of the map used the GeoServer
(Web Map), as shown in
            <xref ref-type="bibr" rid="ref7">(Huang and Xu, 2011)</xref>
            and based on shapefiles7 to the five regions of
Brazil. These occurrences are posts containing
hashtags cited previously in subsection 4.1.
          </p>
          <p>5Check the location and time of the
demonstrations of March 13 — Congress in focus - Link for
news:
http://congressoemfoco.uol.com.br/noticias/confira-ohorario-e-o-local-das-manifestacoes-de-13-de-marco/.</p>
          <p>6Manifestations against the coup are
scheduled for this Thursday (31/03) - Link for news:
http://www.pragmatismopolitico.com.br/2016/03/manifestac
oes-contra-o-golpe-estao-agendadas-para-esta-quinta-feira3103.html.</p>
          <p>7Shapefiles - It is a well-known format
for storing geospatial resources in files, site:
http://www.esri.com/library/whitepapers/pdfs/shapefile.pdf</p>
          <p>Analyzing Figure 3, one can see that in the
Midwest, Southeast and South map there most records
in favor of impeachment. Assuming the map of
the Northeast region, little more of most records
are of opposed to impeachment. The map of the
northern region is the only one of the five regions
presenting the same results for the reviews.</p>
          <p>It is important to note that other research related
to the impeachment process have already been
carried out since 2015 in Brazil, when the first
evidences to the process. One of those researches are
very similar to the one presented in the research in
this study, being presented in the Veja8 magazine.
In it the magazine exposes results of a research
on social networks by the company Torabit9, in
which 49.3 % of posts on social networks are
favorable to impeachment and only 31.7 % contrary.
Considering the results of the report and the
proposed approach, it can be seen that the present
work presents valid trends in relation to the
impeachment process. It is remarkable that the
proposed work informs trends by region, which does
not happen with the work done by Torabit9.</p>
          <p>Seeking to standardize the presentation of data
in the map plotted by the proposed approach (see
Figure 3) we used a graphic seeking to make it
understandable, as shown in Figure 4.
Analyzing Figure 4, it can be seen, in simplified way,
the percentages by region for each of the opinions,
whether favorable or contrary to the impeachment.</p>
          <p>It can be said that the proposed paper presents
information by regions, which can be proven
through traditional research survey. It happens
because these studies uses past data, while the
proposed work can use past or current data. Monthly
data was used in the study of proposed case in
March 2016. This collection and analysis of daily
data can identify possible trends and allows
targeting of strategic actions in general. These actions
carried out by favorable movements or contrary to
impeachment.</p>
          <p>It is important to note that the case study could
be carried out in relation to other periods for the
Twitter posts. In this new study you can be
dispensed the phases by training and testing, since the
849% of mentions on social networks are
pro-impeachment, study shows — Radar
Online — VEJA.com - Link for news:
http://veja.abril.com.br/blog/radar-on-line/sem-categoria/49das-mencoes-em-redes-sociais-sao-pro-impeachmentmostra-estudo/.</p>
          <p>9Page Home - Torabit - Link for site:
http://www.torabit.com.br/.
models were obtained in the previous study and
the same could be reused for other periods. With
the application of this new study results should
verify possible trends for the process of
impeachment for the selected period.
7</p>
          <p>Conclusion
It is concluded the proposed approach achieved
the goal of providing a solution based on opinion
mining to identify policy trends according to
public opinion. The result obtained with the proposed
work to collect and process data from the Twitter
is valid and resembles with the other work.</p>
          <p>Probably, the results of this study conclude that
the data of social networks, such as the Twitter
(available through its API), can be used for public
opinion research purposes that go beyond a simple
mechanism for broadcast content. Remember that
these networks provide a range of opportunities to
detect where and when a topic of interest is being
discussed. Monitoring on a particular topic and
location, allows researchers to compare it with other
collected data using different means. As it was
shown in Case Study proposed.</p>
          <p>Possibly, the results can be improved through
the use of other methods for feature extraction or
combination of these, such as: Latent Semantic
Indexing Principal Component Analysis and others.
These improvements can come with
implementations of these methods in future work.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>I.</given-names>
            <surname>Arel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.C.</given-names>
            <surname>Rose</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.P.</given-names>
            <surname>Karnowski</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Deep machine learning - a new frontier in artificial intelligence research [research frontier]</article-title>
          .
          <source>Computational Intelligence Magazine</source>
          , IEEE,
          <volume>5</volume>
          (
          <issue>4</issue>
          ):
          <fpage>13</fpage>
          -
          <lpage>18</lpage>
          , Nov.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Leo</given-names>
            <surname>Breiman</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Random forests</article-title>
          . Mach. Learn.,
          <volume>45</volume>
          (
          <issue>1</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          , October.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>David P.</given-names>
            <surname>Currie</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>The first impeachment: The constitution's framers and the case of senator william blount</article-title>
          .
          <source>American Journal of Legal History</source>
          ,
          <volume>42</volume>
          (
          <issue>4</issue>
          ):
          <fpage>427</fpage>
          -
          <lpage>429</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Nugroho</given-names>
            <surname>Dwi</surname>
          </string-name>
          Prasetyo and
          <string-name>
            <given-names>Claudia</given-names>
            <surname>Hauff</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Twitter-based election prediction in the developing world</article-title>
          .
          <source>In Proceedings of the 26th ACM Conference on Hypertext &amp;#</source>
          <volume>38</volume>
          ;
          <string-name>
            <surname>Social</surname>
            <given-names>Media</given-names>
          </string-name>
          ,
          <source>HT '15</source>
          , pages
          <fpage>149</fpage>
          -
          <lpage>158</lpage>
          , New York, NY, USA. ACM.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Usama</given-names>
            <surname>Fayyad</surname>
          </string-name>
          , Gregory Piatetsky-shapiro,
          <source>and Padhraic Smyth</source>
          .
          <year>1996</year>
          .
          <article-title>From data mining to knowledge discovery in databases</article-title>
          .
          <source>AI Magazine</source>
          ,
          <volume>17</volume>
          :
          <fpage>37</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Eibe</given-names>
            <surname>Frank and Remco R. Bouckaert</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Naive bayes for text classification with unbalanced classes</article-title>
          .
          <source>In Proceedings of the 10th European Conference on Principle and Practice of Knowledge Discovery in Databases, PKDD'06</source>
          , pages
          <fpage>503</fpage>
          -
          <lpage>510</lpage>
          , Berlin, Heidelberg. Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A method of using geoserver to publish economy geographical information</article-title>
          .
          <source>In Control, Automation and Systems Engineering (CASE)</source>
          , 2011 International Conference on, pages
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          , July.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Thorsten</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Text categorization with suport vector machines: Learning with many relevant features</article-title>
          .
          <source>In Proceedings of the 10th European Conference on Machine Learning</source>
          ,
          <source>ECML '98</source>
          , pages
          <fpage>137</fpage>
          -
          <lpage>142</lpage>
          , London, UK. Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Jungherr</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Tweets and votes, a special relationship: The 2009 federal election in germany</article-title>
          .
          <source>In Proceedings of the 2Nd Workshop on Politics, Elections and Data</source>
          ,
          <source>PLEAD '13</source>
          , pages
          <fpage>5</fpage>
          -
          <lpage>14</lpage>
          , New York, NY, USA. ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Jayashri</given-names>
            <surname>Khairnar</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mayura</given-names>
            <surname>Kinikar</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Machine learning algorithms for opinion mining and sentiment classification</article-title>
          .
          <source>International Journal of Scientific and Research Publications</source>
          ,
          <volume>3</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Vasileios</given-names>
            <surname>Lampos</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>On voting intentions inference from Twitter content: a case study on UK 2010 General Election</article-title>
          .
          <source>arXiv preprint arXiv:1204</source>
          .
          <fpage>0423</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Mahmut</given-names>
            <surname>Onur</surname>
          </string-name>
          <article-title>Karslıog˘lu Nuhcan Akc¸it</article-title>
          , Emrah Tomur.
          <year>2014</year>
          .
          <article-title>Geographical information systems participating into the pervasive computing</article-title>
          .
          <source>In GEOProcessing</source>
          <year>2014</year>
          ,
          <source>The Sixth International Conference on Advanced Geographic Information Systems</source>
          , Applications, and Services, pages
          <fpage>129</fpage>
          -
          <lpage>137</lpage>
          . ThinkMind, March.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Bo</given-names>
            <surname>Pang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lillian</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Opinion mining and sentiment analysis</article-title>
          .
          <source>Found. Trends Inf. Retr.</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          - 2):
          <fpage>1</fpage>
          -
          <lpage>135</lpage>
          , January.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Robertson</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Understanding inverse document frequency: On theoretical arguments for idf</article-title>
          .
          <source>Journal of Documentation</source>
          ,
          <volume>60</volume>
          (
          <issue>5</issue>
          ):
          <fpage>503</fpage>
          -
          <lpage>520</lpage>
          ,
          <year>July</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Marina</given-names>
            <surname>Sokolova</surname>
          </string-name>
          and
          <string-name>
            <given-names>Guy</given-names>
            <surname>Lapalme</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>A systematic analysis of performance measures for classification tasks</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>45</volume>
          (
          <issue>4</issue>
          ):
          <fpage>427</fpage>
          -
          <lpage>437</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Ian</given-names>
            <surname>Sommerville</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Software Engineering: (Update) (8th Edition) (International Computer Science)</article-title>
          .
          <article-title>Addison-Wesley Longman Publishing Co</article-title>
          ., Inc., Boston, MA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Tiara</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          <string-name>
            <surname>Sabariah</surname>
            , and
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Effendy</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Sentiment analysis on twitter using the combination of lexiconbased and support vector machine for assessing the performance of a television program</article-title>
          . pages
          <fpage>386</fpage>
          -
          <lpage>390</lpage>
          , May.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>The research of web mining in ecommerce</article-title>
          .
          <source>In Management and Service Science (MASS)</source>
          , 2011 International Conference on, pages
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          , Aug.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>