<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Subjective Well-Being and Social Media: A Semantically Annotated Twitter Corpus on Fertility and Parenthood</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Public Policy University of Florence</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bocconi University</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Emilio Sulis, Cristina Bosco, Viviana Patti Mirko Lai Dipartimento di Informatica Delia Irazu ́ Herna ́ ndez Far ́ıas University of Turin University of Turin, Italy Italy Univ. Polite`cnica de Vale`ncia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This article describes a Twitter corpus of social media contents in the Subjective Well-Being domain. A multilayered manual annotation for exploring attitudes on fertility and parenthood has been applied. The corpus was further analysed by using sentiment and emotion lexicons in order to highlight relationships between the use of affective language and specific sub-topics in the domain. This analysis is useful to identify features for the development of an automatic tool for sentiment-related classification tasks in this domain. The gold standard is available to the community.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Italiano. L’articolo descrive la creazione
di un corpus tratto da Twitter sui temi del
Subjective Well-Being, fertilita` e
genitorialita`. Un’analisi lessicale ha mostrato il
legame tra l’uso di linguaggio affettivo e
specifiche categorie di messaggi. Questo
esame e` utile per se e per l’addestramento
di sistemi di classificazione automatica sul
dominio. Il gold standard e` disponibile su
richiesta.
are the target of sentiment in the fertility-SWB
domain. The relationship between big data and
official statistics is increasingly a subject of
attention
        <xref ref-type="bibr" rid="ref10 ref17 ref20 ref25">(Mitchell et al., 2013; Reimsbach-Kounatze,
2015; Sulis et al., 2015; Zagheni and Weber,
2005)</xref>
        . In this work we focus on Twitter data
for two main reasons. First, Twitter
individuals’ opinions are posted spontaneously (not
responding to a question) and often as a reaction
to some emotional driven observation. Moreover,
using Twitter we can incorporate additional
measures of attitudes towards children and
parenthood, with a wider geographical coverage than
what is the case for traditional survey. Sentiment
analysis in Twitter has been also used to
monitor political sentiment
        <xref ref-type="bibr" rid="ref22">(Tumasjan et al., 2010)</xref>
        ,
to extract critical information during times of
mass emergency
        <xref ref-type="bibr" rid="ref20 ref23 ref5">(Verma et al., 2011; Buscaldi
and Herna´ndez Far´ıas, 2015)</xref>
        , or to analyse user
stance in political debates on controversial topics
        <xref ref-type="bibr" rid="ref12 ref19 ref19 ref4">(Stranisci et al., 2016; Bosco et al., 2016;
Mohammad et al., 2015)</xref>
        . A comprehensive overview of
sentiment analysis with annotated corpora is
offered in
        <xref ref-type="bibr" rid="ref15 ref4">(Nissim and Patti, 2016)</xref>
        . Focusing on
Italian, among the existing resources we mention
the Senti-TUT corpus
        <xref ref-type="bibr" rid="ref3">(Bosco et al., 2013)</xref>
        and
the TWITA corpus
        <xref ref-type="bibr" rid="ref1 ref11 ref18 ref3">(Basile and Nissim, 2013)</xref>
        that
were recently exploited in the SENTIment
POLarity Classification (SENTIPOLC) shared task
        <xref ref-type="bibr" rid="ref2">(Basile et al., 2014)</xref>
        . The corpus described in this
paper enriches the scenario of datasets available
for Italian, enabling also a finer grained analysis
of sentiment related phenomena in a novel domain
related to parenthood and fertility.
henceforth), which were retrieved through the
Twitter Streaming API and applying the Italian
filter proposed within the TWITA project
        <xref ref-type="bibr" rid="ref1 ref11 ref18 ref3">(Basile
and Nissim, 2013)</xref>
        . The TWITA14 dataset
included 259,893,081 tweets (4,766,342 geotagged).
We applied a multi-step methodology in order to
filter and select those relevant tweets concerning
fertility and parenthood.
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Filtering steps on the dataset</title>
      <p>A number of filtering steps have been applied
for selecting from TWITA14 a corpus of tweets
where users talk about fertility and parenthood
(TW-SWELLFER corpus, henceforth). We could
not rely on the exploitation of one or few
hashtags or other elements that allow identifying posts
on fertility and parenthood. In fact, these
topics are somehow spread in the dataset and
messages may contain relevant information on such
subjects even if the main topic of the post is
different. Therefore, we are facing a situation where,
on the one hand, the set of the data that are
potentially relevant for our specific analysis is wider
than usual; on the other hand, it is more
difficult to identify the presence of information
related to the topics we are interested in. This
leaded us to adopt a multi-step thematic filtering
approach. In a first step (Keyword-based
filtering step), eleven hashtags1 and other 19 keywords
have been chosen for selecting tweets of interest,
including 8 roots to consider diminutives,
singulars and plurals. This list is the result of a
combination of a manual content analysis on 2,500
tweets sampled at completely random (taken as
a starting point) and a linguistic analysis on
synonyms. We obtain a total amount of 3.9
million tweets. A second filtering step consisted in
removing noisy tweets from corpus (User-based
filtering step), as the off-topic ones (messages
not concerning individual expression on fertility
and parenthood topics). Tweets posted by
company/institutions/newspapers accounts have been
deleted. Finally, duplicated tweets not marked as
RT were deleted (Duplicate-based filtering step).
The resulting TW-SWELLFER corpus consists of
2,760,416 tweets.</p>
      <p>1#papa`, #mamma, #babbo, #incinta, #primofiglio,
#secondofiglio, #futuremamme, #maternita`, #paternita`,
#allattamento, #gravidanza
2.2</p>
    </sec>
    <sec id="sec-3">
      <title>Annotation scheme</title>
      <p>
        Given the TW-SWELLFER dataset, we developed
and applied an annotation model aimed at studying
not only the sentiment expressed in the tweets, but
also specific parenthood-related topics discussed
in Twitter that are the target of the sentiment.
To build our annotation model, we relied on a
standard annotation scheme on sentiment
polarity (POLARITY), by exploiting the same labels
POS, NEG, NONE and MIXED provided the
organizers of the shared task for sentiment analysis
in Twitter for Italian
        <xref ref-type="bibr" rid="ref2">(Basile et al., 2014)</xref>
        . Also
the presence/absence of irony has been marked in
order to be able to reason on sentiment polarity
also in case of use of figurative devices.
Annotating the presence of ironic devices is a challenging
task because the inferring process of this figure of
speech does not always lie on semantic and
syntactic elements of texts
        <xref ref-type="bibr" rid="ref18 ref19 ref21 ref7 ref8">(Ghosh et al., 2015; Reyes
et al., 2013; Herna´ndez Far´ıas et al., 2016)</xref>
        , but
often requires contextual knowledge
        <xref ref-type="bibr" rid="ref24">(Wilson, 2006)</xref>
        .
In order to mark irony, we introduced two
polarized ironic labels: HUMNEG, for ironic tweets
with negative polarity, and HUMPOS for ironic
tweets with positive polarity. Finally, a set of
labels marks the specific semantic areas (or
SUBTOPICS) of the tweets related to the parenthood
domain. This part of the annotation scheme is very
important since somehow provides us with a
semantic grid in order to analyse which are the
aspects of parenthood that are discussed on Twitter.
For the annotation of sub-topics we considered 7
labels, suggested by a group three experts on the
SWELLFER (subjective well-being and fertility)
domain, after a manual analysis of a subset of the
tweets:
      </p>
      <p>TOBEPA - To be parents. This tag is
introduced to mark when the user generically
comments about his status of parent.</p>
      <p>TOBESO - To be sons. This tag marks the
sons point of view, i.e. on when the user is
a son that comments on the parent-son
relationship.</p>
      <p>DAILYLIFE - Daily life. This tag marks
messages commenting on recurring situation in
everyday life for what concerns the
relationship between parents and children.</p>
      <p>JUDGOTHERPA - Judgment over other
parents behaviour. The tag allows to mark
comments on educations of children, for
instance comments of behaviours which does
not seems to be appropri - ated for the parent
role.</p>
      <p>FUTURE - Children’ future. This tag is used
for tweets where parents do express
sentiments about the future of children.</p>
      <p>BECOMPA - To become parents. This tag is
introduced to mark tweets where users speak
about the prospect or fear of being parents.
POL - Political side. This tag is introduced to
mark tweets talking about laws having impact
on being parents.</p>
      <p>Finally, two additional tags
(IN-TOPIC/OFFTOPIC) have been added to allow annotators to
mark if the tweet is relevant. The addition of this
tag was necessary because of the noise still present
in the dataset. Moreover, in this way, the manual
annotation will produce also data to be used in
order to create a supervised topic classifier from the
whole TW-SWELLFER corpus. This opens the
way to the exploitation of the corpus for a
finegrained sentiment analysis, by identifying
different aspects and topics of the Twitter debate on
parenthood and the sentiment expressed towards each
aspect/topic.
2.3</p>
    </sec>
    <sec id="sec-4">
      <title>Manual annotation</title>
      <p>
        A random sample of 5,566 tweets from
TWSWELLFER has been collected. On this sample
we applied crowdsourcing for manual annotation
via the Crowdflower platform already used in
literature
        <xref ref-type="bibr" rid="ref13">(Nakov et al., 2016)</xref>
        . We relied on
CrowdFlower controls to exclude unreliable annotators
and spammers based on hidden tests, which we
created by developing a set of gold-standard test
questions equipped with gold reasons2. The
annotator’s task was, first, to mark if the post is IN- or
OFF-TOPIC (or unintelligible), and then to mark
for IN-TOPIC posts, on the one hand, the polarity
and presence of irony, on the other hand, the
subtopics. Precise guidelines were provided to the
annotators. Overall, for each tweet at least three
independent annotations were provided3. In order to
2Test questions resulted from the agreement of three
expert annotators.
      </p>
      <p>
        3We selected the CrowdFlower’s dynamic judgment
option: having the goal of collecting at least 3 reliable
annotations for each tweet, the system was collecting up to a
maximum of 5 annotations (to deal with cases when row’s
conselect the true label we used majority voting.
In-topic vs off-topic: manual annotation on this
aspect resulted in 2,355 in-topic tweets (42.3%)
and 3,136 off-topic (56.3%); the remaining 75
tweets were unknown or null (cases of
disagreement). Thanks to the preliminary filtering steps,
the proportion of in-topic tweets is pretty high
compared to common results from different
Twitter based content and opinion analysis
        <xref ref-type="bibr" rid="ref6">(Ceron et
al., 2014)</xref>
        .
      </p>
      <sec id="sec-4-1">
        <title>Polarity, irony, sub-topics (in-topic tweets): at</title>
        <p>the end of the manual annotation process we
collected 1,545 labeled with the same tags for all the
layers.</p>
        <p>Notice that in the analysis in the next section will
report results also on tweets labeled as IN-TOPIC
after the manual annotation (2,355), but where
annotators did not agree on the polarity, irony and
subtopics labels. We refer to those tweets as
NULL messages.</p>
        <p>Summarizing, the TWSWELLFER-GOLD
corpus includes 1,545 IN-TOPIC tweets labeled with
the same tags for all the layers (POLARITY,
IRONY and SUBTOPICS).
3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Analysis of the gold standard</title>
        <p>Regarding IN-TOPIC tweets, the 26.4% has
been labeled as positive and 22.3% as negative
(See Fig.1), giving us a guidance on what might
be the general feeling in Twitter about the
research topics on happiness and parenthood. The
irony issue is limited to a 15.7% of all the
messages and negative irony prevails (10.1% of
negative ironic tweets and 5.6% of positive ironic
tweets), while neutral tweets are just the 8.3%.
fidence score is low). In our jobs we set 0.7 as minimum
accuracy threshold.</p>
        <p>The amount of mixed tweets is limited to 1.2%
(remaining 26% are labelled as NULL because
of annotators disagreement). Regarding these
results, it appears that positive and negative
feelings towards family, parenthood and fertility
appear more or less equally spread through Twitter
Italy. Even if the positive posts are a little bit
more than the negative ones, ironic tweets must
be considered: most of them are negative ironic
posts (i.e., insulting/damaging the target)
balancing the slight difference between pure positive and
negative tweets. Furthermore, this particular topic,
combined with the Twitter nature which provides
short direct message, discourages people to stand
in the grey (neutral) area, as could happens in other
cases: about the 90% of the tweets shows an
explicit polarity, meaning people take a side and
express their opinions.</p>
        <p>Which are these opinions and about what?
Going further with the analysis and looking also at the
contents, so taking into consideration the “topic
specification attribute and its values (Fig. 2), the
largest category refers to sons tweets (TOBESO)
(40.3%), in which children are discussing and
posting about being children and/or about relating
themselves with parents. Parents tag (TOBEPA)
settles on 15% and becoming tag (BECOMEPA)
on 10%. Remaining categories have minor impact,
all being in between 1% and 6%.
3.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Sentiment and emotion analysis</title>
      <p>The exam of the corpus includes a lexical
analysis on different aspects of affect: sentiment and
emotions. The distribution of terms in each group
of messages reveals interesting patterns. Adopting
sentiment lexical resources4 the whole polarity of
messages is computed summing positive and
negative terms. A normalization is finally performed,
i.e. dividing the polarity value by the number of
terms in each group. In particular, the four lexica
considered count more positive terms in positive
messages. Similarly, negative terms are more
frequent in negative messages. Ironic messages
reveal a similar pattern, even if smoothed. Table 1
presents some of these results.</p>
      <p>
        In addition, the emotion lexicon indicates a
larger frequency of terms related to anger, sadness,
fear and disgust in negative messages than in
positive ones (See Fig. 3). On the contrary, positive
messages contain more terms related to joy,
anticipation and surprise. Some suggestions can be
derived in the comparison of polarity categories and
the corresponding ironic ones. For instance, terms
related to joy are more frequent in ironic negative
messages than in negative ones. It is an insight of
the polarity reversal phenomena, where a shift is
produced by the adoption of a seemingly positive
statement, to reflect a negative one
        <xref ref-type="bibr" rid="ref21">(Sulis et al.,
2016)</xref>
        .
      </p>
      <p>The analysis of topic specification messages
reveals a positive polarity for messages concerning
TOBEPA (to be parents), while BECOMEPA (to
become parents) has a more negative polarity (See
Table 1). Focusing on emotion lexicon, TOBEPA
has an higher incidence of Joy words (Fig. 4).</p>
      <p>
        Messages concerning educations of children
(JUDGOTHERPA) contain a high frequency of
anger and disgust term. The category TOBESO (to
4EmoLex
        <xref ref-type="bibr" rid="ref1 ref11 ref18 ref3">(Mohammad and Turney, 2013)</xref>
        as well as
an own-house Italian version of LIWC
        <xref ref-type="bibr" rid="ref16">(Pennebaker et al.,
2001)</xref>
        , Hu&amp;Liu
        <xref ref-type="bibr" rid="ref9">(Hu and Liu, 2004)</xref>
        , AFINN
        <xref ref-type="bibr" rid="ref14">(Nielsen, 2011)</xref>
        .
Lexicons were translated from English in
        <xref ref-type="bibr" rid="ref20 ref5">(Buscaldi and
Herna´ndez Far´ıas, 2015)</xref>
        .
be sons) is more controversial, having the higher
frequency of negative terms as fear, but also trust,
as well as having the lower frequency of Joy terms.
Coherently, anticipation is more frequent in the
BECOMEPA group of messages. Summarizing,
it seems that children are more critics toward
parents. On the contrary, parents seem express an
attitude more positive towards children.
4
      </p>
      <sec id="sec-5-1">
        <title>Conclusions and Future Work</title>
        <p>The contribution of this paper is the exploration
of opinions and semantic orientation about
fertility and parenthood by scrutinizing about 3 million
Italian tweets. This analysis is useful to identify
features for the development of an automatic
system to address automatic classification tasks in this
domain. The corpus is available to the community.
Its development constitutes a first step and a
precondition to a further analysis that can be applied
on such contents in order to extract, from
semantically enriched data, measures of SWB constructed
in an indirect way. This will hopefully improve
our understanding of attitudes on fertility and
parenthood.</p>
        <p>We are currently extending the corpus by
exploring the very interesting debate around the
“Fertility Day’s initiative” from the Italy’s
Minister of Health Beatrice Lorenzin, which had a
remarkable echo on social media such as Twitter,
with a substantial number of (also sarcastic)
messages with hashtag #fertilityday posted.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Acknowledgments</title>
        <p>The authors gratefully acknowledge financial
support from the European Research Council under
the European ERC Grant Agreement n.
StG313617 (SWELL-FER: Subjective Well-being and
Fertility, P.I. Letizia Mencarini).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          and
          <string-name>
            <given-names>Malvina</given-names>
            <surname>Nissim</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Sentiment analysis on italian tweets</article-title>
          .
          <source>In Proceedings of the 4th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis</source>
          , pages
          <fpage>100</fpage>
          -
          <lpage>107</lpage>
          , Atlanta, Georgia. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Andrea Bolioli, Malvina Nissim, Viviana Patti, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Overview of the Evalita 2014 SENTIment POLarity Classification Task</article-title>
          .
          <source>In Proc. of EVALITA</source>
          <year>2014</year>
          , pages
          <fpage>50</fpage>
          -
          <lpage>57</lpage>
          , Pisa, Italy. Pisa University Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          , Viviana Patti, and
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Bolioli</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Developing corpora for sentiment analysis: The case of irony and senti-tut</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          ,
          <volume>28</volume>
          (
          <issue>2</issue>
          ):
          <fpage>55</fpage>
          -
          <lpage>63</lpage>
          ,
          <year>March</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Bosco</surname>
          </string-name>
          , Mirko Lai, Viviana Patti, and
          <string-name>
            <given-names>Daniela</given-names>
            <surname>Virone</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Tweeting and being ironic in the debate about a political reform: the french annotated corpus twitter-mariagepourtous</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ), pages
          <fpage>1619</fpage>
          -
          <lpage>1626</lpage>
          , Portoroz, Slovenia. ELRA.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Davide</given-names>
            <surname>Buscaldi</surname>
          </string-name>
          and Delia Irazu´ Herna´ndez Far´ıas.
          <year>2015</year>
          .
          <article-title>Sentiment analysis on microblogs for natural disasters management: A study on the 2014 genoa floodings</article-title>
          .
          <source>In Proceedings of the 24th International Conference on World Wide Web, WWW '15 Companion</source>
          , pages
          <fpage>1185</fpage>
          -
          <lpage>1188</lpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Ceron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Curini</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.M.</given-names>
            <surname>Iacus</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Social Media e Sentiment Analysis: L'evoluzione dei fenomeni sociali attraverso la Rete</article-title>
          . SxI - Springer for Innovation / SxI - Springer per l'Innovazione. Springer Milan.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Aniruddha</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Guofu</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Tony</given-names>
            <surname>Veale</surname>
          </string-name>
          , Paolo Rosso, Ekaterina Shutova, Antonio Reyes, and
          <string-name>
            <given-names>Jhon</given-names>
            <surname>Barnden</surname>
          </string-name>
          .
          <year>2015</year>
          . Semeval-2015 task 11:
          <article-title>Sentiment analysis of figurative language in Twitter</article-title>
          .
          <source>In Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval</source>
          <year>2015</year>
          ), pages
          <fpage>470</fpage>
          -
          <lpage>475</lpage>
          , Denver, Colorado, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Delia</given-names>
            <surname>Irazu</surname>
          </string-name>
          ´
          <article-title>Herna´ndez Far´ıas, Viviana Patti</article-title>
          , and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Irony detection in Twitter: The role of affective content</article-title>
          .
          <source>ACM Transaction of Internet Technology</source>
          ,
          <volume>16</volume>
          (
          <issue>3</issue>
          ):
          <volume>19</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          :
          <fpage>24</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Minqing</given-names>
            <surname>Hu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bing</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Mining and summarizing customer reviews</article-title>
          .
          <source>In Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '04</source>
          , pages
          <fpage>168</fpage>
          -
          <lpage>177</lpage>
          , Seattle, WA, USA. ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Mitchell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Frank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. D.</given-names>
            <surname>Harris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Dodds</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Danforth</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The geography of happiness: Connecting Twitter sentiment and expression, demographics, and objective characteristics of place</article-title>
          .
          <source>PLoS ONE</source>
          ,
          <volume>8</volume>
          (
          <issue>5</issue>
          ),
          <fpage>05</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Saif M. Mohammad</surname>
          </string-name>
          and
          <string-name>
            <surname>Peter D. Turney</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Crowdsourcing a Word-Emotion Association Lexicon</article-title>
          .
          <source>Computational Intelligence</source>
          ,
          <volume>29</volume>
          (
          <issue>3</issue>
          ):
          <fpage>436</fpage>
          -
          <lpage>465</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Saif M. Mohammad</surname>
            , Xiaodan Zhu, Svetlana Kiritchenko, and
            <given-names>Joel</given-names>
          </string-name>
          <string-name>
            <surname>Martin</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Sentiment, emotion, purpose, and style in electoral tweets</article-title>
          .
          <source>Information Processing and Management</source>
          ,
          <volume>51</volume>
          :
          <fpage>480</fpage>
          -
          <lpage>499</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Preslav</given-names>
            <surname>Nakov</surname>
          </string-name>
          , Alan Ritter, Sara Rosenthal, Fabrizio Sebastiani, and
          <string-name>
            <given-names>Veselin</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Semeval2016 task 4: Sentiment analysis in twitter</article-title>
          .
          <source>In Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016)</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          , San Diego, California, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Finn</surname>
            <given-names>A</given-names>
          </string-name>
          ˚rup Nielsen.
          <year>2011</year>
          .
          <article-title>A new ANEW: evaluation of a word list for sentiment analysis in microblogs</article-title>
          .
          <source>In Proceedings of the ESWC2011 Workshop on 'Making Sense of Microposts': Big things come in small packages</source>
          , volume
          <volume>718</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <fpage>93</fpage>
          -
          <lpage>98</lpage>
          , Heraklion, Crete, Greece.
          <source>CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Malvina</given-names>
            <surname>Nissim</surname>
          </string-name>
          and
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Semantic aspects in sentiment analysis</article-title>
          .
          <source>In Fersini Elisabetta</source>
          , Bing Liu, Enza Messina, and Federico Pozzi, editors,
          <source>Sentiment Analysis in Social Networks, chapter 3</source>
          , pages
          <fpage>31</fpage>
          -
          <lpage>48</lpage>
          . Elsevier.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>James W. Pennebaker</surname>
            ,
            <given-names>Martha E.</given-names>
          </string-name>
          <string-name>
            <surname>Francis</surname>
            , and
            <given-names>Roger J.</given-names>
          </string-name>
          <string-name>
            <surname>Booth</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Linguistic Inquiry and Word Count: LIWC 2001</article-title>
          . Mahway: Lawrence Erlbaum Associates,
          <fpage>71</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Christian</given-names>
            <surname>Reimsbach-Kounatze</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The proliferation of big data and implications for official statistics and statistical agencies</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Antonio</given-names>
            <surname>Reyes</surname>
          </string-name>
          , Paolo Rosso, and
          <string-name>
            <given-names>Tony</given-names>
            <surname>Veale</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A multidimensional approach for detecting irony in twitter</article-title>
          . Lang. Resour. Eval.,
          <volume>47</volume>
          (
          <issue>1</issue>
          ):
          <fpage>239</fpage>
          -
          <lpage>268</lpage>
          ,
          <year>March</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Stranisci</surname>
          </string-name>
          , Cristina Bosco, Delia Irazu´
          <article-title>Herna´ndez Far´ıas, and</article-title>
          <string-name>
            <given-names>Viviana</given-names>
            <surname>Patti</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Annotating sentiment and irony in the online italian political debate on #labuonascuola</article-title>
          .
          <source>In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ), Paris, France, may.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Emilio</given-names>
            <surname>Sulis</surname>
          </string-name>
          , Mirko Lai, Manuela Vinai, and
          <string-name>
            <given-names>Manuela</given-names>
            <surname>Sanguinetti</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Exploring sentiment in social media and official statistics: a general framework</article-title>
          .
          <source>In Proceedings of the 2nd International Workshop on Emotion and Sentiment in Social and Expressive Media</source>
          , co-located
          <source>with AAMAS</source>
          <year>2015</year>
          , Istanbul, Turkey, May 5,
          <year>2015</year>
          ., volume
          <volume>1351</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <fpage>96</fpage>
          -
          <lpage>105</lpage>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Emilio</given-names>
            <surname>Sulis</surname>
          </string-name>
          , Irazu´ Herna´ndez Far´ıas, Paolo Rosso, Viviana Patti, and
          <string-name>
            <given-names>Giancarlo</given-names>
            <surname>Ruffo</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Figurative messages and affect in twitter: Differences between #irony, #sarcasm and #not</article-title>
          .
          <source>Knowledge-Based Systems</source>
          ,
          <volume>108</volume>
          :
          <fpage>132</fpage>
          -
          <lpage>143</lpage>
          .
          <article-title>New Avenues in Knowledge Bases for Natural Language Processing</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Andranik</given-names>
            <surname>Tumasjan</surname>
          </string-name>
          , Timm Sprenger, Philipp Sandner, and
          <string-name>
            <given-names>Isabell</given-names>
            <surname>Welpe</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Predicting elections with Twitter: What 140 characters reveal about political sentiment</article-title>
          .
          <source>In International AAAI Conference on Web and Social Media.</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Sudha</given-names>
            <surname>Verma</surname>
          </string-name>
          , Sarah Vieweg,
          <string-name>
            <given-names>William</given-names>
            <surname>Corvey</surname>
          </string-name>
          , Leysia Palen, James Martin,
          <string-name>
            <surname>Martha Palmer</surname>
          </string-name>
          , Aaron Schram, and
          <string-name>
            <given-names>Kenneth</given-names>
            <surname>Anderson</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Natural language processing to the rescue? extracting ”situational awareness” tweets during mass emergency</article-title>
          . In
          <source>International AAAI Conference on Web and Social Media.</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Deirdre</given-names>
            <surname>Wilson</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The pragmatics of verbal irony: Echo or pretence?</article-title>
          <source>Lingua</source>
          ,
          <volume>116</volume>
          (
          <issue>10</issue>
          ):
          <fpage>1722</fpage>
          -
          <lpage>1743</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Emilio</given-names>
            <surname>Zagheni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ingmar</given-names>
            <surname>Weber</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Demographic research with non-representative internet data</article-title>
          .
          <source>International Journal of Manpower</source>
          ,
          <volume>36</volume>
          (
          <issue>1</issue>
          ):
          <fpage>13</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>