<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Classi cation of Book Blurbs According to the Emotional Tags of the Social Network Zazie</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Valentina Franzoni</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentina Poggioni</string-name>
          <email>poggionig@dmi.unipg.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabiana Zollo</string-name>
          <email>fabiana.zollo@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Mathematics and Computer Science, University of Perugia</institution>
          ,
          <addr-line>Perugia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Sentiment Analysis and Opinion Mining are receiving increasing attention in many sectors because knowing and predicting opinions of people is considered a strategic added value. In the last years an increasing attention has also been devoted to Emotion Recognition, often by developing automated systems that can associate user's emotions to texts, music or artworks. Zazie is an Italian social network for readers that introduces a new dimension on book characterization, the emotional icon tagging. Each book, besides user's comments and reviews, can be tagged with special icons, the moods, that are emotional tags chosen by the users. The aim of this work is to study the feasibility of an automated classi cation of books in Zazie according to the emotional tags, by means of the lexical analysis of book blurbs. A supervised learning approach is used to determine if a correlation between the characteristics of a book blurb and the emotional icons associated to the book by the users exists.</p>
      </abstract>
      <kwd-group>
        <kwd>sentiment analysis</kwd>
        <kwd>emotion recognition</kwd>
        <kwd>automated classication</kwd>
        <kwd>machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In the last years an increasing attention has been addressed to Sentiment
Analysis and Emotion Recognition, often by developing automated systems that can
associate users' emotions to text, music or artwork and can interpret the
subjective nature of emotional content. Several e orts have been devoted to music
emotion classi cation and recognition whether from a research or a commercial
point of view [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The underlying idea in most of these works is to
analyze the track, representing it by means of a number of features and applying
machine learning algorithms to a given dataset to infer models. A critical point
in these works is the choice of the emotional model used to represent emotions
and moods [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]; the choices are generally split in continuous models (e.g., Valence
and Arousal [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) or discrete models (e.g., Ekman categories [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]).
      </p>
      <p>
        Moreover, in the wide scenario of text mining, several attempts have been
made in order to associate emotions and moods to blogs [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], tales [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], newspapers
titles [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] in several domains and contexts.
      </p>
      <p>
        A particular and interesting domain that, to the best of our knowledge, has
not already been investigated is the association of emotions to books. The idea
of building a model for classifying books from an emotional point of view was
born from Zazie [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], a social network for readers that, di erently from other
similar projects e.g., aNobii or Goodreads, introduces a new dimension for books
description, the emotional icon tagging.
      </p>
      <p>Starting from this context, it would be very innovative to provide Zazie users
with an emotion-driven search within the social network. The necessity of such
an automated system arises also from the presence of a lot of books that have
not been tagged yet by the users, for which there is not any information besides
the characteristics stored in the database i.e., title, author or publisher.</p>
      <p>
        The rst step of this research was focused on the selection of relevant
attributes among those usable and available in Zazie to describe a book. We
decided to analyze the book blurb because it can contain relevant emotional
information. Since the blurb is generally written for attracting the reader, it can
emphasize and highlight some book aspects and it can abound with emotional
terms: it seems to be suitable for the automated recognition of emotions. On
the other hand, a possible drawback is the introduction of a bias caused by the
excessive use of words with a high emotional meaning, so the problem is not
trivial. Moreover, the blurb represents an information always available on Zazie,
regardless of user's opinions, reviews or tags. The main original contribution of
this work is to determine if the book blurb re ects the same emotions that the
reader can nd in the book itself. The emotional model used for moods
representation is directly provided by Zazie by means of its emotional tags (icons)
and can be easily correlated to the well known discrete emotional models, such
as the ones de ned by Ekman [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] or Plutchik [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>This paper is organized as follows: in the next section the social network Zazie
and its model for emotional tagging are described; then, in Section 3 related
works are presented. The features extraction phase and the dataset creation are
presented in Sections 4 and 5, while the experimental results are described and
discussed in Section 6. The paper ends with some conclusions and ideas for future
developments and improvements.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Zazie</title>
      <p>Social networks for books recently had a widespread di usion e.g., aNobii,
Shelfari, Bookish, Goodreads with the common aim of creating communities of
readers, allowing them to share their opinions about books and to obtain suggestions
and advices. Zazie is an Italian social network for books, which is constructed on
the same model of aNobii, but di ers from it and from the other social networks
of book readers for the opportunity to share emotions through icon tags: in
addition to classic information as the average vote, readers' reviews and comments,
Zazie provides users with icons, called moods, which introduce a new emotional
dimension.</p>
      <p>Each book in Zazie can be tagged with two moods icons, selected in a set
of 25 di erent climates related to the reader's opinion about the book or to the
emotion induced by the book (7 examples are shown in Figure 1).</p>
      <p>For the aim of this work only a subset of these 25 moods is taken into
account: the attention has been devoted to the icons representing the emotions
induced by the book: angry, cry, love, sad, smile, think.
The di usion of social media has contributed to generate a data ow that
constitutes an important information container and that has determined an increasing
interest in Sentiment Analysis, whose di usion stems from the facility of fruition
of these information and from their numerous applications e.g., for social
behaviour studies, nancial services, social and political events.</p>
      <p>
        In 2013 a preliminary study on automated classi cation of books has been
proposed by the authors in the master thesis [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], where Zazie is presented and
the emotional analysis of the blurb is introduced.
      </p>
      <p>
        In 2005 Gilad Mishne presented a study for the automated classi cation of
blogs, basing on the moods tagged by the authors [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Starting from a huge
collection of posts, Mishne demonstrated how the accuracy increased signi cantly
with the increase of the quantity of data available for training and, although
low, the results did not di er substantially from human performances in
implementing the same task.
      </p>
      <p>
        In the same year, Cecilia Ovesdotter Alm et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] presented the SNoW
architecture and explored the problem of the automated classi cation of 22 fairy
tales of the Grimm brothers by means of Support Vector Machine (SVM) with
respect to the basic emotions by Ekman [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The results of the experiments
were encouraging.
      </p>
      <p>
        In 2008 Rada Mihalcea and Carlo Strapparava [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] also used the Ekman
emotions to present a series of experiments regarding the automatic analysis of
emotions, contained in titles of newspapers. They described the construction of
a large annotated dataset with respect to six basic emotions: anger, disgust, fear,
joy, sadness and surprise. The authors proposed di erent methods of
knowledgebased and corpus-based automated identi cation of these emotions in a text,
trying to determine which were the best.
      </p>
      <p>
        In 2012 Erik Cambria et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] presented a project aimed to supply
employees of the marketing environment for a new social media tool, allowing the
management of semantic information and providing in this way the opportunity to
capture the polarity of opinions and the emotional information associated with
user-generated content. In particular, the authors decided to consider reviews
related to mobile phones, because of the good quality and the large amount of
comments available on the Web. Firstly in Cambria's work the comments have
been analyzed using the Sentic Computing [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] multidisciplinary approach to
Opinion Mining; then, the information has been codi ed for Sentiment
Analysis on the basis of di erent ontologies; nally, the resulting knowledge base was
made available for classi cation using an ad hoc website.
      </p>
      <p>
        In the same year [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] Kirk Roberts et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] introduced a corpus collected
from Twitter with annotated micro-blog posts (i.e., tweets) annotated with seven
emotions: anger, disgust, fear, joy, love, sadness, and surprise, analyzing how
emotions are distributed in the data annotated and comparing it to the
distributions in other emotion-annotated corpora. Moreover, they used the annotated
corpus to train a classi er that automatically discovers the emotions in tweets
and presented an analysis of the linguistic style used for expressing emotions.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] Matteo Baldoni et al. presented ArsEmotica, an application software
for associating the predominant emotions to artistic resources of a social
tagging platform. The aim of the work was to extract a rich emotional semantics of
tagged resources through an ontology driven approach, exploiting and combining
available computational and sentiment lexicons with an ontology of emotional
categories. In [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] Federico Bertola and Viviana Patti presented some
achievements on the topic of social tagging related to artworks and museums within the
ArsEmotica framework. Their focus was on eliciting sharable emotional
meanings from visitors' tags in online collections, by interactively involving the users
of the virtual communities in the process of capturing the latent emotions
behind the tags, relying on methods and tools from a set of disciplines ranging
from Semantic and Social Web to NLP. The aim was the creation of a semantic
social space where artworks can be dynamically organized according to a new
ontology of emotions inspired by the Plutchik's model of human emotions.
      </p>
      <p>
        The Plutchik's model is also used in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], where Jared Suttles and Nancy
Ide present an experiment to identify emotions in tweets, classifying emotions
according to a set of eight basic bipolar emotions de ned by the Plutchik's
wheel of emotions. This allowed to treat the multi-class problem of emotion
classi cation as a binary problem for four opposing emotion pairs. They applied
distant supervision, which has been shown to be an e ective way to overcome
the need for a large set of manually labeled data to produce accurate classi ers.
4
      </p>
    </sec>
    <sec id="sec-3">
      <title>Preprocessing and Features Extraction</title>
      <p>In addition to the features that can be naturally associated to a book, we focused
on the book blurb i.e., the content in the back cover, in order to analyze it from
an emotional point of view, aiming to extract a set of emotions that can represent
the book itself. The choice of taking into account the blurb instead of analysing
directly the whole content of the book derives from the obvious infeasibility of
processing a so long text. Moreover, the entire content of the book could be not
available, because of the dataset or for copyright issues. The idea to verify if
the emotions extracted from the blurb can be relevant compared to the moods
tagged from the users, is promising. However the blurb analysis is not immediate
and di erent factors in uence the choice of methods to be adopted.</p>
      <p>
        Similarly to the case of newspapers or online newsletter titles [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], the blurb
is written for attracting the reader and consequently it makes use of emotional
terms, and seems to be suitable for the automated recognition of emotions. The
blurb is particularly concise, unlike other kinds of texts, as blog posts, tales or
articles; thereby it is not appropriate to think in terms of words frequency or in
terms of measures which are directly correlated to the length of the text.
Therefore MultiWordNet [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], an extension of WordNet [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] including information on
Italian and English words, and in particular WordNet-A ect [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] have been used
in order to extract a series of emotions starting from the blurb, taking advantage
of the existing relations between WordNet synsets.
4.1
      </p>
      <p>Blurb Analysis.</p>
      <p>The blurb analysis has been realized in three main phases: preprocessing,
extraction of emotions and reduction of emotions.</p>
      <p>
        Preprocessing. In this phase a series of passages in order to normalize the text
of the blurb were carried on, obtaining a suitable output for e cient
processing. Firstly a stop words deletion was applied to all the terms of ordinary usage
and not incisive within the analysis process e.g., articles and prepositions. Then,
a phase of tokenization was carried on, extracting words ignoring punctuation
marks and digits. Once drew all the words, it was necessary to reduce each
inected form to its canonical form, called lemma. This procedure, called
lemmatization, has been realized by means of Morph-it! [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], a morphological resource
for the Italian language. In Italian, in fact, there are some more linguistic issues
to face than the English language. For instance, adjectives are declined in many
ways, depending whether they refer to males or females, singular or plural: the
lemmatization phase is basic to reduce the noise due to this variability. When a
certain word can belong to more than one grammatical category, depending on
its role within the sentence (e.g., noun or verb), all the related lemmata are kept.
The output at the end of the preprocessing phase consists in a list of lemmata.
Extraction of Emotions. Once terminated the preprocessing phase, the
extraction of emotions by means of WordNet-A ect was realized. Firstly it was
necessary to retrieve for each lemma the WordNet synsets associated to it, using the
multilinguale lexical databse MultiWordNet.
      </p>
      <p>At this stage the a ective domain WordNet-A ect was exploited in order to
obtain all the emotions associated to the synsets, ltering out the terms which did
not convey a ective information and taking into account multiple occurrences
of the same emotion.</p>
      <p>
        Let us see an example: the blurb of The Count of Montecristo. The emotions
extracted and their frequency are the following:
negative-concern[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], horror[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], anxiety[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], distress[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], enthusiasm[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
negative-fear[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], love[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], affection[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], hate[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], comfortableness[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
Reduction of Emotions. It is necessary to highlight that the emotional hierarchy
of WordNet-A ect is particularly pronged, with 296 nodes, and the result of the
emotion extraction phase can be excessively detailed. For this reason the set of
emotions has to be reduced and two di erent approaches have been implemented
and tested.
      </p>
      <p>
        At rst the reduction has been made to the 32 emotions corresponding to
the third level of hierarchy i.e., the subtree rooted in emotion. Each extracted
emotion was replaced with the emotion reached by climbing the hierarchy up to
the third level consisting of:
1. love, a ection, liking, enthusiasm, gratitude, self-pride, levity, calmness,
fearlessness, positive-expectation, positive-fear, positive-hope, joy (positive-emotion)
2. negative-fear, sadness, general-dislike, ingratitude, shame, compassion,
humility, despair, anxiety, daze (negative-emotion)
3. think, gravity, surprise, ambiguous -agitation, ambiguous-fear, pensiveness,
ambiguous-expectation (ambiguous-emotion)
4. apathy, neutral-unconcern (neutral-emotion)
This issue could be also faced by an ontology driven approach as in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
The related ongoing work is discussed in the last section.
      </p>
      <p>
        The new set of emotions for The Count of Montecristo is:
affection[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], anxiety[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], enthusiasm[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], general-dislike[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], joy[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], love[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
negative-fear[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
      </p>
      <p>
        Then, a further reduction phase can be implemented associating the 32
emotions of the third level of WordNet-A ect to an extended set of Ekman
emotions[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], formed by eight emotional categories happiness, anger, disgust, fear,
sadness, surprise, neutral, ambiguous.
      </p>
      <p>
        Taking in account these eight emotional categories, the set of emotions associated
to the The Count of Montecristo is further reduced and includes:
disgust[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], fear[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], happiness[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Dataset</title>
      <p>Dataset structure. The database provided by Zazie was constituted of 38374 records,
each one representing the association of a tag in the mood set to a book by an user.
Each record is represented by 7 elds (user id, book isbn, book title, book pages,
book publisher, book blurb, mood).
Database ltering. The ltering phase has been implemented in ve steps:
{ Only the books that have received su cient attention from the community
have been selected. This criterion has been applied ltering all the books that
received less than 5 tags i.e., only the books that appear in the database with
at least 5 records have been kept. After this step the database contains 19819
records distributed as shown in Fig. 2.
{ Records are grouped with respect to (book isbn, mood) in order to compute
how many users have chosen the mood tag for the book book isbn. After this
step the database contains 9644 records representing the books having the
structure (book isbn, mood, #occ).
{ A lter based on the standard deviation values allows to select only the
books associated with tags that are e ectively representative. All the books
having value lower than 1.5 were discarded.
{ In order to avoid to associate a mood with too little occurrences, all those
records having the tag frequency less than the arithmetic mean were
discarded. After this phase the database contains 691 records.
{ Only the records having mood value in the set M=fangry, cry, love, sad,
smile, thinkg of the tags we are studying has been selected. After this step
the database contains 236 records distributed as in Fig. 3
Although the initial database provided by Zazie developers was large, the dataset
resulting by applying the techniques described in the previous sections is
relatively small, containing 236 instances. Furthermore, the distribution among the
classes is not uniform: there are two predominant classes (think with 95 instances
and smile with 87 instances), and four minor classes with 54 instances. An
ongoing work using a larger initial database and designing di erent ltering methods
and rules is being implemented.</p>
      <p>A series of preliminary tests was run in order to determine an appropriate
value allowing to select a su cient number of books where some tags were more
popular than others, for example considering the arithmetic mean and the
standard deviation of the moods frequency. It is important to note that this values
depend on the speci c dataset.</p>
      <p>The following tables show an example of discarded and kept books; it is
possible to note that, in table on the left, the frequencies for each tag are all
similar, not allowing to determine a characteristic mood, while, in the table on
the right, a book having a high standard deviation is shown. In this second
case it is clear that the frequencies distribution is more heterogeneous and that
the two most characteristic moods can be easily individuated.</p>
      <p>ISBN Mood Freq Mean ISBN Mood Freq Mean
978804548836 cool 2 1.75 0.50 978806176556 cry 3 10.75 9.39
978804548836 angry 2 1.75 0.50 978806176556 angry 3 10.75 9.39
978804548836 lightning 2 1.75 0.50 978806176556 culture 15 10.75 9.39
978804548836 lol 1 1.75 0.50 978806176556 think 22 10.75 9.39
Experiments were carried on in order to prove if an automated classi cation of
book blurbs based on Zazie emotional tags is possible and can actually be used
with a satisfactory accuracy. However other experiments and the design of other
techniques to build a reliable dataset are ongoing works.</p>
      <p>
        In this group of experiments the classes are identi ed by the selected moods
fsmile, love, sad, think, angry, cryg. Previous experiments presented in [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]
considered the ve most frequent moods i.e., suggest, think, wow, smile and
lightning. The classi cation accuracy, showed in Table 2, was not satisfactory. A
deeper analysis made clear that a motivation could lie in the meaning of the
tags; in fact, tags as wow, lightning and suggest can not be related without
ambiguity to an emotional term, in particular they can be used both for positive
both for negative emotions, and so they do not seem to be suitable to be related
to the emotional content of blurbs.
      </p>
      <p>
        Among the information characterizing a book which is available in the Zazie
database, the author and the emotions extracted by the blurb analysis have been
used as the sample features. Di erent from [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], the publisher and the number
of pages have been discarded, because they are associated to a particular edition
of the book and do not characterize the book as its general literary work.
      </p>
      <p>Therefore, in these experiments, each record in the dataset represents a book
and is characterized by either 34 or 10 features:
{ the author (nominal attribute)
{ the emotions extracted from the blurb (numerical attributes) valued by their
occurrences (32 or 8 depending on which strategy for emotion reduction is
applied)
{ the mood tag (nominal attribute) representing the class attribute
Note that further informations on the books are not used because our aim is to
prove that an automated classi cation on Zazie is possible, with an acceptable
accuracy, using only the information that is actually available in Zazie itself.
An automated classi cation that uses other information and features, even if
important, was out of our goal.</p>
      <p>
        The experiments had been carried on by means of the software Weka [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ],
that supplies for the implementation of many machine learning algorithms and
several measures for the model evaluation.
      </p>
      <p>
        The experimentation has been realized through the cross validation technique
with ten folds using, in particular, algorithms based on decision trees.
Preliminary tests, also reported in [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], were run also using bayesian classi ers. Besides
the good results it has demonstrated, the decision tree approach was preferred
in this group of experiments because it returns a model (i.e. the tree) that is
more readable and analyzable.
      </p>
      <p>Models have been evaluated by the accuracy, recall and precision measures,
de ned as follow:
{ Accuracy = T P=N where T P is the number of instances correctly classi ed
and N is the total number of instances.
{ Recall = N1C Pi=1::NC Recalli where N C is the number of classes, Recalli =
T PiT+PFi Ni and T Pi and F Ni are respectively the instances correctly classi ed
as members of class i and the instances wrongly classi ed as not belonging
to the class i.
{ P recision = N1C Pi=1::NC P recisioni where P recisioni = T PTi+PFi Pi and
F Pi is the number of instances wrongly classi ed as members of class i.</p>
      <p>
        The results for accuracy, recall and precision for the dataset described in
Section 5 are presented in Table 3 where comparisons among the used algorithms
are shown: best accuracy levels are obtained with J48 [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] and BFTree[
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]
algorithms that have equivalent performances, while, also in contrast with the results
contained in Table 2, LADTree [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and RandomForest [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] did not perform as
well as expected. The analysis of motivations of these di erences is ongoing but
preliminary results show that algorithms having an unpruned version perform
better. Moreover it seems, from the results, that there is not a great di erence
between the emotional model with 32 emotions (Tab.3a) and the one with 8
emotions (Tab.3b). It can be noted that, with respect to the results presented
in Table 2, a signi cant improvement was obtained reaching more than 70% of
accuracy.
      </p>
      <p>Algorithm
J48 -U -M 2
BayesNet default
LADTree -B 10
The aim of this work was to study the feasibility of an automated classi cation of
books in Zazie according to the emotional tags by means of the lexical analysis of
book blurbs. A supervised learning approach was used and an experimentation
was implemented.</p>
      <p>Although experiments were preliminary they are very encouraging, especially
considering that improvements are expected from the ongoing works on database
ltering and emotion extraction.</p>
      <p>The blurb is con rmed to be a good source of emotional information about
a book and it actually can be analyzed with the aim of sentiment analysis and
emotion recognition. To the best of our knowledge this is the rst attempt to
apply sentiment analysis to books classi cation.</p>
      <p>Further developments are split into di erent directions.</p>
      <p>On the one hand an improved dataset has to be built: now some classes are
not su ciently represented and most of the misclassi cation errors arises for this
reason.</p>
      <p>
        Also the set of features characterizing the samples has to be extended and
the approach presented in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] using WordNet hypernyms and signi cant word
is under implementation; furthermore an ontology driven approach that uses
the ArsEmotica ontology presented in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] is under consideration, to perform a
di erent emotion extraction phase.
      </p>
      <p>On the other hand, from the perspective of Zazie, a user feedback process
could be implemented in order to con rm or contradict the mood chosen by the
classi er and an emotion-driven search engine could be developed in Zazie.</p>
      <p>After nishing o this promising but still preliminar and explorative
approach, a more general classi er of books, which can process other domains
di erent from Zazie's network, can be also a future goal.</p>
      <p>
        Furthermore, given appropriate tools capable to collect information from the
results of queries on general search engines (e.g., Google, Bing, Yahoo Search) or
specialised repositories for books (e.g., Google Books, Amazon), a further
application could be directed to the use in the preprocessing phase of the web based
proximity measures analysed in [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] and [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] which can return the similarity
between two or more words and have been already applied to semantics-driven
search engines.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>We are grateful to Marco Ghezzi and Zazie developers, Joe and David, for the
collaboration to this research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Y. E.</given-names>
            <surname>Kim</surname>
          </string-name>
          et al.:
          <article-title>Music Emotion Recognition: a State of the Art Review</article-title>
          .
          <source>ISMIR 2010 - 11th International Society for Music Information Retrieval Conference</source>
          , (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.-H.</given-names>
          </string-name>
          <article-title>and</article-title>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H. H.:</given-names>
          </string-name>
          <article-title>Machine recognition of music emotion: A review</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ),
          <volume>1</volume>
          {
          <fpage>30</fpage>
          , (May
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Habu</given-names>
            <surname>Music</surname>
          </string-name>
          . http://www.habumusic.com
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. Stereomood. http://www.stereomood.com</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Barthet</surname>
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Multidisciplinary perspectives on music emotion recognition: Implications for content and context-based models</article-title>
          .
          <source>Proc. CMMR</source>
          ,
          <volume>492</volume>
          {
          <fpage>507</fpage>
          , (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Russell</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>A circumplex model of a ect</article-title>
          .
          <source>J. of Pers. and Social Psy</source>
          .
          <volume>39</volume>
          (
          <issue>6</issue>
          ),
          <volume>1161</volume>
          {
          <fpage>1178</fpage>
          (
          <year>1980</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Thayer</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          :
          <article-title>Multiple indicators of a ective responses to music</article-title>
          .
          <source>Dissert. Abst. Int.</source>
          ,
          <volume>47</volume>
          (
          <issue>12</issue>
          ), (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Gilad</given-names>
            <surname>Mishne</surname>
          </string-name>
          :
          <article-title>Experiments with Mood Classi cation in Blog Posts. Style2005, Stylistic Analysis of Text for Information Access</article-title>
          , (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. C. Ovesdotter Alm, at al.:
          <article-title>Emotions from Text: Machine Learning for Text-based Emotion Prediction</article-title>
          .
          <source>Proc. of HLT and EMNL Conferences</source>
          ,
          <volume>579</volume>
          {
          <fpage>586</fpage>
          , (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Zazie</surname>
          </string-name>
          . http://www.zazie.it
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. Paul Ekman:
          <article-title>Facial Expression</article-title>
          and Emotion. American Psychologist,
          <volume>48</volume>
          (
          <issue>4</issue>
          ),
          <volume>384</volume>
          {
          <fpage>392</fpage>
          , (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>C. Strapparava</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Mihalcea</surname>
          </string-name>
          :
          <article-title>Learning to Identify Emotions in Text</article-title>
          . SAC, (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. E. Cambria,
          <string-name>
            <given-names>M.</given-names>
            <surname>Grassi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hussain</surname>
          </string-name>
          and
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Havasi: Sentic Computing for social media marketing</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          ,
          <volume>59</volume>
          , 557{
          <fpage>577</fpage>
          , (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>M. Baldoni</surname>
          </string-name>
          , et al.:
          <article-title>From tags to emotions: Ontology-driven sentiment analysis in the social semantic web</article-title>
          .
          <source>Intelligenza Arti ciale 6</source>
          (
          <issue>1</issue>
          ):
          <fpage>41</fpage>
          -
          <lpage>54</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>K.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Roach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , J. Guthrie, and
          <string-name>
            <surname>S. M.</surname>
          </string-name>
          <article-title>Harabagiu: Empatweet: Annotating and detecting emotions on Twitter</article-title>
          .
          <source>In Proceedings of the LREC12</source>
          , Istanbul, Turkey, pp.
          <volume>3806</volume>
          {
          <fpage>3813</fpage>
          , (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. H. Liu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lieberman</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ted</given-names>
            <surname>Selker</surname>
          </string-name>
          :
          <article-title>A model of textual a ect sensing using real-world knowledge</article-title>
          .
          <source>IUI</source>
          <year>2003</year>
          :
          <fpage>125</fpage>
          -
          <lpage>132</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>J.</given-names>
            <surname>Suttles</surname>
          </string-name>
          and Nancy Ide:
          <article-title>Distant Supervision for Emotion Classi cation with Discrete Binary Values</article-title>
          .
          <source>CICLing (2)</source>
          :
          <fpage>121</fpage>
          -
          <lpage>136</lpage>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>F.</given-names>
            <surname>Bertola</surname>
          </string-name>
          and
          <string-name>
            <surname>V.</surname>
          </string-name>
          <article-title>Patti: Emotional Responses to Artworks in Online Collections</article-title>
          .
          <source>In Proceedings of PATCH</source>
          <year>2013</year>
          :
          <article-title>Personal Access to Cultural Heritage</article-title>
          ,
          <source>UMAP Workshop</source>
          , volume
          <volume>997</volume>
          <source>of CEUR Workshop Proceedings</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. E.
          <string-name>
            <surname>Pianta</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bentivogli</surname>
            and
            <given-names>C.</given-names>
          </string-name>
          <article-title>Girardi: MultiWordNet: Developing an aligned multilingual database</article-title>
          .
          <source>Proceedings of the 1st International WordNet Conference</source>
          ,
          <volume>293</volume>
          {
          <fpage>302</fpage>
          , (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>George</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>Miller: WordNet: a lexical database for English</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>38</volume>
          (
          <issue>11</issue>
          ),
          <volume>39</volume>
          {
          <fpage>41</fpage>
          , (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>C.</given-names>
            <surname>Strapparava</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Valitutti</surname>
          </string-name>
          ,
          <article-title>WordNet-A ect: an a ective extension of WordNet</article-title>
          .
          <source>In Proc. of 4th International Conference on Language Resources and Evaluation (LREC</source>
          <year>2004</year>
          ),
          <volume>1083</volume>
          {
          <fpage>1086</fpage>
          ,
          <year>2004</year>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. E. Zanchetta and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Baroni: Morph-it! A free corpus-based morphological resource for the Italian language</article-title>
          .
          <source>Corpus Linguistics</source>
          <year>2005</year>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ), University of Birmingham, Birmingham, UK (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. Fabiana Zollo:
          <article-title>Classi cazione automatica di libri rispetto ai tag emozionali del social network Zazie</article-title>
          .
          <source>Master Thesis</source>
          , Department of Mathematics and Computer Science, Universita degli Studi di Perugia,
          <source>Italy</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>M. Hall</surname>
            , E. Frank,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Reutemann</surname>
          </string-name>
          and
          <string-name>
            <surname>Ian H. Witten: The WEKA Data Mining</surname>
          </string-name>
          <article-title>Software: An Update</article-title>
          .
          <source>SIGKDD Explorations</source>
          ,
          <volume>11</volume>
          (
          <issue>1</issue>
          ), (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Ross</surname>
          </string-name>
          <article-title>Quinlan: C4.5: Programs for Machine Learning</article-title>
          . Morgan Kaufmann Publishers, (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26. G. Holmes,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pfahringer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kirkby</surname>
          </string-name>
          , E. Frank, and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Hall: Multiclass alternating decision trees</article-title>
          .
          <source>ECML</source>
          , Springer,
          <volume>161</volume>
          {
          <fpage>172</fpage>
          , (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>J. Friedman</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Hastie</surname>
          </string-name>
          , R. Tibshirani:
          <article-title>Additive logistic regression : A statistical view of boosting</article-title>
          .
          <source>In Annals of statistics</source>
          .
          <volume>28</volume>
          (
          <issue>2</issue>
          ),
          <volume>337</volume>
          {
          <fpage>407</fpage>
          , (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Leo</surname>
          </string-name>
          <article-title>Breiman: Random Forests</article-title>
          .
          <source>In Machine Learning</source>
          .
          <volume>45</volume>
          (
          <issue>1</issue>
          ),
          <volume>5</volume>
          {
          <fpage>32</fpage>
          , (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>M. Grassi</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Cambria</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Hussain</surname>
            , and
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Piazza</surname>
          </string-name>
          .
          <article-title>Sentic web: A new paradigm for managing social media a ective information</article-title>
          .
          <source>Cognitive Computation</source>
          <volume>3</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>480</fpage>
          -
          <lpage>489</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <given-names>V.</given-names>
            <surname>Franzoni</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Milani: PMING Distance: A Collaborative Semantic Proximity Measure</article-title>
          .
          <source>WI{IAT</source>
          , vol.
          <volume>2</volume>
          ,
          <issue>442</issue>
          {
          <fpage>449</fpage>
          , IEEE/WIC/ACM (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>C. H. C. Leung</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Milani</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Franzoni: Collective Evolutionary Concept Distance Based Query Expansion for E ective Web Document Retrieval</article-title>
          .
          <source>Computational Science and Its Applications</source>
          ,
          <volume>657</volume>
          {
          <fpage>672</fpage>
          , Springer, (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>