<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Style of a Successful Story: a Computational Study on the Fanfiction Genre</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Mattei</string-name>
          <email>a.mattei3@studenti.unipi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dominique Brunato</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta</string-name>
          <email>felice.dellorlettag@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Pisa Istituto di Linguistica Computazionale “Antonio Zampolli” (ILC-CNR) ItaliaNLP Lab -</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents a new corpus for the Italian language representative of the fanfiction genre. It comprises about 55k usergenerated stories inspired to the original fantasy saga “Harry Potter” and published on a popular website. The corpus is large enough to support data-driven investigations in many directions, from more traditional studies on language variation aimed at characterizing this genre with respect to more traditional ones, to emerging topics in computational social science such as the identification of factors involved in the success of a story. The latter is the focus of the presented case-study, in which a wide set of multi-level linguistic features has been automatically extracted from a subset of the corpus and analysed in order to detect the ones which significantly discriminate successful from unsuccessful stories</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Computational Sociolinguistics is an emergent
interdisciplinary field aimed at exploiting
computational approaches to study the relationship
between language and society
        <xref ref-type="bibr" rid="ref10">(Nguyen et al., 2016)</xref>
        .
One of the primary factors driving its foundation is
the widespread diffusion of social media and other
user-generated data available online, which has
promoted massive research on computer-mediated
communication from several perspectives. For
instance, scholars working in the field of genre
and register variation have relied on quantitative
approaches to inspect the peculiarities of social
media language, with the purpose of providing
      </p>
      <p>
        Copyright c 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
a characterization of this new genre with respect
to more traditional ones
        <xref ref-type="bibr" rid="ref11 ref7">(Paolillo, 2001; Herring
and Androutsopoulos, 2015)</xref>
        . In the NLP
community, the writing style of user-generated data has
been analyzed through computational stylometry
approaches for addressing tasks broadly related to
author profiling
        <xref ref-type="bibr" rid="ref4">(Daelemans, 2013)</xref>
        , such as
gender and age detection
        <xref ref-type="bibr" rid="ref12 ref8">(Peersman et al., 2011;
Koppel et al., 2002)</xref>
        . The vast majority of this work has
taken into account contents published on few
microblogging platforms considered as more
representative of the contemporary user-generated
mediascape, e.g. Twitter. More recently, the
attention has been oriented to the language used by
online communities whose members share a
common interest towards an object, an activity – and
more in general any area of human interest –
allowing scholars to shed light on the growing
phenomenon of fandom
        <xref ref-type="bibr" rid="ref13">(Sindoni, 2015)</xref>
        . One of the
most prominent expressions of fandom is
fanfiction (fanfic, fic or FF), i.e. fiction written by fans
of a TV series, movie, book etc., using existing
characters and situations to develop new plots. In
many languages dedicated websites exist where
users can publish their own literary works inspired
to the original book they are fans of.
      </p>
      <p>
        From a computational linguistics standpoint,
one perspective from which fanfiction has been
investigated aimed to infer the relationship between
user-generated stories and their original source,
e.g. comparing the representation of characters
according to their gender, as well as to model reader
reactions to stories
        <xref ref-type="bibr" rid="ref14">(Smitha and Bamman, 2016)</xref>
        .
Inspired to that study, which was based on a large
dataset of stories mainly in English, we collect a
new corpus of fanfic stories1, which, to our
knowledge, is the first one for the Italian language. We
rely on this corpus to carry out an investigation
1Terms of service forbid us to distribute this data.
However, the tools used to gather it are available at https:
//github.com/AndreMatte97/Fanfiction
aimed at shedding light on the possibility of
computationally modeling the expected success of a
fanfic story, based on the assumptions of linguistic
profiling and stylometry research.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Dataset collection</title>
      <p>The corpus comprises texts collected from
efpfanfic.net, a portal active since 2001 which allows
users to publish stories and to comment on them.
The website is made up of two sections: one for
original stories and the other for fanfictions. We
considered only the latter and we limited the
collection to stories based on the fantasy saga by the
British writer J.K. Rowling, “Harry Potter”. This
choice was motivated by the main purpose of our
analysis, i.e. characterizing the success of a novel
with respect to its writing style rather than as an
effect of the various subject matters it deals with. At
the same time, the preference given to a very
popular book allowed us to keep a consistent number of
potential readers and reviewers across the corpus,
still having a large sample of texts to analyze. The
data collection was performed through web
scraping, with two spiders written in Python using the
open-source Scrapy framework. The first spider
crawls the list of stories in the category of choice
and extracts their first chapters together with some
metadata, including the URLs of the subsequent
chapters. The second spider takes these addresses
as input and downloads texts and additional
information about all the chapters after the firsts. In the
dataset created this way, the record for each
chapter includes: ID and Reference ID, combinations
used by the website to identify the webpage of
each chapter. We use the ID of the first chapter as
a reference to group together records belonging to
the same story; Title; Rating, an estimate given by
the author about the rawness of themes and scenes
contained in his story; Date of posting; Author’s
nickname; Number of chapters in the story; Text;
Total number of reviews received by the story,
divided in positive, negative and neutral; Number
of reviews received by the single chapter, as well
as the text of the most recent ones. The crawlers
downloaded 54,717 stories, for a total of 19,7310
chapters and a mean of approximately 3.6 chapter
per story, which is consistent with the one
calculated taking into account every entry on the
website. The obtained corpus was divided into folders,
each containing stories with the same number of
chapters.</p>
    </sec>
    <sec id="sec-3">
      <title>The success of a fanfiction story: an exploratory study</title>
      <p>
        Based on the newly created dataset, we carried
out a computational stylometric analysis aimed at
studying whether there is a connection between
the success of a fanfic story and its writing style.
Such a connection has been demonstrated for more
canonical literary works covering novel and movie
domains
        <xref ref-type="bibr" rid="ref15 ref6">(Ganjigunte et al., 2013; Solorio et al.,
2017)</xref>
        , showing that stylometry is a viable
approach also in scenarios different from authorship
attribution and verification.
      </p>
      <p>
        The methodological framework of our
investigation is linguistic profiling
        <xref ref-type="bibr" rid="ref16 ref9">(Montemagni, 2013;
van Halteren, 2004)</xref>
        , a NLP-based approach in
which a large set of linguistically-motivated
features automatically extracted from text are used to
obtain a vector-based representation of it. Such
representations can be then compared across texts
representative of different textual genres and
varieties to identify the peculiarities of each. For
the purpose of our analysis, we split the original
dataset into two varieties corresponding to
“successful” and “unsuccessful” stories. To define
success we follow an approach similar to that used by
Solorio et al. (2017), which is based on the
number of reviews obtained by each story. In this
regard, we decided to include all reviews, not only
the positive ones, which can undoubtedly testify a
favorable attitude by the reader for the story. Two
main reasons motivated our choice: first, we
noticed that the overwhelming majority of collected
reviews are written to convey appreciation, with
just 0.73% among a total of nearly 900k reviews
being negative; therefore, from a statistical point
of view, we can reasonably get rid of the
distinction between various kinds of reviews and simply
take into consideration the overall amount of
feedback received. Secondly, also a negative feedback
proves that a given story has been read and aroused
some interest in the reader. With this in mind, we
define as “unsuccessful” those stories that did not
receive any reviews, thus being largely ignored by
their readers. Conversely, the “successful”
category includes all stories with the same number
of chapters having received a review count higher
than the average of all stories of that length. We
also decided to limit the focus of this analysis to
single-chapter fanfictions written before 2018, so
as to avoid the inclusion of stories not yet
concluded. The resulting classes comprise 2101
unsuccessful texts and 14486 successful ones, with a
threshold for success amounting to 5 reviews.
Table 1 shows an example of stories classified in the
two categories.
      </p>
      <p>
        All texts were pre-processed by means of
regular expressions, with the aim of removing
errors and inconsistencies in the use of
punctuation, capitalization and special characters, in order
to increase the reliability of automatic linguistic
annotation and the process of feature extraction,
which were performed using the Profiling-UD tool
        <xref ref-type="bibr" rid="ref1">(Brunato et al., 2020)</xref>
        .
      </p>
      <p>In what follows we first provide an overview of
the linguistic features used for our statistical
analysis and then we discuss the ones that turned out
to be more prominent in successful writing.
3.1</p>
      <sec id="sec-3-1">
        <title>Linguistic Features</title>
        <p>
          The set of features is based on the one described
in Brunato et al. (2020) and counts more than 150
features, distributed across distinct levels of
linguistic annotation and computed according to the
Universal Dependencies (UD) annotation
framework. These features have be shown to be
effective in a variety of different scenarios, all related to
modeling the ‘form’ of a text, rather than the
content: e.g., from the assessment of sentence
complexity by humans
          <xref ref-type="bibr" rid="ref2">(Brunato et al., 2018)</xref>
          to the
identification of the native language of a speaker
from his/her productions in a second language
(L2)
          <xref ref-type="bibr" rid="ref3">(Cimino et al., 2018)</xref>
          . Specifically, they can
be grouped into the following main phenomena:
        </p>
        <p>Raw Text Features: Document length
computed as the total number of tokens and of
sentences ((#Tokens, #Sentences in Table 2); average
sentence length and token length, calculated in
tokens and in characters, respectively (Sent length,
Word length).</p>
        <p>
          Lexical Richness: Distribution of words and
lemmas belonging to the Basic Italian Vocabulary
          <xref ref-type="bibr" rid="ref5">(De Mauro, 2000)</xref>
          (BIV Tok, BIV Types) and to
the internal repertories (i.e. fundamental, high
usage and high availability, BIV Fund; BIV
HighUS; BIV High-AV); Type/Token Ratio, a feature
of lexical variety computed as the ratio between
the number of lexical types and the number of
tokens in the first 100 and 200 words of text (TTR
Lemma); Lexical density.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Morpho-Syntactic Information: Distribution</title>
        <p>of all grammatical categories, with respect to the
Universal part-of-speech tagset (UPOS * and the
language specific tagset (XPOS *); Distribution of
verbs according to tense, mood and person, both
for main and auxiliar verbs (aux *; V *)).</p>
      </sec>
      <sec id="sec-3-3">
        <title>Verbal Predicate Structure: Average distribu</title>
        <p>tion of verbal roots and of verbal heads for
sentences (VerbHead); features related to the arity of
verbs (i.e. average number of dependents for
verbal head, distribution of verbs by arity).</p>
        <p>Global and Local Parsed Tree: Average depth
of the syntactic tree (MaxDepth); average depth of
embedded complement chains headed by a
preposition; average length of dependency links and of
the maximum link (Links Len; Max Link Length);
relative order of the subject and object with respect
to the verb;</p>
        <p>Syntactic relations: Distribution of typed UD
dependency relations (dep *);</p>
        <p>Use of Subordination: Distribution of main
and subordinate clauses (Main clause, Subord
clause), average length of subordinate chains,
distribution of subordinate chains by length.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Data Analysis</title>
      <p>For each considered feature we calculated the
average value and the standard deviation in the
two classes. We the assessed whether the
variation between mean values is significant using the
Wilcoxon rank sum test. We found that 57% (i.e.
126 out of the 219) of features are differently
distributed in a significant way between successful
and unsuccessful stories. In Table 2 we report an
extract of the most interesting ones.</p>
      <p>As it can be seen, successful stories are on
average longer in terms of number of tokens and
sentences (1, 2), although these sentences are
generally shorter (3), suggesting that readers appreciate
more a plain writing style. However, when lexical
factors are considered, the preference is given to
texts exhibiting less frequent words, as suggested
by the slightly lower distribution of words
belonging to the Basic Italian Vocabulary (5,6) and
especially to the Fundamental one (7). Inflectional
morphology also appears as a domain of
variation between the two classes. Successful
fanfictions employ quite more often verbs in the second
person (15), a feature typical of narrative writing
related to direct speech. On the contrary, we
observe a higher distribution of third person verb,
specifically auxiliaries, both singular (14) and
plural (13), in less successful texts, which can hint at
a preference for reported speech.
Unsuccessful</p>
      <p>Example (Italian)
La citta` di Edimburgo era sommersa da
una cascata d’acqua. Pioveva. Pioveva
da giorni e giorni, senza sosta. Il cielo
era illuminato di lampi e scosso da tuoni.</p>
      <p>Le strade erano vuote. Per la prima
volta da giorni, allo scoccare della
mezzanotte, la pioggia cesso` di colpo. Il
silenzio piombo` sui quartieri che
sembrarono improvvisamente piu` bui. E in
quel silenzio penetrante, l’unico rumore
che si riusciva a distinguere era un
tactac-tac leggero e discontinuo. Proveniva
da una finestra. La finestra di una
lussuosa casa in centro, l’unica luce accesa
a quell’ora. Joanne era davanti al
computer, fonte di quel tremolio e scriveva.</p>
      <p>Batteva le dita sulla tastiera per alcuni
istanti, poi si fermava, rileggeva,
cancellava e riscriveva. Andava avanti cos`ı
da giorni. I suoi occhi erano stanchi, ma
la sua mente lavorava frenetica.
Mancava poco2.</p>
      <p>Il cielo era tetro cosparso di nuvole che
sembravano volere annunciare un
acquazzone, il vento ulula forte facendo
sbattere le finestre violentemente, come
se volesse gridare, liberarsi da una
rabbia repressa. La donna dai lunghi capelli
rosso scuro continuava a fissare la
devastazione attraverso il vetro che ora si era
appannato dal suo stesso respiro. Aveva
lo sguardo malinconico non piu`
illuminato da quella dolce espressione che il
riso le donava. Una mano le si poggio`
sulla spalla e giro` pian piano il volto
verso la persona amata che con un ritmo
lento comincio` ad accarezzarle le gote
che assunsero un colorito roseo alla sua
pelle pallida. Chiuse gli occhi come
per assaporare quel dolce tocco che ora
si era spostato nei suoi capelli. “Non
guardare piu` oltre il vetro” Mormoro` la
voce con una nota di preoccupazione,
apparteneva a James, marito di Lily la
donna dai lunghi capelli rossi3.</p>
      <p>Example (English)
The city of Edimburgh was flooded by
a cascade of water. It was raining. It
had been raining for days and days,
relentlessly. The sky was lit by lightning
and shaken by thunder. The streets were
empty. For the first time in days, at
the stroke of midnight, the rain stopped
abruptly. Silence fell upon the districts
that suddenly seemed darker. And in
that piercing silence, the only noise that
could be recognized was a faint and
irregular tac-tac-tac. It was coming from
a window. The window of a luxurious
house in the city centre, the only light
still on at that time. Joanne was in front
of the computer, source of that
trembling and was writing. She tapped her
fingers on the keyboard for a few
moments, then stopped, reread, deleted and
rewrote. She had been going on like this
for days. Her eyes were tired, but her
mind was working frantically. Almost
there.</p>
      <p>The sky was bleak strewn with clouds
that seemed to want to announce a
downpour, the wind howls loudly
making the windows slam violently, as if
it wanted to scream, to free itself from
a suppressed anger. The woman with
the long dark red hair kept staring the
devastation through the glass that was
now clouded by her own breath. Her
melancholic gaze was no longer lit up
by that sweet look that laughter gave
her. A hand rested on her shoulder and
slowly turned her face towards the loved
one who started slowly caressing her
cheeks which took on a rosy tone on
her pale skin. She closed her eyes, as
if to savor that sweet touch that had now
moved into her hair. “Don’t look beyond
the glass anymore” Whispered the voice
with a note of concern, it belonged to
James, husband of Lily the woman with
long red hair.
ditionally we can see that balanced marks (24), i.e.
syntactic categories, there is a significant
differparenthesis and quotation marks, occur more in
ence in the usage of the most common punctuation
successful texts, strengthening our previous claim
marks, commas (25) and full stops (26), which
about a more frequent presence of direct speech
are quite more frequent in highly-reviewed
fanin this class. At syntactic level, dependency
relafictions. These features relate themselves to the
tions are slightly shorter in successful texts, both
previously observed difference in terms of
docuconsidering the average value of all
dependenment length, as texts with more sentences
necescies (29) and the value of the maximum
depensarily use punctuation marks to divide them.
Addency link (30). In readability assessment
stud2The full story can be found at https://efpfanfic.
net/viewstory.php?sid=607026&amp;i=1</p>
      <p>3The full story can be found at https://efpfanfic.
net/viewstory.php?sid=27412&amp;i=1
ies, longer syntactic dependencies are typically
found in complex texts, and the same holds for
deeper syntactic trees. Both these features have
lower values in highly-reviewed stories,
suggestsignificantly between successful and
unsuccessful stories. All differences are significant at p &lt;
0.001, except for features marked with an asterisk,
which have p &lt; 0.05.
ing that the style of successful writing is
characterized by a simpler syntactic structure.
Interestingly, these results, although preliminary, go in
the opposite direction to those reported by
Ganjigunte et al. (2013) for successful literary works in
English, which where found to be less correlated
with text readability scores. Finally, subordinate
clauses (33) occur slightly more often than main
clauses (32) in unsuccessful texts, while there is a
nearly even split between hypotaxis and parataxis
in successful ones.</p>
      <p>To deepen our analysis, we also computed the
coefficient of variation
* for all features varying
significantly between the two classes, where
* is
the ratio between the standard deviation
mean . This allowed us to evaluate the
dispersion of values around the average in a standardized
way, and thus to compare the stability of features
pertaining to data measured on different scales. A
feature that is much scattered in a class of texts
and highly stable in the other has a greater chance
of being a meaningful representative of the latter.</p>
      <p>In Figure 1 we show the average variability in
the two classes of the four groups of features
distinguished according to the level of annotation
they were extracted from. As a whole, we
noticed that successful texts display less variability
in nearly every considered feature: 117 of them
(92%) are more stable in this class. In successful
stories, features with greater stability compared to
the other class are mainly raw text, e.g. number of
sentences, number of tokens and syntactic ones,
e.g. verbal heads per sentence and average depth
of syntactic trees. Among the few features which
are more stable in poorly received texts, we find
instead verbal predicate features, such as the
distributions of past tenses and of indicative moods,
in addition to the frequency of usage of cardinal
numbers. The set of lexical features is instead the
most stable one for both classes.
class of features, both for successful and
unsuccessful texts.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we presented a NLP-based
stylometric analysis on the emerging genre of
fanfiction aimed at characterizing the writing style of a
successful story. We collected a new large-scale
corpus which – to the best of our knowledge – is
the first one of this genre for Italian. We showed
that successful stories, defined as those receiving
a number of reviews higher that the average, are
characterized by a variety of linguistic features at
different levels of granularity and that these
features are more uniformly distributed within them.</p>
      <p>In the future, we would like to broad the
perspective to other genres in order to study whether
there are linguistic predictors of successful
writing which are constant across different genres, as
well as across concepts somehow similar to
success, such as virality and engagement.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Brunato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cimino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Venturi</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Profiling-UD: a Tool for Linguistic Profiling of Texts</article-title>
          .
          <source>Proceedings of The 12th Language Resources and Evaluation Conference, European Language Resources Association</source>
          ,
          <fpage>7145</fpage>
          -
          <lpage>7151</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Brunato</surname>
          </string-name>
          , L. De Mattei,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Iavarone</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Is this Sentence Difficult? Do you Agree?</article-title>
          <source>Proceedings of Conference on Empirical Methods in Natural Language Processing (EMNLP</source>
          <year>2018</year>
          ),
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Cimino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brunato</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Sentences and Documents in Native Language Identification</article-title>
          .
          <source>Proceedings of 5th Italian Conference on Computational Linguistics (CLiCIT)</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , Turin.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>W.</given-names>
            <surname>Daelemans</surname>
          </string-name>
          .
          <year>2013</year>
          <article-title>Explanation in Computational Stylometry. Gelbukh A. (eds) Computational Linguistics and Intelligent Text Processing</article-title>
          .
          <source>CICLing 2013, Lecture Notes in Computer Science</source>
          , vol
          <volume>7817</volume>
          . Springer, Berlin, Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Tullio De Mauro</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Grande dizionario italiano dell'uso (GRADIT)</article-title>
          . Torino, UTET.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>V.</given-names>
            <surname>Ganjigunte Ashok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Feng</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yejin</given-names>
            <surname>Choi</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Success with style: Using writing style to predict the success of novels</article-title>
          .
          <source>Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <fpage>1753</fpage>
          -
          <lpage>1764</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>S.C.</given-names>
            <surname>Herring</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Androutsopoulos</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Computer-mediated discourse 2.0. The handbook of discourse</article-title>
          , 2nd ed.
          <source>Deborah Tannen</source>
          , Heidi E. Hamilton, Deborah Schiffrin, eds. John Wiley Sons.,
          <volume>1753</volume>
          -
          <fpage>1764</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Koppel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Argamon</surname>
          </string-name>
          and
          <string-name>
            <surname>A. Rachel Shimoni</surname>
          </string-name>
          <year>2002</year>
          .
          <article-title>Automatically Categorizing Written Texts by Author Gender</article-title>
          .
          <source>Lit. Linguistic Comput.</source>
          ,
          <volume>17</volume>
          ,
          <issue>4</issue>
          ,
          <fpage>401</fpage>
          -
          <lpage>412</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Tecnologie linguisticocomputazionali e monitoraggio della lingua italiana</article-title>
          .
          <source>Studi Italiani di Linguistica Teorica e Applicata (SILTA)</source>
          ,
          <fpage>145</fpage>
          -
          <lpage>172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , A.S. Dog˘ruo¨z,
          <string-name>
            <given-names>C.P.</given-names>
            <surname>Rose´</surname>
          </string-name>
          , and F.M.G. de Jong.
          <year>2016</year>
          .
          <article-title>Computational Sociolinguistics: A Survey</article-title>
          .
          <source>Computational Linguistics</source>
          , Vol.
          <volume>42</volume>
          , No.
          <volume>3</volume>
          ,
          <fpage>537</fpage>
          -
          <lpage>593</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>John</given-names>
            <surname>Paolillo</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Language variation on Internet Relay Chat: A social network approach</article-title>
          .
          <source>Journal of Sociolinguistics</source>
          ,
          <volume>5</volume>
          ,
          <fpage>180</fpage>
          -
          <lpage>213</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Peersman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Daelemans</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L. Van</given-names>
            <surname>Vaerenbergh</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Predicting Age and Gender in Online Social Networks</article-title>
          .
          <source>Proceedings of the 3rd International Workshop on Search and Mining User-Generated Contents</source>
          ,
          <fpage>37</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>M.G. Sindoni</surname>
          </string-name>
          <year>2011</year>
          .
          <article-title>'I Really Have No Idea What Non-Fandom People Do with Their Lives.' A Multimodal and Corpus-Based Analysis of Fanfiction</article-title>
          . Lingue e Linguaggi, (
          <volume>13</volume>
          ),
          <year>2015</year>
          ,
          <fpage>277</fpage>
          -
          <lpage>300</lpage>
          , doi.org/10.1285/i22390359v13p277.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Smitha</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Bamman</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Beyond Canonical Texts: A Computational Analysis of Fanfiction</article-title>
          .
          <source>Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP</source>
          <year>2016</year>
          , Austin, Texas, USA, November 1-
          <issue>4</issue>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Solorio</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Montes-y-Go´mez, Suraj Maharjan</article-title>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>Ovalle and Fabio A. Gonza´lez. 2017. Multi-task Approach to Predict Likability of Books</article-title>
          .
          <source>Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics</source>
          ,
          <fpage>1217</fpage>
          -
          <lpage>1227</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>H. van Halteren</surname>
          </string-name>
          <year>2004</year>
          .
          <article-title>Linguistic profiling for author recognition and verification</article-title>
          .
          <source>Proceedings of the Association for Computational Linguistics</source>
          ,
          <fpage>200</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>