<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journalism ScientificProse Narrative Educational
1/2 2/3 3/4 4/5 5/6 1/6 1/2 2/3 3/4 4/5 5/6 1/6 1/2 2/3 3/4 4/5 5/6 1/6 1/2 2/3 3/4 4/5 5/6 1/6</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Lost in Text. A Cross-Genre Analysis of Linguistic Phenomena within Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chiara Buongiovanni</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Gracci</string-name>
          <email>f.gracci1g@studenti.unipi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dominique Brunato</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta</string-name>
          <email>felice.dellorlettag@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Pisa</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Moving from the assumption that formal, rather than content features, can be used to detect differences and similarities among textual genres and registers, this paper presents a new approach to the linguistic profiling methodology, which focuses on the internal parts of a text. A case study is presented showing that it is possible to model the degree of variance within texts representative of four traditional genres and two levels of complexity for each.1</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        The combined use of corpus-based and
computational linguistics methods to investigate language
variation has become an established line of
research. The heart of this research is the so-called
‘linguistic profiling’, a technique in which a large
number of counts of linguistic features
automatically extracted from parsed corpora are used as
a text profile and can then be compared to
average profiles for groups of texts
        <xref ref-type="bibr" rid="ref12">(van Halteren,
2004)</xref>
        . Although it has been originally developed
for authorship verification and recognition,
linguistic profiling has been successfully applied to
the study of genre and register variation, following
Biber’s claim that “linguistic features from all
levels function together as underlying dimensions of
variation, with each dimension defining a different
set of linguistic relations among registers”
        <xref ref-type="bibr" rid="ref2">(Biber,
1993)</xref>
        . By modeling the ‘form’ of a text through
large sets of linguistic features extracted from
representative corpora, it has been possible not only
to enhance automatic classification of genres
        <xref ref-type="bibr" rid="ref11">(Stamatatos et al., 2001)</xref>
        , but also to get a better
un1Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
derstanding of the impact of features in classifying
genres and text varieties
        <xref ref-type="bibr" rid="ref5">(Cimino et al., 2017)</xref>
        .
      </p>
      <p>
        This paper moves in this framework but
presents a new approach of linguistic profiling,
in which the unit of analysis is not the document
as a whole entity, but the internal parts in which
it is articulated. In this respect, our perspective
is similar to the one proposed by
        <xref ref-type="bibr" rid="ref7">Crossley et al.
(2011)</xref>
        , who developed a supervised classification
method based on linguistically motivated features
to discriminate paragraphs with a specific
rhetorical purpose within English students’ essays.
However, differently from that work, we focus on
Italian and enlarge the analysis to four traditional
textual genres and two levels of language complexity
for each. The aim is i) to explore to what extent the
internal structure of a text can be modeled via
linguistic features automatically extracted from texts
and ii) to study whether the variance across
different parts of a text changes according to genre and
level of complexity within genre.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Corpora and approach</title>
      <p>
        Our investigation was carried out on four genres:
Journalism, Educational writing, Scientific prose
and Narrative. For each genre, we selected the
two corpora described in Brunato and Dell’Orletta
(2017), which represent a ‘complex’ and a
‘simple’ language variety for that genre, where the
level of complexity was established according to
the expected reader. Specifically, the journalistic
genre comprises a corpus of articles published
between 2000 and 2005 on the general newspaper
La Repubblica and a corpus of easy-to-read
articles from Due Parole, a monthly magazine
written in a controlled language for readers with
basic literacy skills or mild intellectual disabilities
        <xref ref-type="bibr" rid="ref10">(Piemontese, 1996)</xref>
        . The corpus belonging to the
Educational genre is articulated into two
collections targeting high school (AduEdu) vs. primary
school (ChiEdu) students. For the scientific prose,
the ‘complex’ variety is represented by a corpus
of 84 scientific articles on different topics, while
the ‘simple’ one by a corpus of 293 Wikipedia
articles, extracted from the Italian Portal ‘Ecology
and Environment’. For the Narrative genre, we
took a dataset specifically developed for research
on automatic text simplification. It consists of 56
texts covering short novels for children and pieces
of narrative writing for high school L2 students
arranged in a parallel fashion, i.e. for each original
text a manually simplified version is available. For
our study, the original texts and the corresponding
simplified versions were chosen as representative
of the complex variety and the simple variety,
respectively.
      </p>
      <p>
        All corpora were automatically tagged by the
part-of-speech tagger described in Dell’Orletta
(2009) and dependency parsed by the DeSR parser
        <xref ref-type="bibr" rid="ref1">(Attardi et al., 2009)</xref>
        to allow the extraction of
more than 80 linguistic features, on which we
relied to investigate our research questions. These
features (detailed in Section 3) capture
linguistic phenomena of a different nature, with a
focus on morpho–syntactic and syntactic structure,
and were selected since they were proven
effective for genre classification in previous works, as
well as in other scenarios all focused on the
analysis of the ‘form’ of the text rather than its content,
such as linguistic complexity, readability
assessment
        <xref ref-type="bibr" rid="ref6">(Collins-Thompson, 2014)</xref>
        , native language
identification
        <xref ref-type="bibr" rid="ref9">(Malmasi et al., 2017)</xref>
        .
      </p>
      <p>
        As a preliminary step for the analyses, all
documents were split into a fixed number of
sections, where each section is composed by a
certain number of paragraphs, roughly corresponding
to the three main parts of the rhetorical structure
of a text (i.e. introductory, body and
concluding paragraphs). According to the literature, for
some genres, such as academic writing, the
distinction into paragraphs is quite rigid and follows
the so-called ‘five-paragraphs’ format
        <xref ref-type="bibr" rid="ref7">(Crossley et
al., 2011)</xref>
        which adheres to the rhetorical goals
of the document, i.e. the first and the last
paragraph correspond respectively to the introduction
and the conclusion, and the three middle ones to
the body part. However, based on a preliminary
investigation of our corpora we preferred to define
a six-section subdivision in order to avoid
flattening too much the distinctions across genres. The
corpora under analysis indeed are made by
documents which are very different in terms of average
length: for instance, scientific articles are on
average longer than others (184 sentences per
document) and this reflects the fact that the body part
is more dense and possibly articulated into more
middle paragraphs. For each document, the six
sections are thus composed by an average number
of sentences that depends on the document length,
ranging from 2 sentences per section, for the
shortest documents, to 35 for the longest ones.
According to this choice, documents shorter than six
sentences were discarded, thus we finally relied
on a corpus of 1168 documents (see Table 1 for
details). As a result of the stage, we represented
each section of a document as a vector of features,
whose values correspond to the average value that
each feature has in all sentences included in the
section.
      </p>
      <p>In order to understand whether and to what
extent the different parts of a text represent
distinctive varieties with a peculiar linguistic structure,
we carried out two statistical analyses. First, we
assessed whether the difference of the feature
values in each section was statistically significant.
Specifically, we performed a pairwise comparison
between each section and the following one (i.e.
1/2, 2/3, 3/4 etc), as well as between the first and
the last section (i.e. 1/6); the latter was
deliberately aimed at verifying whether our set of features
alone is able to distinguish between the
introductory and the closing part of a document, the two
more distant sections of a text which are supposed
to have a more codified structure. Secondly, we
verified whether there is a correlation between the
values of features in the two sections under
comparison. For both analyses, all data were
calculated across and within genre. The cross-genre
analysis was focused on genre only, thus
considering the two corpora representative of the
complex and simple variety as a unique one for each
genre. In the second scenario, the two corpora
were kept distinct to investigate if there is an effect
of genre that is preserved despite language
complexity changes.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Linguistic features</title>
      <p>The set of features extracted from previously
identified sections are distinguished into three
different categories, according to the level of annotation
from which they derive.</p>
      <p>Raw Text Features: they include the average
word and sentence length (char tok and n tokens
in Table 2), calculated as the number of characters
per token and of tokens per sentence, respectively.</p>
    </sec>
    <sec id="sec-4">
      <title>Morpho-syntactic Features: i.e. distribution</title>
      <p>of unigrams of part-of-speech distinct into 14
coarse-grained pos tags (cpos ) and the 37
finegrained tags (pos ) according to the ISST-TANL
annotation.</p>
      <p>Syntactic Features: these features model
grammatical phenomena of different types, i.e:
- the probability of syntactic dependency types e.g.
subject (dep subj), direct object (dep dobj),
modifiers, calculated as the distribution of each type
out of the total dependency types according to the
ISST-TANL dependency tagset;
- the length of dependency links, i.e. the
average length of all dependency links (each one
calculated as the number of words occurring
between the syntactic head and the dependent)
(avg links l) and of the maximum dependency
link (max links l);
- the order of constituents with respect to the
syntactic head: as a proxy of canonicity effects, it
is calculated the relative position of the subject,
object and adverb with respect to the verbal head
and the position of the adjective with respect to the
nominal head;
- the parse tree structure, in terms of features
calculating: the depth of the whole parse tree
(sent depth) (in terms of the longest path from
the root of the dependency tree to some leaf); the
width of the parse tree (sent width), measured as
the highest number of nodes placed on the same
level; the average number of dependents for all
verbal and nominal heads (avg dependent);
- subordination features: within the group of
syntactic features, a in–depth analysis was devoted
to model subordination phenomena by measuring:
the average distribution of subordinate clauses for
sentence (avg sub clause), the percentage of
subordinate clauses with respect to the main clause (%
sub main) and the percentage of embedded
subordinate clauses, i.e. subordinate clauses
dependent on other embedded subordinate clauses (%
sub minor); for each type, it is also calculated
the average depth (subord depth) and weight
(subord width) of the parse tree generated by the
subordinate clauses and their relative order with
respect to the clause on which they depend.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Data Analysis</title>
      <p>Table 2 illustrates the main findings we obtained.</p>
      <p>Specifically, it shows all features which turned out
to have a statistically significant variation in at
least one of the six pairwise comparisons, or a
correlation score &gt; 0.3 according to the Spearman’s
correlation coefficient. A first clear result is that
the higher number of features varying in a
statistically significant way occurs in the journalistic and
scientific genre, both considered as whole (i.e. row
g for each feature) and with respect to the language
complexity variety (rows s and c). The opposite
trend is reported for educational texts, which is
probably due to the heterogeneous nature of this
genre that includes documents of different textual
typologies (course books, pieces of literature etc.).</p>
      <p>
        If journalism and scientific prose are the two
genres with the highest internal variance, the
comparison between sections allows us to get a better
understanding of this data. Specifically, for both
genres, the majority of significant variations are
observed between the first and the second section
and between the first and the last one. This
suggests that the introduction is a stylistic unit with
a peculiar linguistic structure with respect to the
body and the conclusion. It is characterized e.g.
by shorter sentences (Figure 1), likely due to the
presence of the title in both newspaper and
scientific articles, and by a distinctive distribution
of Parts–of–speech (Figure 2). With this respect,
this data are consistent with other studies in the
literature, e.g.
        <xref ref-type="bibr" rid="ref13">(Voghera, 2005)</xref>
        , and also with
previous findings we obtained on the same
corpora
        <xref ref-type="bibr" rid="ref4">(Brunato et al., 2016)</xref>
        , showing that scientific
prose and newswire texts rely more on the nominal
style. However, with the proposed approach, we
were able to go further in this analysis,
highlighting that noun/verb ratio is always higher in the first
section than all other ones. Besides, at least for
newspaper articles, this feature appears as a genre
marker which is not affected by language
complexity, since the same tendency is observed when
the ‘simple’ and the ‘complex’ corpus are
analyzed independently. The same does not hold for
other features related to syntax and, in particular,
to the use of subordination. In this case, the ‘shift’
between the introduction and the subsequent part
of texts yields significant variations only for
articles of Repubblica. Specifically, the first section
contains less embedded sentences (sent depth: 1st
sect: 5.55; 2nd sect: 7.76), and a lower presence of
subordinate clauses, which appear as structurally
simpler e.g. in terms of depth (subord depth: 1st
sect: 1.67; 2nd sect: 3.5) and width (subord width:
1st sect: 0.94; 2nd sect: 1.97). Conversely, for the
simple variant of this genre (i.e. the articles of the
easy-to-read newspaper 2Parole), we do not
observe significant changes affecting these features;
this is not particularly surprising since
subordination is always less represented in this corpus with
respect to all the other ones.
      </p>
      <p>Leaving aside the similar tendencies
characterizing the introduction, Journalistic and Scientific
prose show a different behavior when we focus
on the internal structure of text. While in this
case much fewer features vary in a significant way,
the majority occurs in the journalistic genre only,
especially between the second and the third
section. Again, they concern a different distribution
of morpho-syntactic categories but also some
syntactic features related to subordination. According
to these data, we can conclude that the journalistic
genre has a more rigorous structure and that it is
possible to capture the boundaries between
different parts by using linguistic features that are not
related to the content of the article.
features</p>
      <p>Rawtextfeatures
g XX X - - - XX XX - - - - XX - - - - - X - - - - -
ntokens s XX - - - - - XX - - - - XX - - - - - X - - - - X
c XX - - - - XX - - - - - - - - - - - X - - - - -
g - - - - - X X - - - - X - - - - - - - - - - -
chartok s - - - - - XX X - - - - XX - - - - - - X - - - -
c - - - - - - - X - - - - - - - - - - - - - - -</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper we have presented a novel approach
to the study of language variation, which
relies on the prerequisites of the linguistic
profiling methodology but with the specific purpose of
modeling the stylistic form of the different parts
within a text. A cross-genre investigation on four
traditional genres in Italian, and two levels of
complexity for each, showed that morpho-syntactic
and syntactic features are differently distributed
across subsections of texts belonging to a
specific genre and language variety. This approach
has important implications for research on genre
variation since it suggests that the
characterization of texts and texts varieties should benefit by
inspecting corpora from this fine-grained
perspective. A better understanding of linguistic
phenomena characterizing the introductory, middle and
conclusive parts of a text is also highly relevant
not only to enhance automatic genre classification
but also for other natural language processing
applications devoted to modeling style: e.g. in
education, as a component of intelligent tutoring
systems able to provide detailed feedback to students
in writing courses or for the automatic generation
of texts with the stylistic properties of a specific
genre and level of complexity.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was partially supported by the
2year project ADA, Automatic Data and
documents Analysis to enhance human-based
processes, funded by Regione Toscana (BANDO
POR FESR 2014-2020).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Attardi</surname>
          </string-name>
          , Felice Dell'Orletta, Maria Simi,
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Turian</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Accurate dependency parsing with a stacked multilayer perceptron</article-title>
          .
          <source>In Proceedings of EVALITA</source>
          <year>2009</year>
          <article-title>- Evaluation of NLP and Speech Tools for Italian 2009</article-title>
          . Reggio Emilia, Italy,
          <year>December 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Douglas</given-names>
            <surname>Biber</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Using register-diversified corpora for general language studies</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>19</volume>
          (
          <issue>2</issue>
          ),
          <fpage>219</fpage>
          -
          <lpage>242</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Brunato and Felice Dell'Orletta</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>On the order of words in Italian: a study on genre vs complexity</article-title>
          .
          <source>International Conference on Dependency Linguistics (Depling</source>
          <year>2017</year>
          ),
          <fpage>18</fpage>
          -
          <lpage>20</lpage>
          September 2017, Pisa, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Brunato</surname>
          </string-name>
          , Felice Dell'Orletta,
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Monitoraggio linguistico di Scritture Brevi: aspetti metodologici e primi risultati. A. Manco e A</article-title>
          . Mancini (eds.),
          <article-title>Scritture Brevi: segni, testi e contesti. Dalle iscrizioni antiche ai tweet, Collana di studi Quaderni di AION-Linguistica, Universita` di Studi di Napoli “L'Orientale”</article-title>
          , Napoli,
          <fpage>149</fpage>
          -
          <lpage>176</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Cimino</surname>
          </string-name>
          , Martijn Wieling, Felice Dell'Orletta, Simonetta Montemagni,
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Identifying Predictive Features for Textual Genre Classification: the Key Role of Syntax</article-title>
          .
          <source>Proceedings of 4th Italian Conference on Computational Linguistics (CLiC-it)</source>
          ,
          <fpage>11</fpage>
          -
          <lpage>13</lpage>
          December,
          <year>2017</year>
          , Rome.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Kevyn</given-names>
            <surname>Collins-Thompson</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Computational Assessment of text readability</article-title>
          .
          <source>Recent Advances in Automatic Readability Assessment and Text Simplification</source>
          . Special issue of
          <source>International Journal of Applied Linguistics</source>
          ,
          <volume>165</volume>
          :
          <fpage>2</fpage>
          , John Benjamins Publishing Company,
          <fpage>97</fpage>
          -
          <lpage>135</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Crossley</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dempsey</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McNamara</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Classifying paragraph types using linguistic features: Is paragraph positioning important</article-title>
          ?
          <source>Journal of Writing Research.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Felice</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Ensemble system for partof-speech tagging</article-title>
          .
          <source>In Proceedings of EVALITA 2009 - Evaluation of NLP and Speech Tools for Italian</source>
          <year>2009</year>
          ,
          <string-name>
            <given-names>Reggio</given-names>
            <surname>Emilia</surname>
          </string-name>
          , Italy,
          <year>December 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Keelan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cahill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tetreault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pugh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hamill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Napolitano</surname>
          </string-name>
          , and
          <string-name>
            <surname>Y. Qian</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>A report on the 2017 native language identification shared task</article-title>
          .
          <source>In Proceedings of the 12th Workshop on Building Educational Applications Using NLP.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Emanuela Piemontese</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>Capire e farsi capire. Teorie e tecniche della scrittura controllata</article-title>
          . Napoli, Tecnodid.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          , Nikos Fakotakis and
          <string-name>
            <given-names>George</given-names>
            <surname>Kokkinakis</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Automatic text categorization in terms of genre and author</article-title>
          .
          <source>Computational Linguistics</source>
          , (
          <volume>26</volume>
          )
          <fpage>471</fpage>
          -
          <lpage>495</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Hans van Halteren</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Linguistic profiling for author recognition and verification</article-title>
          .
          <source>In Proceedings of the Association for Computational Linguistics (ACL04)</source>
          ,
          <fpage>200207</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Miriam</given-names>
            <surname>Voghera</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>La misura delle categorie sintattiche</article-title>
          . In Chiari Isabella / De Mauro Tullio (eds.)
          <article-title>Parole e numeri. Analisi quantitative dei fatti di lingua</article-title>
          , Aracne, Roma,
          <fpage>125</fpage>
          -
          <lpage>138</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>