<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Recurrent Deep Neural Network Model to measure Sentence Complexity for the Italian Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giosu</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lo Bos</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>nni Pil</string-name>
          <email>giovanni.pilato@icar.cnr.it</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dipartimento di Matematica e Informatica, Universita degli studi di Palermo</institution>
          ,
          <country country="IT">ITALY</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ICAR-CNR - National Research Council of Italy</institution>
          ,
          <addr-line>Palermo</addr-line>
          ,
          <country country="IT">ITALY</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Text simpli cation (TS) is a natural language processing task devoted to the modi cation of a text in such a way that the grammar and structure of the phrases is greatly simpli ed, preserving the underlying meaning and information contents. In this paper we give a contribution to the TS eld presenting a deep neural network model able to detect the complexity of italian sentences. In particular, the system gives a score to an input text that identi es the con dence level during the decision making process and that could be interpreted as a measure of the sentence complexity. Experiments have been carried out on one public corpus of Italian texts created speci cally for the task of TS. We have also provided a comparison of our model with a state of the art method used for the same purpose.</p>
      </abstract>
      <kwd-group>
        <kwd>Text Simpli cation Neural Networks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Text simpli cation (TS) is a process that aims at reducing the linguistic complexity
by modifying syntax structure and substituting lemmas in a text. The result of
TS is a new text that keeps the original meaning but that is more easily readable
and understandable.</p>
      <p>
        TS is useful for many di erent kinds of people such as who are not mother
tongue, or have language disabilities, with low educational level and so on. For
example, children a ected by deafness have to face many reading di culties
caused by linguistic problems arisen in their youth [
        <xref ref-type="bibr" rid="ref17 ref19">17, 19</xref>
        ] or people a ected by
dyslexia have comprehension di culties in reading infrequent and long words.
Furthermore, although the increasing investments for school and instruction are
helping the growing of personal culture, there still is a huge percentage of people
with low literal skills that are unable to understand common texts. For example,
Italy is one of the countries with a considerable number of people with low
linguistic competencies [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        Although, there has been a lot of research on TS for English language, there is a
lack of resources for the Italian language that could be used to build TS systems.
Algorithms that work well for English language have lead to poor performance for
the Italian one underling the di erences between the two languages. Despite these
di culties many works have been made trying to face di erent NLP problems
[
        <xref ref-type="bibr" rid="ref1 ref6 ref7">1, 6, 7</xref>
        ] suggesting that much remains to be done for what concern the automatic
analysis of Italian texts.
      </p>
      <p>In this paper, we give a contribution to the TS eld using neural networks
for developing a system capable of classifying Italian sentences in 2 classes
according to their complexity. The system gives a score that represents the level
of con dence during the decision process and that could be interpreted as a
measure of a sentence complexity for low literacy skills readers.</p>
      <p>
        In the domain of TS, as underlined in the work by Shardlow [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], words like
complex and simple should be used keeping in mind that their meaning is relative
to each other and the di culty of a sentence is related to a determined class of
people that could have di erent needs. Unfortunately, since the nature of the
corpus we have used, our system is not specialized for any class of people but
acquire a more general knowledge about complex and simple meaning.
The paper is organized as follow: in section 2 we describe some of the works
related to TS, in section 3 we will describe the system and our approach of
facing the problem, in section 4 we will explain the methodology of carrying out
the tests and results, in section 5 we will give conclusion.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>
        The problem of evaluating a sentence complexity is a research eld that has been
faced using di erent methodologies. Historically, measures have relied on a set
of structural text features like the length of the sentence, the number of words
syllable or the number of characters. For example, for the case of Italian language
Flesch-Vacca [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and GulpEase [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] are the most common used tools to score the
complexity of a text. The former is an adaptation of Flesch-Kincaid measure
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] that is function of the average number of syllables per word and the average
number of sentence words, while the latter is based on the average number of
word characters and the average number of words per sentence. Unfortunately,
it's a common opinion that these classes of indexes are not able to cover the
multiple aspects of complexity, in fact, for example, both indexes consider longer
sentences as more di cult to read, which could be not the real case. The weak
e ciency of such kind of measures has led to the development of new measures
that take into account a simple words dictionary [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Apart from what has been produced so far in the contest of TS, it needs
to consider another level of text complexity that is not only words related,
but take into account the syntactic structure of the sentence in addition of
his lemmas. In fact, some syntactical structures causes a loss of meaning for
person with cognitive impairments such as aphasia. Thus, a speci c assistance
for these people is the reorganization of the sentence structure which makes
it easier to comprehend. A Neural Network (NN) model based on Long Short
Term Memory (LSTM) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] units could be able to evaluate both syntactical
and lexical aspects of complexity, learning the features of complex and simple
sentences autonomously from data.
      </p>
      <p>
        Nowadays, the most important index for assessing sentence complexity is
READIT [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. READ-IT is a classi er based on Support Vector Machine (SVM) which
considers many linguistic features that represent the complexity of the sentence.
READ-IT has been training using "La Repubblica" as examples of complex
sentences and "Due Parole" as example of simple sentences. The SVM receives
as input a vector that represents linguistic features which are divided in three
classes: Lexical Features, Morpho-syntactic Features and Syntactic Features. Each
one of these classes include many measures which describe the sentence under
di erent points of view. For example, Lexical class includes elements related to
the sentence lexical aspects such as the ratio between the number of lexical types
and the number of tokens or the presence of easy terms in the sentence.
Morphosyntactic class contains features that look into the morphology and syntax of the
sentence measuring, amongst other, lexical density and verbal mood. Finally,
Syntactic class contains features that take into account elements only related to
the sentence's syntax such as the depth of parse tree, distribution of subordinate
clause, distribution of main clause and so on.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Proposed methodology</title>
      <p>
        The proposed methodology is able to understand the rules that characterize the
di culty of a sentence. It belongs to the class of Neural Networks [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] which, in
the recent past, have shown good results in many di erent linguistic elds.
Our system is based on a particular class of NNs called Recurrent Neural Networks
(RNNs) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] which ts well the problem of analyzing data sequences. This kind
of networks are able to examine the symbols of the sequence step by step but
taking into account what it has been previously analyzed. Thus, the output of
the network is function of sequence elements but also of their structure.
From the Natural Language Processing (NLP) point of view the sentences can
be structured as a sequence of words and punctuation in which their positions
identify the syntactic structure. In this context, a RNN is able to take into
account both lemmas and syntactic structure to establish the complexity of the
sentence.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Preprocessing</title>
        <p>A sentence is divided into a sequence of tokens. In our model, tokens are words
and punctuation symbols. Splitting a sentence as sequence of tokens is a well
known technique for the representation of a sentence. Even if stop-words and
punctuation are often neglected, in our case all kind of words could a ect the
sentence complexity.</p>
        <p>
          Because of the low number of pairs of sentences in our corpus, we have avoided
to insert an embedding layer able to build his own representation of tokens in
a n-dimensional space, choosing to use a pre-trained word representation. In
particular, each token is converted in 300-dimensional vector through the use of
the dictionary [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The authors have used FastText [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], a library for e cient
learning of word representations and sentence classi cation, trained on Common
Crawl [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] and Wikipedia to create a pre-trained word vector representation for
157 languages, including Italian.
        </p>
        <p>At the end of preprocessing, the sentence is a sequence of 300-dimensional vector
that represents the meaning and the structure of the sentence.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Architecture</title>
        <p>
          Our Network architecture is based on LSTM [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] which has been widely used
for many sequence modeling tasks, thanks to its abilities of facing the problem
of vanishing gradient [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and of remembering the dependencies among elements
inside a sequence which are distant from each other.
        </p>
        <p>Each element of the sequence, in our case the representations of words and
punctuations, is analyzed by a layer of 512 LSTM units. The outcome of this layer
is then processed by fully connected layer composed by two neurons adopting
the softmax as activation function, nally we have applied L2 regularization.
The network architecture is shown in gure 1.</p>
        <p>The softmax function expresses the probability that a sentence belongs to one of
the two classes and it might be interpreted as a cumulative score that measures
the complexity of the sentence taking into account both lemmas and syntactic
structure.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Parameters</title>
        <p>We have observed good results when limiting the source sentences to 20 tokens
and training the network for 8 epochs.</p>
        <p>
          The loss function used is the well known cross-entropy, that has been minimized
choosing the RMSPROP [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] algorithm on balanced minibatch of size 25.
For what concerns the regularization L2 we have used a weight value of 0:01.
The parameters of the network have been obtained through a set of trials. In
detail, we noted that the best results are obtained when training the network
for 8 epochs, that if exceeded cause over tting. For what concerns the choice on
the number of tokens, we have not observed valuable improvements choosing a
number greater than 20.
LSTM
LSTM
LSTM
LSTM
LSTM
LSTM
c
s
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <sec id="sec-4-1">
        <title>Corpus</title>
        <p>
          Since there is a lack of corpus for the Italian language suited for training machine
learning algorithms we have used, to our knowledge, the biggest available dataset
created for italian text simpli cation [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>The corpus contains about 63:000 pairs of sentences in which each original
sentence has the corresponding translation in its simpli ed form. The paired
sentences containing structural transformations that identify how to simplify a
sentence. Although the corpus has been created using an automatic procedure,
it has been tested deeply with the help of the humans specialists. Some of
simpli cation rules inside the corpus are deletion or substitution of words from a
source sentence that make it di cult to understand, insertion of other words that
can explain better the meaning of the sentence, reordering of complex sentence
words in which lemmas position is changed in order to make the sentence more
easily understandable.</p>
        <p>We have trained the NN model used as simple sentences examples all the simpli ed
sentences in the corpus, while, the others are used as complex sentences examples.
The set of sentences allow the NN to understand what are the common patterns
in the complex class and in the simple class.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Experiments</title>
        <p>Since few pairs of sentences are present in our corpus, we decided to use the
K-FOLD cross-validation (K-FOLD) with K = 10 to evaluate our approach.
K-FOLD is a validation method for assessing the abilities of a statistical model
in order to generalize his knowledge to an independent dataset. It partitions
randomly the dataset into K equal sized subsets: the method select K-1 subsets
that are used to train the model while the last one is used to validate it.
We have trained K models considering two classes of sentences: in need of
simpli cation (class positive), simpli ed (class negative).</p>
        <p>To quantify the obtained results we have calculated Precision, Recall, True
Positive Ratio (TPR) and True Negative Ratio (TNR) for each iteration of
KFOLD. Recall and Precision give information about the percentage of positive
class elements that the model is able to correctly classify and how many times
it makes mistakes labelling an element as belonging to the positive class. TPR3
and TNR express informations about how e ective is the classi er for identifying
the correct classes for elements of both classes. Finally, the results have been
averaged on the K executed iterations. Table 1 shows the results obtained by
our Network.</p>
        <p>RECALL PRECISION True Positive Ratio True Negative Ratio</p>
        <p>0.83 0.86 0.83 0.87</p>
        <p>
          To deeper understand the potentiality of our system, we have compared its
performance with READ-IT [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Both our network and READ-IT are created
to classify sentences based on their features. In detail they o er a score that is
the sentence probability of belonging to one of the two classes. A score having
a value close to 1 represents a high probability that the sentence needs to be
simpli ed, otherwise a score close to 0 identifying no need of simpli cation.
Note that the READ-IT tool outputs only the sentence probability of beeing
complex. For this reason, it is possible to compute an accuracy measure such
as precision, recall, TPR and TNR only when setting a threshold value for the
probability. Due to the di culty in estimating such value, we have decided to
compare our method with READ-IT by using the Area Under the Receiver
Operating Characteristic (AUROC) curve [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. The receiver Operating Characteristic
(ROC) curve is a graphical plot that shows the classi cation abilities of a binary
classi er when its discrimination threshold changes. If T is the set of thresholds
that contains equal separated elements from 0 to 1, 8t 2 T the ROC curve
plots True Positive Rate (on the y axes) and False Positive Rate (in the x
axes) calculated taking t as discrimination threshold. The AUROC represents the
expectation that a uniformly drawn random positive is ranked before a uniformly
drawn random negative.
3 TPR is calculated at the same way of RECALL.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Discussion</title>
        <p>The comparison in Figure 2 shows high values of AUROC for both models; in
details our model achieves performances comparable to the READIT ones by
showing noteworthy results for the same corpus with the same conditions.</p>
        <p>ROC CURVE ON AVERAGE FOR READIT AREA=0.915</p>
        <p>ROC CURVE ON AVERAGE FOR RNN AREA=0.921
0.2
0.4 FPR 0.6
0.8
1.0
0.2
0.8</p>
        <p>1.0
0.4 FPR 0.6
Although our system and READ-IT reach almost the same results they are
deeply di erent. The main di erence is due to the input data that the models
accept. Our model is able to di erentiate the classes of sentences taking as input
the only raw texts. Otherwise, READ-IT needs to calculate many measures that
are given as input to the SVM model.</p>
        <p>Our idea is to build an easy-to-use model able to learn, on its own, linguistic
features that identi es simple and complex objects building its knowledge on
the dataset of annotated italian sentences. In addition, we think that the use of
NN as the main model for understanding the complexity of sentences can better
emulate the reasoning of human beings.</p>
        <p>This model is open to a variety of application but we think that one of the most
important is as a measure for reliable evaluation of automated text-simpli cation
methodologies.</p>
        <p>The lack of our model is related, like all NN-based approaches, to the di culty of
understanding the behavior of the NN. Its structure does not allow to understand
which are the features of a sentence that make it complex, hence at this moment
the model is not able to advice actions for simplify the sentence.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We have introduced a neuronal system able to automatically detect features
that make italian sentences complex or simple for low literacy skills readers.
The approach is completely data driven and makes use of raw text only. The
main application of the system is that it provides a measure to quantify the
complexity of a sentence in order to detect if a simpli cation is needed. In
addition, our system can be used as an embedded part of another system as
a stand-alone module.</p>
      <p>According to our test the system represents a good alternative to measure the
sentence complexity level since its performance are coherent with the READ-IT
ones. In spite of the system is not able to advice a simpli cation strategy we
believe that it could represents a core component of a more complex system able
to automatically simplify a generic text.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alfano</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lenzitti</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lo Bosco</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perticone</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>An automatic system for helping health consumers to understand medical texts</article-title>
          . pp.
          <volume>622</volume>
          {
          <issue>627</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bishop</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          :
          <article-title>Neural Networks for Pattern Recognition</article-title>
          . Oxford University Press, Inc., New York, NY, USA (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>arXiv preprint arXiv:1607.04606</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Brunato</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dell'Orletta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venturi</surname>
          </string-name>
          , G.:
          <article-title>Paccss-it: A parallel corpus of complex-simple sentences for automatic text simpli cation</article-title>
          .
          <source>In: Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <volume>351</volume>
          {
          <fpage>361</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chall</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dale</surname>
          </string-name>
          , E.:
          <article-title>Readability revisited: The new Dale-Chall readability formula</article-title>
          .
          <source>Brookline Books</source>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chiavetta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lo Bosco</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pilato</surname>
          </string-name>
          , G.:
          <article-title>A lexicon-based approach for sentiment classi cation of amazon books reviews in italian language</article-title>
          . vol.
          <volume>2</volume>
          , pp.
          <volume>159</volume>
          {
          <issue>170</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Chiavetta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lo Bosco</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pilato</surname>
          </string-name>
          , G.:
          <article-title>A layered architecture for sentiment classi cation of products reviews in italian language</article-title>
          . In: Monfort,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Krempels</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.H.</given-names>
            ,
            <surname>Majchrzak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.A.</given-names>
            ,
            <surname>Traverso</surname>
          </string-name>
          , P. (eds.)
          <source>Web Information Systems and Technologies</source>
          . pp.
          <volume>120</volume>
          {
          <fpage>141</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Dell'Orletta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montemagni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venturi</surname>
          </string-name>
          , G.:
          <article-title>Read-it: Assessing readability of italian texts with a view to text simpli cation</article-title>
          .
          <source>In: Proceedings of the second workshop on speech and language processing for assistive technologies</source>
          . pp.
          <volume>73</volume>
          {
          <fpage>83</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Franchina</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vacca</surname>
          </string-name>
          , R.:
          <article-title>Adaptation of esh readability index on a bilingual text written by the same author both in italian and english languages</article-title>
          .
          <source>Linguaggi</source>
          <volume>3</volume>
          ,
          <issue>47</issue>
          {
          <fpage>49</fpage>
          (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Goodfellow</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Courville</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Deep Learning</article-title>
          . MIT Press (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Learning word vectors for 157 languages</article-title>
          . arXiv preprint arXiv:
          <year>1802</year>
          .
          <volume>06893</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swersky</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Neural networks for machine learning lecture 6a overview of mini-batch gradient descent (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9(8)</source>
          ,
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ling</surname>
            ,
            <given-names>C.X.</given-names>
          </string-name>
          :
          <article-title>Using auc and accuracy in evaluating learning algorithms</article-title>
          .
          <source>IEEE Transactions on knowledge and Data Engineering</source>
          <volume>17</volume>
          (
          <issue>3</issue>
          ),
          <volume>299</volume>
          {
          <fpage>310</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Kincaid</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fishburne</surname>
            <given-names>Jr</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.P.</given-names>
            ,
            <surname>Rogers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.L.</given-names>
            ,
            <surname>Chissom</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.S.:</surname>
          </string-name>
          <article-title>Derivation of new readability formulas (automated readability index, fog count and esch reading ease formula) for navy enlisted personnel</article-title>
          .
          <source>Tech. rep., Naval Technical Training Command Millington TN Research Branch</source>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lucisano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piemontese</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          :
          <article-title>Gulpease: una formula per la predizione della di colta dei testi in lingua italiana</article-title>
          .
          <source>Scuola e citta</source>
          <volume>3</volume>
          (
          <issue>31</issue>
          ),
          <volume>110</volume>
          {
          <fpage>124</fpage>
          (
          <year>1988</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Marschark</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spencer</surname>
            ,
            <given-names>P.E.</given-names>
          </string-name>
          :
          <article-title>The Oxford handbook of deaf studies, language, and education</article-title>
          , vol.
          <volume>2</volume>
          . Oxford University Press (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. OECD:
          <article-title>Inchiesta sulle competenze degli adulti primi risultati (</article-title>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Paul</surname>
          </string-name>
          , P.V.:
          <article-title>Language and deafness</article-title>
          . Jones &amp; Bartlett
          <string-name>
            <surname>Learning</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Shardlow</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A survey of automated text simpli cation</article-title>
          .
          <source>International Journal of Advanced Computer Science and Applications</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ),
          <volume>58</volume>
          {
          <fpage>70</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>21. www.commoncrawl.org:</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>