<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>How “deep” is learning word inflection?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pisa (Italy) francoalberto.cardillo@ilc.cnr.it</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pisa (Italy) claudia.marzi@ilc.cnr.it</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Marcello Ferro Istituto di Linguistica Computazionale ILC-CNR</institution>
          ,
          <addr-line>Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vito Pirrelli Istituto di Linguistica Computazionale ILC-CNR</institution>
          ,
          <addr-line>Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Machine learning offers two basic strategies for morphology induction: lexical segmentation and surface word relation. The first one assumes that words can be segmented into morphemes. Inducing a novel inflected form requires identification of morphemic constituents and a strategy for their recombination. The second approach dispenses with segmentation: lexical representations form part of a network of associatively related inflected forms. Production of a novel form consists in filling in one empty node in the network. Here, we present the results of a recurrent LSTM network that learns to fill in paradigm cells of incomplete verb paradigms. Although the process is not based on morpheme segmentation, the model shows sensitivity to stem selection and stem-ending boundaries.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. La letteratura offre due strategie
di base per l’induzione morfologica. La
prima presuppone la segmentazione delle
forme lessicali in morfemi e genera parole
nuove ricombinando morfemi conosciuti;
la seconda si basa sulle relazioni di una
forma con le altre forme del suo
paradigma, e genera una parola sconosciuta
riempiendo una cella vuota del paradigma. In
questo articolo, presentiamo i risultati di
una rete LSTM ricorrente, capace di
imparare a generare nuove forme verbali a
partire da forme giï¿oe note non segmentate.
Ciononostante, la rete acquisisce una
conoscenza implicita del tema verbale e del
confine con la terminazione flessionale.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>Morphological induction can be defined as the task
of singling out morphological formatives from
fully inflected word forms. These formatives are
understood to be part of the morphological
lexicon, where they are accessed and retrieved, to be
recombined and spelled out in word production.
The view requires that a word form be segmented
into meaningful morphemes, each contributing a
separable piece of morpho-lexical content.
Typically, this holds for regularly inflected forms, as
with Italian cred-ut-o ’believed’ (past participle,
from CREDERE), where cred- conveys the
lexical meaning, and -ut-o is associated with
morphosyntactic features. A further assumption is that
there always exists an underlying base form upon
which all other forms are spelled out. In an
irregular verb form like Italian appes-o ’hung’ (from
APPENDERE), however, it soon becomes difficult
to separate morpholexical information (the verb
stem) from morpho-syntactic information.</p>
      <p>
        A different formulation of the same task
assumes that the lexicon consists of fully-inflected
word forms and that morphology induction is
the result of finding out implicative relations
between them. Unknown forms are generated by
redundant analogy-based patterns between known
forms, along the lines of an analogical
proportion such as: rendere ‘make’ :: reso ‘made’ =
appendere ‘hang’ :: appeso ‘hung’. Support
to this view comes from developmental
psychology, where words are understood as the
foundational elements of language acquisition, from
which early grammar rules emerge
epiphenomally
        <xref ref-type="bibr" rid="ref29">(Tomasello, 2000; Goldberg, 2003)</xref>
        . After all,
children appear to be extremely sensitive to
subregularities holding between inflectionally-related
forms
        <xref ref-type="bibr" rid="ref11 ref12 ref17 ref24 ref9">(Bittner et al., 2003; Colombo et al., 2004;
Da˛browska, 2004; Orsolini and Marslen-Wilson,
1997; Orsolini et al., 1998)</xref>
        . Further support is
lent by neurobiologically inspired computer
models of language, blurring the traditional dichotomy
between processing and storage
        <xref ref-type="bibr" rid="ref15 ref22">(Elman, 2009;
Marzi et al., 2016)</xref>
        . In particular we will
consider here the consequences of this view on issues
of word inflection by recurrent Long Short Term
Memory (LSTM) networks (Malouf, in press).
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>The cell-filling problem</title>
      <p>
        To understand how word inflection can be
conceptualised as a word relation task, it is useful to
think of this task as a cell-filling problem
        <xref ref-type="bibr" rid="ref1 ref2">(Blevins
et al., 2017; Ackerman and Malouf, 2013;
Ackerman et al., 2009)</xref>
        . Inflected forms are
traditionally arranged in so-called paradigms. The full
paradigm of CREDERE ’believe’ is a labelled set of
all its inflected forms: credere, credendo, creduto,
credo etc. In most cases, these forms take one and
only one cell, defined as a specific combination
of tense, mood, person and number features: e.g.
crede, PRES IND, 3S. In all languages, words
happen to follow a Zipfian distribution, with very few
high-frequency words, and a vast majority of
exceedingly rare words (Blevins et al., 2017). As a
result, even high-frequency paradigms happen to
be attested partially, and learners must then be able
to generalise incomplete paradigmatic knowledge.
This is the cell-filling problem: given a set of
attested forms in a paradigm, the learner has to guess
other missing forms in the same paradigm.
      </p>
      <p>The task can be simulated by training a
learning model on a number of partial paradigms, to
then complete them by generating missing forms.
Training consists of &lt;lemma_paradigm cell,
inflected form&gt; pairs. A lemma is not a form (e.g.
credere), but a symbolic proxy of its lexical
content (e.g. CREDERE). Word inflection consists of
producing a fully inflected form given a known
lemma and an empty paradigm cell.
2.1</p>
      <p>Methods and materials
Following Malouf (in press), our LSTM
network (Figure 1) is designed to take as input
a lemma (e.g. CREDERE), a set of
morphosyntactic features (e.g. PRES_IND, 3, S) and
a sequence of symbols (&lt;crede&gt;)1 one symbol
st at a time, to output a probability distribution
1‘&lt;’ and ‘&gt;’ are respectively the start-of-word and the
end-of-word symbols
lexeme
(50)</p>
      <p>…
symbol (t)
(33)
tense
mood
(5)
person</p>
      <p>(4)
number
(3)
z(t)</p>
      <p>
        LSTM(t)
1:n
over the upcoming symbol st+1 in the sequence:
p(st+1jst;CREDERE, PRES_IND, 3, S). To
produce the form &lt;crede&gt;, we take the start symbol
‘&lt;’ as s1, use s1 to predict s2, then use the
predicted symbol to predict s3 and so on, until ‘&gt;’ is
predicted. Input symbols are encoded as mutually
orthogonal one-hot vectors with as many
dimensions as the overall number of different symbols
used to encode all inflected forms. The
morphosyntactic features of tense, person and number are
given different one-hot vectors, whose dimensions
equal the number of different values each
feature can take.2 All input vectors are encoded by
trainable dense matrices whose outputs are
concatenated into the projection layer z(t), which is
in turn input to a layer of LSTM blocks (Figure
1). The layer takes as input both the
information of z(t), and its own output at t–1.
Recurrent LSTM blocks are known to be able to capture
long-distance relations in time series of symbols
        <xref ref-type="bibr" rid="ref17 ref19">(Bengio et al., 1994; Hochreiter and Schmidhuber,
1997; Jozefowicz et al., 2015)</xref>
        , avoiding classical
problems with training gradients of Simple
Recurrent Networks
        <xref ref-type="bibr" rid="ref13 ref18">(Jordan, 1986; Elman, 1990)</xref>
        .
      </p>
      <p>
        We tested our model on two comparable sets
of Italian and German inflected verb forms
(Table 1), where paradigms are selected by sampling
the highest-frequency fifty paradigms in two
reference corpora
        <xref ref-type="bibr" rid="ref20 ref5">(Baayen et al., 1995; Lyding et al.,
2014)</xref>
        . For both languages, a fixed set of cells was
2Note that an extra dimension is added when a feature can
be left uninstatiated in particular forms, as is the case with
person and number features in the infinitive.
language
chosen from each paradigm: all present indicative
forms (n=6), all past tense forms (n=6), infinitive
(n=1), past participle (n=1), German present
participle/Italian gerund (n=1).3 The two sets are
inflectionally complex: they exhibit extensive stem
allomorphy and a rich set of affixations,
including circumfixation (German ge-mach-t ’made’,
past participle). Most importantly, the
distribution of stem allomorphs is entirely accountable in
terms of equivalence classes of cells, forming
morphologically heterogenous, phonologically poorly
predictable, but fairly stable sub-paradigms
        <xref ref-type="bibr" rid="ref25">(Pirrelli, 2000)</xref>
        . Selection of the contextually
appropriate stem allomorph for a given cell thus requires
knowledge of the form of the allomorph and of its
distribution within the paradigm.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Results and discussion</title>
      <p>To meaningfully assess the relative
computational difficulty of the cell-filling task, we
calculated a simple baseline performance, with 695
forms of our original datasets selected for
training, and 55 for testing.4For this purpose, we used
the baseline system for Task 1 of the
CoNLLSIGMORPHON-2017 Universal Morphological
Reinflection shared task.5The model changes the
infinitive into its inflected forms through rewrite
rules of increasing specificity: e.g. two Italian
forms such as badare ‘to look after’ and bado ‘I
look after’ stand in a BASE :: PRES_IND_3S
relation. The most general rule changing the
former into the latter is -are -&gt; -o, but more specific
rewrite rules can be extracted from the same pair:
3The full data set is available at http://www.
comphyslab.it/redirect/?id=clic2017_data.
Each training form is administered once per epoch, and the
number of epochs is a function of a “patience” threshold.
Although a uniform distribution is admittedly not realistic,
it increases the entropy of the cell-filling problem, to define
some sort of upper bound on the complexity of the task.</p>
      <p>4Test forms were selected to constitute a benchmark for
evaluation. We made it sure that a representative sample of
German and Italian irregulars were included for evaluation,
provided that they could be generalised on the basis of the
training data available.</p>
      <p>5https://github.com/sigmorphon/
conll2017(written by Mans Hulden).</p>
      <p>German test
CoNLL baseline
128-blocks
256-blocks
512-blocks
Italian test
baseline
128-blocks
256-blocks
512-blocks
all
-dare -&gt; -do, -adare -&gt; -ado, -badare -&gt; -bado.
The algorithm then generates the PRES_IND_3S of
- say - diradare ’thin out’, by using the rewrite rule
with the longest left-hand side matching diradare
(namely -adare &gt; -ado). If there is no matching
rule, the base is used as a default output.</p>
      <p>The algorithm proves to be effective for
regular forms in both languages (Table 2). However,
per-word accuracy drops dramatically on German
irregulars (0.23), and Italian irregulars (0.5). The
same table shows accuracy scores on test data
obtained by running 128, 256 and 512 LSTM blocks.
Each model instance was run 10 times, and overall
per-word scores are averaged across repetitions.6</p>
      <p>
        The CoNLL baseline is reminiscent of
Albright and Hayes’ (2003) Minimal Generalization
Learner, inferring Italian infinitives from first
singular present indicative forms
        <xref ref-type="bibr" rid="ref4">(Albright, 2002)</xref>
        . In
the present case, however, the inference goes from
the infinitive (base) to other paradigm cells. The
inference is much weaker in German, where stem
allomorphy is more consistently distributed within
each paradigm. In Appendix, Table 3 contains
a list of all German forms wrongly produced by
the CoNLL baseline, together with per-word
accuracy of our models. Most wrong forms are
inflected forms requiring ablaut, which turn out to
be over-regularised by the CoNLL baseline (e.g.
*stehtet for standet, *beginntet for begannt). It
appears that, in German, a purely syntagmatic
approach to word production, deriving all inflected
forms from an underlying base, has a strong bias
towards over-regularisation. Simply put, the
orthotactic/phonotactic structure of the German stem
6The per-word score is 1 (correct), or 0 (wrong).
      </p>
      <p>Italian training (256−cell LSTM)</p>
      <p>Italian training (512−cell LSTM)
1
is less criterial for stem allomorphy than the
Italian one. LSTMs are considerably more robust in
this respect. Memory resources allowing, they
can keep track of local syntagmatic constraints
as well as more global, paradigmatic constraints,
whereby all paradigmatically-related forms
contribute to fill in gaps in the same paradigm. For
example, knowledge that a paradigm contains a few
stem allomorphs is good reason for an LSTM to
produce a stem allomorph in other (empty) cells.
The more systematic the distribution of stem
alternants is across the paradigm, the easier for the
learner to fill in empty cells. German conjugation
proves to be paradigmatically well-behaved.</p>
      <p>An LSTM recurrent network has no information
about the morphological structure of input forms.
Due to the predictive nature of the production task
and the LSTM re-entrant layer, however, the
network develops a left-to-right sensitivity to
upcoming symbols, with per-symbol accuracy being a
function of the network confidence about the next
output symbol. To assess the correlation between
per-symbol accuracy and “perception” of the
morphological structure, we used a Linear Mixed
Effects (LME) model of how well structural features
of German and Italian verb forms interpolate the
“average” network accuracy in producing an
upcoming symbol (1 for a hit, 0 for a miss) in both
training and test. The marginal plots of Figure 2
show that there is a clear structural effect of the
distance to the stem-ending boundary of the
symbol currently being produced, over and above the
length of the input string. Besides, stems and
suffixes of regulars exhibit different accuracy slopes
compared with stems and suffixes of irregulars.
Intuitively, production of an inflected form by a
LSTM network is fairly easy at the beginning of
the stem, but it soon gets more difficult when
approaching the morpheme boundary, particularly
with irregulars. Accuracy reaches the minimum
value on the first symbol of the inflectional
ending, which marks a point of structural
discontinuity in an inflected verb form. From that position,
accuracy starts increasing again, showing a
characteristically V-shaped trend. Clearly, this trend
is more apparent with test words (Figure 2,
bottom), where stems and endings are recombined in
novel ways. The same results hold for German.
On the other hand, no evidence of structure
sensitivity was found in a LME model of the baseline
output for both German and Italian.</p>
      <p>
        The cell-filling problem is an ecological,
developmentally motivated task, based on evidence of
fully inflected forms. Although other (simpler)
models have been proposed to account for
formmeaning mapping in Morphology
        <xref ref-type="bibr" rid="ref26 ref6">(Baayen et al.,
2011; Plaut and Gonnerman, 2000, among
others)</xref>
        , we do not know of any other artificial
neural networks that can simulate word inflection as
a cell-filling task. Unlike more traditional
connectionist architectures
        <xref ref-type="bibr" rid="ref27">(Rumelhart and
McClelland, 1986)</xref>
        , recurrent LSTMs do not presuppose
the existence of underlying base forms, but they
learn possibly alternating stems upon exposure
to full forms. Admittedly, the use of
orthogonal one-hot vectors for lemmas, unigram temporal
series for inflected forms, and abstract
morphosyntactic features as a proxy of context-sensitive
functional agreement effects, are crude
representational short-hands. Nonetheless, in tackling
the task, LSTMs prove to be able to orchestrate
“deep” knowledge about word structure, well
beyond pure surface word relations: namely
stemaffix boundaries, paradigm organisation and
degrees of regularity in stem formation. Acquisition
of different inflectional systems may require a
different balance of all these pieces of knowledge.
      </p>
      <p>Adele E Goldberg. 2003. Constructions: a new
theoretical approach to language. Trends in cognitive
sciences, 7(5):219–224.</p>
      <p>Robert Malouf. in press. Generating morphological
paradigms with a recurrent neural network.
Morphology.</p>
      <p>Appendix A. Comparative test results
base form</p>
      <p>target
bleiben
dï¿oerfen</p>
      <p>sein
mï¿oessen
bestehen
sprechen
geben
sehen</p>
      <p>tun
stehen
fahren
finden
dï¿oerfen</p>
      <p>fahren
beginnen
kommen
liegen
sehen
bringen
fragen
gehen
haben
nehmen
nennen
sagen
tragen
bitten
denken
geben
scheinen</p>
      <p>setzen
sprechen
werden
bliebt
gedurft
seiend
gemusst
bestandet
spricht
gibt
siehst
tatet
standet
fï¿oehrst
fandet
darf
fuhrst
begannt
kamst
lagt
saht
brachtet
fragtet
gingt
hattet
nahmt
nanntet
sagtet
trï¿oegst</p>
      <p>baten
dachtest</p>
      <p>gabst
schienst
setztet
sprachst
wurdet</p>
      <p>CoNNL
baseline</p>
      <p>blieb
gedï¿oerfen</p>
      <p>seind
gemï¿oessen
bestehtet
sprecht
gebt
sehst
tut
stehtet
fahrst
findet
dï¿oerfe
fahrtest
beginntet
kommst
liechtet
sehtet
brinchtet
frugt
gehtet
habt
neht
nenntet
sugt
tragst
bitten
denkest</p>
      <p>gebst
scheintest</p>
      <p>setzet
sprechtest
werdet</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Farrell</given-names>
            <surname>Ackerman</surname>
          </string-name>
          and
          <string-name>
            <given-names>Robert</given-names>
            <surname>Malouf</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Morphological organization: The low conditional entropy conjecture</article-title>
          .
          <source>Language</source>
          ,
          <volume>89</volume>
          (
          <issue>3</issue>
          ):
          <fpage>429</fpage>
          -
          <lpage>464</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Farrell</given-names>
            <surname>Ackerman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>James P.</given-names>
            <surname>Blevins</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Robert</given-names>
            <surname>Malouf</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Parts and wholes: Patterns of relatedness in complex morphological systems and why they matter</article-title>
          . In James P. Blevins and Juliette Blevins, editors,
          <source>Analogy in Grammar: Form and Acquisition</source>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Adam</given-names>
            <surname>Albright</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bruce</given-names>
            <surname>Hayes</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Rules vs. analogy in english past tenses: A computational/experimental study</article-title>
          .
          <source>Cognition</source>
          ,
          <volume>90</volume>
          (
          <issue>2</issue>
          ):
          <fpage>119</fpage>
          -
          <lpage>161</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Adam</given-names>
            <surname>Albright</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Islands of reliability for regular morphology: Evidence from italian</article-title>
          .
          <source>Language</source>
          , pages
          <fpage>684</fpage>
          -
          <lpage>709</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Harald R. Baayen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Piepenbrock</surname>
            , and
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Gulikers</surname>
          </string-name>
          ,
          <year>1995</year>
          .
          <article-title>The CELEX Lexical Database (CD-ROM)</article-title>
          .
          <source>Linguistic Data Consortium</source>
          , University of Pennsylvania, Philadelphia, PA.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Harald R. Baayen</surname>
          </string-name>
          , Petar Milin, Dusica Filipovic´ Ðurd¯evic´,
          <string-name>
            <surname>Peter Hendrix</surname>
            , and
            <given-names>Marco</given-names>
          </string-name>
          <string-name>
            <surname>Marelli</surname>
          </string-name>
          .
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>An amorphous model for morphological processing in visual comprehension based on naive discriminative learning</article-title>
          .
          <source>Psychological review</source>
          ,
          <volume>118</volume>
          (
          <issue>3</issue>
          ):
          <fpage>438</fpage>
          -
          <lpage>481</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          1994.
          <article-title>Learning long-term dependencies with gradient descent is difficult</article-title>
          .
          <source>IEEE transactions on neural networks</source>
          ,
          <volume>5</volume>
          (
          <issue>2</issue>
          ):
          <fpage>157</fpage>
          -
          <lpage>166</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Dagmar</given-names>
            <surname>Bittner</surname>
          </string-name>
          , Wolfgang U. Dressler, and
          <string-name>
            <surname>Marianne</surname>
          </string-name>
          Kilani-Schoch, editors.
          <year>2003</year>
          .
          <article-title>Development of Verb Inflection in First Language Acquisition: a crosslinguistic perspective</article-title>
          . Mouton de Gruyter, Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <year>2017</year>
          .
          <article-title>The zipfian paradigm cell filling problem</article-title>
          .
          <source>In Ferenc Kiefer</source>
          ,
          <string-name>
            <given-names>James P.</given-names>
            <surname>Blevins</surname>
          </string-name>
          , and Huba Bartos, editors,
          <source>Morphological Paradigms and Functions</source>
          . Brill, Leiden.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Colombo</surname>
          </string-name>
          , Alessandro Laudanna, Maria De Martino, and
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Brivio</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Regularity and/or consistency in the production of the past participle? Brain and language</article-title>
          ,
          <volume>90</volume>
          (
          <issue>1</issue>
          ):
          <fpage>128</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Ewa</given-names>
            <surname>Da</surname>
          </string-name>
          ˛browska.
          <year>2004</year>
          .
          <article-title>Rules or schemas? evidence from polish</article-title>
          .
          <source>Language and cognitive processes</source>
          ,
          <volume>19</volume>
          (
          <issue>2</issue>
          ):
          <fpage>225</fpage>
          -
          <lpage>271</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey L Elman</surname>
          </string-name>
          .
          <year>1990</year>
          .
          <article-title>Finding structure in time.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Cognitive</given-names>
            <surname>Science</surname>
          </string-name>
          ,
          <volume>14</volume>
          (
          <issue>2</issue>
          ):
          <fpage>179</fpage>
          -
          <lpage>211</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey L Elman</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>On the meaning of words and dinosaur bones: Lexical knowledge without a lexicon.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Cognitive</surname>
            <given-names>science</given-names>
          </string-name>
          ,
          <volume>33</volume>
          (
          <issue>4</issue>
          ):
          <fpage>547</fpage>
          -
          <lpage>582</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jürgen</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation</source>
          ,
          <volume>9</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Jordan</surname>
          </string-name>
          .
          <year>1986</year>
          .
          <article-title>Serial order: A parallel distributed processing approach</article-title>
          .
          <source>Technical Report 8604</source>
          , University of California.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Rafal</given-names>
            <surname>Jozefowicz</surname>
          </string-name>
          , Wojciech Zaremba, and
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>An empirical exploration of recurrent network architectures</article-title>
          .
          <source>In Proceedings of the 32nd International Conference on Machine Learning (ICML15)</source>
          , pages
          <fpage>2342</fpage>
          -
          <lpage>2350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Verena</given-names>
            <surname>Lyding</surname>
          </string-name>
          , Egon Stemle, Claudia Borghetti, Marco Brunello, Sara Castagnoli, Felice Dell' Orletta, Henrik Dittmann, Alessandro Lenci, and
          <string-name>
            <given-names>Vito</given-names>
            <surname>Pirrelli</surname>
          </string-name>
          .
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <article-title>The paisá corpus of italian web texts</article-title>
          .
          <source>Proceedings of the 9th Web as Corpus Workshop (WaC-9)@ EACL</source>
          <year>2014</year>
          , pages
          <fpage>36</fpage>
          -
          <lpage>43</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Claudia</given-names>
            <surname>Marzi</surname>
          </string-name>
          , Marcello Ferro, Franco Alberto Cardillo, and
          <string-name>
            <given-names>Vito</given-names>
            <surname>Pirrelli</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Effects of frequency and regularity in an integrative model of word storage and processing</article-title>
          .
          <source>Italian Journal of Linguistics</source>
          ,
          <volume>28</volume>
          (
          <issue>1</issue>
          ):
          <fpage>79</fpage>
          -
          <lpage>114</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          1997.
          <article-title>Universals in morphological representation: Evidence from italian</article-title>
          .
          <source>Language and Cognitive Processes</source>
          ,
          <volume>12</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Margherita</given-names>
            <surname>Orsolini</surname>
          </string-name>
          , Rachele Fanari, and
          <string-name>
            <given-names>Hugo</given-names>
            <surname>Bowles</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Acquiring regular and irregular inflection in a language with verb classes</article-title>
          .
          <source>Language and cognitive processes</source>
          ,
          <volume>13</volume>
          (
          <issue>4</issue>
          ):
          <fpage>425</fpage>
          -
          <lpage>464</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Vito</given-names>
            <surname>Pirrelli</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Paradigmi in morfologia.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>David C Plaut and Laura M Gonnerman</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Are non-semantic morphological effects incompatible with a distributed connectionist approach to lexical processing? Language</article-title>
          and
          <string-name>
            <given-names>Cognitive</given-names>
            <surname>Processes</surname>
          </string-name>
          ,
          <volume>15</volume>
          (
          <issue>4</issue>
          /5):
          <fpage>445</fpage>
          -
          <lpage>485</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>David E.</given-names>
            <surname>Rumelhart and James L. McClelland</surname>
          </string-name>
          .
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <article-title>On learning the past tenses of english verbs</article-title>
          . In David E. Rumelhart, James L.
          <article-title>McClelland, and</article-title>
          the PDP Research Group, editors,
          <source>Parallel Distributed Processing. Explorations in the Microstructures of Cognition</source>
          , volume
          <volume>2</volume>
          Psychological and Biological Models, pages
          <fpage>216</fpage>
          -
          <lpage>271</lpage>
          . MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Tomasello</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>The item-based nature of children's early syntactic development</article-title>
          .
          <source>Trends in cognitive sciences, 4</source>
          (
          <issue>4</issue>
          ):
          <fpage>156</fpage>
          -
          <lpage>163</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>