<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mapping the Constructicon with SYMPAThy. Italian Word Combinations between fixedness and productivity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro Lenci Sara Castagnoli</string-name>
          <email>alessandro.lenci@ling.unipi.it s.castagnoli@unibo.it gianluca.lebani@for.unipi.it francesca.masini@unibo.it marco.senaldi@sns.it</email>
          <email>s.castagnoli@unibo.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Malvina Nissim</string-name>
          <email>m.nissim@rug.nl</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Copyright c by the paper's authors. Copying permitted for private and academic purposes.</string-name>
          <email>gianluca.lebani@for.unipi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Gianluca E. Lebani</institution>
          ,
          <addr-line>Marco S. G. Senaldi Francesca Masini</addr-line>
          ,
          <institution>University of Pisa University of Bologna</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>In Vito Pirrelli, Claudia Marzi, Marcello Ferro (eds.): Word Structure and Word Usage. Proceedings of the NetWordS Final</institution>
          ,
          <addr-line>Conference, Pisa, March 30-April 1, 2015, published at http://ceur-ws.org</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Groningen</institution>
        </aff>
      </contrib-group>
      <fpage>144</fpage>
      <lpage>149</lpage>
      <abstract>
        <p>This work introduces SYMPAThy, a data representation model in which the combinatorial properties of a lexical item are described by merging surface and deeper linguistic information. The proposed approach is then evaluated by comparing, for a sample list of verbal idioms, a set of SYMPAThy-based fixedness indexes against the relevant speaker-elicited indexes available in the descriptive norms collected by Tabossi et al. (2011).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Word combinatorics and constructions</title>
      <p>
        By “Word Combinations” (WoCs) we broadly
refer to the range of constructions typically
associated with a lexical item. In Construction
Grammar, constructions (Cxn) are
conventionalized form-meaning pairings that can vary in
both complexity and schematicity
        <xref ref-type="bibr" rid="ref10 ref6 ref8">(Fillmore et al.,
1988; Goldberg, 2006; Hoffmann and Trousdale,
2013)</xref>
        . The Constructicon spans from fully
specified structures (kick the bucket ) to complex,
productive abstract structures such as argument
patterns (e.g., the Ditransitive Cxn “Subj V Obj1
Obj2”, she baked him a cake), passing through
“intermediate” Cxns with different degrees of
schematicity, complexity and productivity (e.g.,
take Obj for granted), in what is known as the
lexicon-syntax continuum. WoCs thus comprise
so-called Multiword Expressions (MWEs), i.e. a
variety of recurrent expressions acting as a
single unit at some level of linguistic analysis, like
phrasal lexemes, idioms, collocations
        <xref ref-type="bibr" rid="ref13 ref13 ref3 ref3 ref9">(Calzolari et
al., 2002; Sag et al., 2002; Gries, 2008)</xref>
        , as well as
the preferred distributional properties of a word at
a more abstract level, i.e. argument structures and
selectional preferences
        <xref ref-type="bibr" rid="ref7">(Goldberg, 1995)</xref>
        .
      </p>
      <p>
        Each lexeme can thus be described as having a
combinatory potential to be defined and observed
at a more constrained, surface POS-pattern level
(P-level) and at the more abstract level of
syntactic structure (S-level). These two levels are often
kept separate, not only theoretically, but also
computationally, as their performance varies according
to the different types of combinations that we want
to track
        <xref ref-type="bibr" rid="ref13 ref3 ref5">(Sag et al., 2002; Evert and Krenn, 2005)</xref>
        .
      </p>
      <p>We advocate a unified and integrated view of a
lexeme’s combinatory potential, in order to
capture both fixed combinations (MWEs of various
types) and more productive aspects of the lexeme’s
distributional behaviour. The theoretical premises
lie in the constructionist view of the mental
lexicon outlined above, whereas a proposal for a
computational implementation is illustrated here.
Specifically, we i) present SYMPAThy, a model
of data representation that takes into account both
surface and deeper linguistic information; ii)
develop and test an index of productivity for Italian
WoCs based on SYMPAThy.
2</p>
    </sec>
    <sec id="sec-2">
      <title>SYMPAThy: a joint approach to WoCs</title>
      <p>We argue that to obtain a comprehensive picture of
the combinatory potential of a word and enhance
extracting efficacy for WoCs, the P-based
approach (which exploits sequences of POS-patterns
and association measures) and the S-based
approach (which exploits syntactic dependencies and
association measures) should be combined. We
illustrate this point with an example based on the
Target Lexeme (TL) gettare ‘throw’ (V).1</p>
      <p>We want to use S-based methods to capture the
fact that V occurs typically within some
syntactic Frames and not others, that for each Frame
we have typical Fillers (lexical items) instantiating
Frame slots, and that each slot is associated with
certain semantic (ontological) classes:2</p>
      <p>
        1All data is from a version of the “la Repubblica” corpus
        <xref ref-type="bibr" rid="ref2">(Baroni et al., 2004)</xref>
        POS tagged with the Part-Of-Speech
tagger described in Dell’Orletta (2009) and dependency parsed
with DeSR
        <xref ref-type="bibr" rid="ref1 ref4">(Attardi and Dell’Orletta, 2009)</xref>
        .
      </p>
      <p>
        2Data extracted by LexIt
        <xref ref-type="bibr" rid="ref11">(Lenci, 2014)</xref>
        . The list is partial:
only the first three Frames are included; Frames with the
re• subj#obj#comp-in
– OBJ Filler: {scompiglio, sasso, corpo, fumo,
cadavere, ...}; {Natural Object, Substance, ...}
– COMP-in Filler: {panico, caos, sconforto, mare,
stagno, cestino, ...}; {Feeling, State, ...}
• subj#obj
– OBJ Filler: {spugna, base, ombra, acqua, luce,
ponte, ...}; {Substance, Artifact, ...}
At this point, we observe that all these words are
typically associated with our TL, but we don’t
know in which way they are all linked to one
another. For instance, we have no elements
for thinking that subj#gettare#acqua#su fuoco is
any different from subj#gettare#acqua#su tavolo
or subj#gettare#ombra#su istituzione. However,
while gettare acqua sul fuoco ‘defuse’ is an
idiom in Italian, gettare acqua sul tavolo only has
a literal meaning (‘throw water on the table’);
subj#gettare#fango#su istituzione is yet different,
since gettare fango su ‘defame’ is a fixed
expression, but the Filler istituzione ‘institution’ is just
one of many possibilities, so the expression is
partially fixed, resulting in something like [ gettare
fango su PERSON/INSTITUTION]. The
significance of gettare acqua sul fuoco with respect
to gettare acqua sul tavolo emerges much more
clearly if we use a P-based method. Extracting
surface material, the former expression will be
ranked higher than the latter (given the pattern “V
N PREPART N”) as the association between all
words is stronger.
      </p>
      <p>So, fine-grained differences do not emerge with
the S-method, while the P-based method fails to
capture the higher-level generalizations we get
with the S-method. In order to get the best of both
worlds, we extracted corpus data into
SYMPAThy (SYntactically Marked PATterns), a database
where information on both levels is stored and
accessible jointly:
• syntactic frames with argument slots and fillers;
• linear order of all elements for each TL;
• POS tag for each element (simple preposition
vs. preposition with article, definite vs.
indefinite article, modal vs. full verb, etc.);
flexive form gettarsi ’throw oneself’ and objectless forms are
excluded.
• morphosyntactic features:
finiteness, tense, etc.
gender, number,
3</p>
    </sec>
    <sec id="sec-3">
      <title>WoC fixedness with SYMPAThy</title>
      <p>Since constructions span along a continuum
between fixedness and productivity, there have been
various attempts at measuring how fixed a given
WoC is, mostly based on surface features. Nissim
and Zaninello (2011) assess the fixedness of a
subset of complex nominals by comparing inflected
and lemmatized forms, and taking into account the
proportion of elements that undergo variation in a
given MWE. Inflection is also used by Squillante
(2014) on noun-adjective expressions, and is
combined with two other measures, interruptibility and
substitutability. Zeldes (2013) extends Baayen’s
morphological productivity approach to argument
structure and estimates the productivity of a
syntactic slot from the number of its hapax noun
fillers. Wulff (2009) uses a set of
morphosyntactic indexes of variations and a collocation-based
index of compositionality as variables in a
regression study to determine fixedness.</p>
      <p>We extend the state of the art of the quantitative
approach to construction fixedness by exploiting
the potentialities of SYMPAThy to develop a
series of corpus-based indexes able to describe the
fixedness of some idiomatic expressions. Our
approach is then evaluated by comparing, for a
sample list of expressions, a composition of our
indexes against the behavioral judgments of
syntactic flexibility collected by Tabossi et al. (2011).
3.1</p>
      <sec id="sec-3-1">
        <title>The combinatory behaviour of a TL</title>
        <p>In the SYMPAThy model, the combinatory space
of a Target Lexeme is assumed to be formed by a
network of Cxns, varying for their degree of
fixedness/productivity. For any given TL such a
representation is built by means of the following
fourstep procedure:
1. its SYMPAThy patterns are extracted from a
reference corpus;
2. the set of single and multiple slot Cxns that TL
combines with are semi-automatically
identified. An example for the verb gettare is
reported and explained in Appendix 1;
3. each construction is associated with a
variational profile formed by a number of statistics
extracted from the SYMPAThy pattern to
estimate: i) the variability of the fillers that
instantiate the syntactic slots of constructions; ii) the
morphological variability of the constructions’
components; iii) the variability with respect to
determiners; iv) the variability with respect to
adjectival and adverbial modifications; v) the
variability in the linear order.
4. variational profiles are then used to measure the
lexical, morphological and syntactic degrees of
freedom of Cxns, providing a multidimensional
quantitative characterization of their level of
fixedness.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Entropy-based Cxn fixedness modeling</title>
        <p>
          In what follows, we devise a way to encode the
variation possibilities shown by Cxns, as well as
a meaningful way to combine them. Specifically,
we distinguish a series of dimensions of variation
and propose to exploit Entropy
          <xref ref-type="bibr" rid="ref14">(Shannon, 1948)</xref>
          to measure how fixed is the behavior of a Cxn in a
given dimension.
        </p>
        <p>Entropy is a measure of randomness, calculated
as the average uncertainty of a single variable:
H(X) = −</p>
        <p>X p(x) log2(p(x))
x∈X
(1)
This measure of randomness can be adapted to our
needs by taking the variable X as being a Cxn of
interest, and the states of the system x as its values
on one dimension of variation. Lower entropy
values are to be understood as evidence of fixedness,
while higher values suggest a more variable
distribution of the states of a given variable, i.e. the
target construction tends to be freer.</p>
        <p>Observed entropy values, however, can span
from 0 to the logarithm of the number of values
that X can assume. As a consequence, entropy
values related to different dimensions of variation
are not comparable, and cannot be combined into
a single fixedness index. We overcome this
limitation by following Wulff (2008) and describing the
randomness of each variability dimension in terms
of relative entropy, computed as the ratio between
the observed entropy from eq.1 and the maximum
entropy Hmax for the variable X:</p>
        <p>Hrel(X) =</p>
        <p>H(X)
Hmax(X)
=</p>
        <p>H(X)
log2(|X|)
This measure, that ranges from 0 to 1, has been
employed as a flexibility measure to describe the
flexibility of a given set of target Cxns along the
following dimensions of variation:
LEXICAL VARIABILITY. The entropy of the
lexical instantiation of the slot positions of a Frame
is calculated by assuming that the states x of
the random variable X are all the possible fillers
that can instantiate a given slot in Cxn (e.g. in
subj#gettare#obj:luce#su X, X can be filled by
vicenda ‘matter’, mistero ‘mystery’, etc.).</p>
      </sec>
      <sec id="sec-3-3">
        <title>MORPHOLOGICAL VARIABILITY. It is cal</title>
        <p>culated as the entropy of the morphological
features manifested by the fillers of a Cxn
(e.g., gettare#ombra-fs ‘cast shadow-singular’;
gettare#ombra-fp ‘cast shadow-plural’).</p>
      </sec>
      <sec id="sec-3-4">
        <title>ARTICLES VARIABILITY. This index encodes</title>
        <p>how variable is the presence or absence of articles
determining the available slots in a Cxn, and, if
appropriate, their type (DEFinite vs. INDefinite):
for instance, gettare#∅+acqua#su DEF+fuoco.</p>
      </sec>
      <sec id="sec-3-5">
        <title>PRESENCE OF MODIFIERS. This index en</title>
        <p>codes how variable is the presence or
absence of adjectives, adverbs or prepositional
phrases modifying the available slots. In this
way, it is possible to account for patterns
like:gettare#molta+acqua#su ∅+fuoco.</p>
      </sec>
      <sec id="sec-3-6">
        <title>DISTANCE VARIABILITY. This index exploits</title>
        <p>information on linear order available in
SYMPAThy to estimate how variable is the distance in
tokens between a TL and the other constituents of a
given lexically specified Cxn.</p>
        <p>In the experiment reported in the next section,
we have combined the single variability measures
Hrel(X) into an overall flexibility index F (X)
corresponding to four possible combinations:
• SUM: F (X) is obtained by summing over all
the single Hrel(X) values;
• AVERAGE: F (X) is the mean of the single</p>
        <p>Hrel(X) values;
• AVERAGEP OS : F (X) is the mean of the
positive Hrel(X) values;
• MAX: F (X) is the highest Hrel(X) value.</p>
        <p>We leave to future research the investigation of
further ways to combine the variability indexes.
(2)
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>In order to evaluate our approach, we set out to test
if our indexes can mimic the intuitive judgments
of native speakers about the fixedness of fully
lexically specified constructions. To do so, we
selected a subset of the idioms in the norms collected
by Tabossi et al. (2011), and tested to what degree
the speaker-elicited flexibility judgments available
in this repository can be modeled by a composition
of our variability indexes.</p>
      <sec id="sec-4-1">
        <title>The descriptive norms by Tabossi et al.</title>
        <p>Tabossi et al. (2011) collected several normative
measures for 245 Italian verbal idiomatic
expressions. Using a group of 740 Italian speakers, they
collected a minimum of 40 elicited judgments for
each idiom on several psycholinguically relevant
variables.</p>
        <p>Among the different kinds of ratings, those
concerning syntactic flexibility have been collected
by inserting each idiomatic expression in a
sentence in which one of the following vfie syntactic
modifications occurred: adverb insertion,
adjective insertion, left dislocation, passive and
movement. Participants were asked to evaluate, on a
7-point scale, how much the meaning of the
idiomatic expression in the syntactically modified
sentence was similar to its unmarked meaning as
expressed in a paraphrase prepared by the authors.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Data extraction</title>
        <p>Out the 245 expressions in Tabossi et al., we
selected the 23 target idioms reported in Appendix 2.
Each such idiom can be represented, in our
approach, as a fully lexically specified transitive Cxn
headed by a given verbal TL, for which the subject
slot is underspecified (e.g. gettare#obj:maschera).
We built the variational profiles of our target
idioms by adopting an adapted version of the
procedure described in Section 3:
1. for each TL, we extracted the SYMPAThy
patterns from the “la Repubblica” corpus;
2. the patterns involving one of our target idioms
were identified and selected;
3. for each idiom, the variability indexes
described in Section 3.2 were calculated. Note
that, given the nature of our experimental
stimuli, the lexical variability index is not relevant;
4. we built a fixedness index for each idiom,
according to the four composition methods in the
previous section.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Results and discussion</title>
        <p>In order to test the cognitive plausibility of the
fixedness indexes extracted from SYMPAThy, we
calculated the Pearson’s Product-Moment
Correlation strength between them and the syntactic</p>
      </sec>
      <sec id="sec-4-4">
        <title>Combination</title>
        <p>SUM
AVERAGE
AVERAGE P OS
MAX
r
flexibility ratings in Tabossi et al. (2011).
Correlation values are reported in Table 1. In all cases,
there is a significant ( p &lt; .05) positive correlation,
ranging between .44 and .47, thus supporting the
psycholinguistic plausibility of our corpus-based
variability indexes.</p>
        <p>These results, albeit preliminary, look
promising especially given the different nature of the
behavioral and corpus-based indexes. On the
one hand, the speakers’ ratings are semantically
driven, since they are thought to model how much
the figurative meaning of a given idiom is sensitive
to its syntactic form. On the other hand, the
automatically corpus-derived information exploited by
our indexes does not take meaning into account.
SUch indexes describe a lexically specified Cxn
that can in principle have an idiomatic as well as
a compositional, literal meaning (even if,
presumably, the latter case is rare in the corpus).
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this study we presented a procedure for
characterizing the combinatorial potential of a lexical
item and the degree of fixedness of the Cxns it
occurs in. Such a procedure has been preliminary
tested on a small sample of idiomatic expressions
and the resulting representation has been evaluated
against the subject-elicited judgments collected by
Tabossi et al. (2011). In the future, we are
planning to extend the inventory of variability
dimensions (addressing also the question of the semantic
compositionality of Cxns), to study their relative
weight and their interactions, and to develop more
sophisticated ways to combine them.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>
        This research was carried out within the CombiNet
project
        <xref ref-type="bibr" rid="ref12">(PRIN 2010-2011 Word Combinations in
Italian: theoretical and descriptive analysis,
computational models, lexicographic layout and
creation of a dictionary, n. 20105B3HE8)</xref>
        funded by
the Italian Ministry of Education, University and
Research (MIUR).
[CAUSE (OBJ, [GO (OBJ, [TO ([ON (COMP)])])])]
      </p>
      <p>ilnks )
ofrm
emanig
#acqu
#sul #Fuoc
[[SUBJ]NP gettare (ADV) (ADJ) acqua sul fuoco]</p>
      <p>Person, Event,...</p>
      <p>SUBJ:
OBJ:
SU:
COMP: fuoco
acqua
sul
‘defuse, minimize a situation’
raget
‘add fuel to the fire’</p>
      <p>Person, Event,...</p>
      <p>benzina
[[SUBJ]NP gettare (ADV) (ADJ) benzina sul fuoco]
lt = GETTARE ‘tHroW’
ofrm
emanig
ofrm
emanig
ofrm
emanig
ofrm
[CAUSE (OBJ, [GO (AWAY)])]</p>
      <p>Frame
... ... ...
... ... ...</p>
      <p>... ... ...</p>
      <p>... ... ...
[[SUBJ]NP gettare (ADV) (ADJ) fango su [COMP]NP]
SUBJ:
OBJ:</p>
      <p>Person, Event,...</p>
      <p>fango (⇒ SG; bare | partitive)
COMP: Person, Institution, ...
‘defame, discredit, blacken the name of’
ofrm
ofrm
emanig
... Questo getta una pesantissima ombra sulla legittimità ... ... rischia di gettare ulteriore fango sul calcio ...
‘This casts a serious shadow on the legitimacy...’
... Il rivale getta ombra sulla salute del leader ...
‘His opponent casts a shadow on the leader’s health’
‘(it) may sully football even more’
... Hanno sempre gettato fango su di noi ...</p>
      <p>‘They have always sullied us’
... Gli amici hanno gettato sulla bara garofani rossi ... ‘Friends threw red carnations on his coffin’
... getta un sasso sull’autostrada ... ‘(s/he) throws a stone in the highway’
The verb gettare ‘to throw’ combines with the highly schematic subj#obj#comp-su Cxn, whose slots
can freely vary with respect to linear order, presence of determiners, modifiers, etc. A semi-productive
instance of this construction is the subj#obj:ombra#comp-su Cxn, with a fixed object slot and a partially
variable oblique slot, which can appear with a semantically limited range of arguments. A fully lexically
specified instance of the same construction is instead the subj#obj: acqua#comp-su:sul-fuoco Cxn, which
has both slots instantiated and limited degree of variability.</p>
      <p>Appendix 2: List of idioms used as experimental stimuli
Gettare la maschera (‘to reveal oneself ’)</p>
      <sec id="sec-6-1">
        <title>Gettare la spugna (‘to give up’)</title>
        <p>Gettare acqua sul fuoco (‘to defuse a situation’)
Gettare olio sul fuoco (‘to inflame a situation ’)
Mettere la mano sul fuoco (‘to stake one’s life on</p>
      </sec>
      <sec id="sec-6-2">
        <title>Perdere il filo (‘ to lose the thread’) Mettere il carro davanti ai buoi (‘to put the cart Prendere il toro per le corna (‘to take the bull by Mettere le carte in tavola (‘to lay one’s cards on</title>
        <p>Prendere una cotta (‘to get a crush on somebody’)
Mettersi il cuore in pace (‘to resign oneself to sth’)
Mettere nero su bianco (‘to put sth down in black
Mettere il dito sulla piaga (‘to hit someone where</p>
      </sec>
      <sec id="sec-6-3">
        <title>Tirare la corda (‘to take sth too far’)</title>
        <p>Mettere i puntini sulle i (‘to be nitpicking’)</p>
      </sec>
      <sec id="sec-6-4">
        <title>Mettere zizzania (‘to sow discord’)</title>
      </sec>
      <sec id="sec-6-5">
        <title>Perdere la testa (‘to lose one’s head’)</title>
        <p>Perdere il treno (‘to miss an opportunity’)
Perdere la bussola (‘to lose one’s bearings’)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[Attardi and Dell'Orletta2009] Giuseppe Attardi and Felice Dell'Orletta</source>
          .
          <year>2009</year>
          .
          <article-title>Reverse revision and linear tree combination for dependency parsing</article-title>
          .
          <source>In Proceedings of NAACL 2009</source>
          , pages
          <fpage>261</fpage>
          -
          <lpage>264</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Baroni et al.2004]
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          , Silvia Bernardini, Federica Comastri, Lorenzo Piccioni, Alessandra Volpi, Guy Aston, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Mazzoleni</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Introducing the La Repubblica Corpus: A Large, Annotated, TEI(XML)-Compliant Corpus of Newspaper Italian</article-title>
          .
          <source>In Proceedings of LREC 2004</source>
          , pages
          <fpage>1771</fpage>
          -
          <lpage>1774</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Calzolari et al.2002]
          <string-name>
            <given-names>Nicoletta</given-names>
            <surname>Calzolari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Charles J.</given-names>
            <surname>Fillmore</surname>
          </string-name>
          , Ralph Grishman, Nancy Ide, Alessandro Lenci, Catherine MacLeod, and Antonio Zampolli.
          <year>2002</year>
          .
          <article-title>Towards best practice for multiword expressions in computational lexicons</article-title>
          .
          <source>In Proceedings of LREC 2002</source>
          , pages
          <fpage>1934</fpage>
          -
          <lpage>1940</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>[Dell'Orletta2009] Felice Dell'Orletta</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Ensemble system for Part-of-Speech tagging</article-title>
          .
          <source>In Proceedings of EVALITA</source>
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Evert and Krenn2005] Stefan Evert and Brigitte Krenn</source>
          .
          <year>2005</year>
          .
          <article-title>Using small random samples for the manual evaluation of statistical association measures</article-title>
          .
          <source>Computer Speech &amp; Language</source>
          ,
          <volume>19</volume>
          (
          <issue>4</issue>
          ):
          <fpage>450</fpage>
          -
          <lpage>466</lpage>
          . Special issue on Multiword Expression.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Fillmore et al.1988
          <string-name>
            <surname>] Charles</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Fillmore</surname>
          </string-name>
          , Paul Kay, and
          <string-name>
            <surname>Mary Catherine O'Connor</surname>
          </string-name>
          .
          <year>1988</year>
          .
          <article-title>Regularity and idiomaticity in grammatical constructions: the case of let alone</article-title>
          .
          <source>Language</source>
          ,
          <volume>64</volume>
          (
          <issue>3</issue>
          ):
          <fpage>501</fpage>
          -
          <lpage>538</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Goldberg1995]
          <string-name>
            <given-names>Adele</given-names>
            <surname>Goldberg</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Constructions. A Construction Grammar Approach to Argument Structures</article-title>
          . The University of Chicago Press, Chicago.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Goldberg2006]
          <string-name>
            <given-names>Adele</given-names>
            <surname>Goldberg</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Constructions at work</article-title>
          . Oxford University Press, Oxford.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Gries2008] Stefan Th. Gries</source>
          .
          <year>2008</year>
          .
          <article-title>Phraseology and linguistic theory: a brief survey</article-title>
          .
          <source>In Sylviane Granger and Fanny Meunier</source>
          , editors,
          <source>Phraseology: an interdisciplinary perspective</source>
          , pages
          <fpage>3</fpage>
          -
          <lpage>25</lpage>
          . John Benjamins, Amsterdam &amp; Philadelphia.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Hoffmann and Trousdale2013] Thomas Hoffmann and Graeme Trousdale, editors.
          <source>2013. The Oxford Handbook of Construction Grammar</source>
          . Oxford University Press, Oxford.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Lenci2014]
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Carving verb classes from corpora</article-title>
          .
          <source>In Raffaele Simone and Francesca Masini</source>
          , editors,
          <source>Word Classes. Nature</source>
          , typology and representations,
          <source>Current Issues in Linguistic Theory</source>
          , pages
          <fpage>17</fpage>
          -
          <lpage>36</lpage>
          . John Benjamins.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[Nissim and Zaninello2011] Malvina Nissim and Andrea Zaninello</source>
          .
          <year>2011</year>
          .
          <article-title>A quantitative study on the morphology of Italian multiword expressions</article-title>
          . Lingue e Linguaggio, X:
          <fpage>283</fpage>
          -
          <lpage>300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Sag et al.2002]
          <article-title>Ivan A</article-title>
          .
          <string-name>
            <surname>Sag</surname>
            , Timothy Baldwin, Francis Bond, Ann Copestake, and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Flickinger</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Multiword expressions: A pain in the neck for NLP</article-title>
          .
          <source>In Proceedings of CICLing 2002</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>[Shannon1948] Claude</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Shannon</surname>
          </string-name>
          .
          <year>1948</year>
          .
          <article-title>A mathematical theory of communication</article-title>
          .
          <source>The Bell System Technical Journal</source>
          ,
          <volume>27</volume>
          (
          <issue>3</issue>
          ):
          <fpage>379</fpage>
          -
          <lpage>423</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Squillante2014]
          <string-name>
            <given-names>Luigi</given-names>
            <surname>Squillante</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Towards an empirical subcategorization of multiword expressions</article-title>
          .
          <source>In Proceedings of the 10th Workshop on Multiword Expressions (MWE)</source>
          , pages
          <fpage>77</fpage>
          -
          <lpage>81</lpage>
          , Gothenburg, Sweden, April. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Tabossi et al.2011]
          <string-name>
            <given-names>Patrizia</given-names>
            <surname>Tabossi</surname>
          </string-name>
          , Lisa Arduino, and
          <string-name>
            <given-names>Rachele</given-names>
            <surname>Fanari</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Descriptive norms for 245 Italian idiomatic expressions</article-title>
          .
          <source>Behavior Research Methods</source>
          ,
          <volume>43</volume>
          :
          <fpage>110</fpage>
          -
          <lpage>123</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Wulff2008]
          <string-name>
            <given-names>Stefanie</given-names>
            <surname>Wulff</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Rethinking Idiomaticity: A Usage-based Approach</article-title>
          . Continuum.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [Wulff2009]
          <string-name>
            <given-names>Stefanie</given-names>
            <surname>Wulff</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Converging evidence from corpus and experimental data to capture idiomaticity</article-title>
          .
          <source>Corpus Linguistics and Linguistic Theory</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>131</fpage>
          -
          <lpage>159</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Zeldes2013]
          <string-name>
            <given-names>Amir</given-names>
            <surname>Zeldes</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Productive argument selection: Is lexical semantics enough? Corpus Linguistics</article-title>
          and
          <source>Linguistic Theory</source>
          ,
          <volume>9</volume>
          (
          <issue>2</issue>
          ):
          <fpage>263</fpage>
          -
          <lpage>291</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>