<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Clustering verbal Objects: manual and automatic procedures compared</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vít Baisa Lexical Computing Ltd. Czech Republic</string-name>
          <email>vit.baisa@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Elisabetta Jezek University of Pavia Department of Humanities</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ilaria Colucci University of Pavia Department of Humanities</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>As highlighted by Pustejovsky (1995, 2002), the semantics of each verb is determined by the totality of its complementation patterns. Arguments play in fact a fundamental role in verb meaning and verbal polysemy, thanks to the sense co-composition principle between verb and argument. For this reason, clustering of lexical items filling the Object slot of a verb is believed to bring to surface relevant information about verbal meaning and the verb-Objects relation. The paper presents the results of an experiment comparing the automatic clustering of direct Objects operated by the agglomerative hierarchical algorithm of the Sketch Engine corpus tool with the manual clustering of direct Objects carried out in the T-PAS resource. Cluster analysis is here used to improve the semantic quality of automatic clusters against expert human intuition and as an investigation tool of phenomena intrinsic to semantic selection of verbs and the construction of verb senses in context.</p>
      </abstract>
      <kwd-group>
        <kwd>Clustering</kwd>
        <kwd>verbal Objects</kwd>
        <kwd>Italian</kwd>
        <kwd>Semantic Types</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Clustering techniques have been used extensively
in recent decades in Linguistics and NLP,
especially in Word Sense related tasks. As a matter of
fact, partitioning data sets on the basis of their
similarity at a distributional level clarifies the
meaning of lexical elements
        <xref ref-type="bibr" rid="ref4">(Brown et al., 1991)</xref>
        .
Partitioning verbal arguments, for example, can
be beneficial to investigate the sense properties
they share but also to explore verbal meaning.
In fact, as highlighted by
        <xref ref-type="bibr" rid="ref16">Pustejovsky (1995</xref>
        ,
2002), the semantics of each verb is determined
by the totality of its complementation patterns and
arguments play a fundamental role in verb
meaning and verbal polysemy, thanks to the sense
cocomposition principle. Id est, the process of
bilateral semantic selection between the verb and its
complement gives rise to a novel sense of the verb
in each context of use (ibidem).
      </p>
      <p>Clustering lexical items filling the argument
positions of a verb is then believed to bring to surface
relevant information about verbal meaning and
the verb-arguments relation. Clustering them, and
especially direct Objects in pro-drop languages
such as Italian, allows hence to investigate how to
better induce, discriminate and disambiguate verb
senses. Because argument fillers share the same
semantic nature, they can be grouped and
generalized with respect to their content and be
associated with semantic types, i.e. empirically
identified semantic classes representing selectional
properties and preferences of verbs.</p>
      <p>
        Clustering of Objects can therefore be used as a
survey tool for the intrinsic phenomena of
semantic classes and, at the same time, as an object of
investigation to improve the clustering automatic
models themselves against human partitioning.
This paper presents the results of an experiment
comparing manual and automatic clustering of
Italian Object fillers to be used in verb-sense
identification and, along with it, it describes the
linguistic phenomena that emerged from the
semantic analysis of non-supervised clusters. The
comparison concerns the agglomerative hierarchical
clustering algorithm of the Sketch Engine corpus
tool1
        <xref ref-type="bibr" rid="ref12">(Kilgarriff et al., 2014)</xref>
        and the manual
clustering carried out in the T-PAS resource2 (
        <xref ref-type="bibr" rid="ref11">Ježek at
al., 2014</xref>
        ), in which verbal senses are identified in
context based on the fillers of the argument
positions (see section 1.1) and are annotated with a
semantic type (ST; see section 1.2) able to identify
them. Thanks to their semantic generalization
properties, ST are also believed to represent a
useful comparative tool between manual and
automatic clustering. After presenting the theoretical
background of the research, section 2 will cover
1 www.sketchengine.eu/
2 tpas.fbk.eu
data, method and work pipeline, while clustering
evaluation via metrics and linguistic analysis will
be presented in section 3.
(1) pilotare
1. [Human] pilotare ([Flying Vehicle] | [Water
Vehicle])
      </p>
    </sec>
    <sec id="sec-2">
      <title>1.1 Clustering verbal Objects fillers</title>
      <p>
        Clustering is a Data Mining task (Kotu &amp;
Deshpande, 2014) in which a grouping process of a set
of objects is carried out, obtaining clusters of
elements which are similar to each other but
dissimilar from the objects of other groups
        <xref ref-type="bibr" rid="ref20">(Xu &amp;
Wunsch, 2008)</xref>
        . In most implementations, clustering is
used with an exploratory function, i.e. it is a
technique applied to data sets for which there is no a
priori knowledge concerning the set membership
of the samples
        <xref ref-type="bibr" rid="ref15">(Lavine and Mirjankar, 2006)</xref>
        . In
these cases, clustering is therefore considered as a
non-supervised procedure with the aim of
providing an insight into the studied data. However, it
can be considered a supervised method and
regarded as a classification task when a manually
created benchmark (a ground truth) is used to
assess the output of the clustering
        <xref ref-type="bibr" rid="ref3">(Bishop, 1995)</xref>
        .
The manually created partition or the manually
defined set of classes is used to validate the
groupings proposed by the automatic algorithm,
through a process defined as external clustering
evaluation
        <xref ref-type="bibr" rid="ref6">(Gan, Ma and Wu, 2007)</xref>
        . The idea
behind this paper is to operate through a procedure
very similar to external evaluation in which the
manual clustering and the automatic one taken
into consideration are mutually compared; but yet
here the aim is not to validate the automatic model
but more to bring out matches and differences
between the partitioning criteria at the basis of the
supervised clustering and the unsupervised one.
The supervised clustering under consideration
here was performed on the lexical items that fill
different argument positions in T-PAS, a resource
of predicate-argument structures for Italian
obtained from corpora (
        <xref ref-type="bibr" rid="ref11">Ježek at al., 2014</xref>
        ). T-PAS
contains, for each argument slot, the specification
of the semantic class to which the fillers found in
that position in the corpus belong. We considered
the direct Object clusters, which therefore contain
the fillers that occupy that slot in the various
occurrences of the corpus. To clarify this, given the
following sense for verb pilotare (to pilot), the
related cluster for the Object position will appear as
follows:
3 See
        <xref ref-type="bibr" rid="ref13">Kilgarriff et al. (2015)</xref>
        for statistics and technical details
on similarity computing.
pilotare_clust1: {macchina (car), moto
(motorbike), barca (boat), caccia (fighter aircraft), nave
(ship)}
The ST defined for the direct Object slot can thus
also be used as a label to semantically identify
what is contained in the cluster.
      </p>
      <p>
        As for automatic clustering, in our comparison we
used the built-in clustering function
        <xref ref-type="bibr" rid="ref13 ref2">(Baisa et al.,
2015)</xref>
        in the Sketch Engine tool (SkE). The model
is based on a hierarchical agglomerative algorithm
that compute the distributional similarity 3
between the Object fillers and groups them in an
unsupervised way, starting from a minimum
similarity value given to the algorithm
        <xref ref-type="bibr" rid="ref12">(Kilgarriff et al.,
2014)</xref>
        . Clusters creation starts with computing
Word Sketches, i.e. automatic, corpus-based
summaries of a word’s grammatical and collocational
behaviour
        <xref ref-type="bibr" rid="ref14">(Kilgarriff et al., 2004)</xref>
        . The results
concerning the direct Object are then grouped
through a bottom-up process in which clusters are
populated through pairings of words. The
inclusion and exclusion criterion is a minimum default
value of 0.15 4 for distributional similarity. The
clusters created in Sketch Engine for pilotare are
the followings, for which, unlike T-PAS, ST
labelling is not available:
(2) pilotare_clust1: {nave (ship), barca (boat)}
pilotare_clust2: {macchina (car), moto
(motorbike)}
      </p>
      <p>pilotare_clust3: {caccia (fighter aircraft)}5
The main difference between T-PAS and SkE
clustering procedures are the
semantic-distributional criteria on which they are based. T-PAS
approach can be defined as verb-oriented: Objects
are primarily clustered on the basis of their verbal
distributional behaviour and ability to activate a
given verbal sense as direct objects. Since all
fillers occupying a given slot for a given sense share
the same relation with the verb, they can be
ontologically and semantically generalized with an ST
on the basis of their common semantic traits. This
generalization allows to make the verbal
selectional constraints visible. On the contrary, SkE
4 We also conducted a similarity value manipulation
experiment, which confirmed what discussed in detail in section 3.
5 Clusters consist at least of 1 word and up to 1000.
performs noun-based clustering: it takes into
account the general distributional behaviour of
fillers, not merely the verbal one. In the process of
creating sets, each filler behaviour is weighed
against the entire reference corpus and with
respect to the frequencies of appearance in different
contexts. The elements clustered together in SkE
are therefore not only similar in their sense and
behaviour as direct objects, but also respect to the
whole nominal class they belong to.</p>
    </sec>
    <sec id="sec-3">
      <title>1.2 T-PAS System of Semantic Types</title>
      <p>
        As mentioned above, in T-PAS argument slots are
linked to ST labels, semantic classes able to
generalize over the sets of lexical items in argument
positions found in the corpus (
        <xref ref-type="bibr" rid="ref11">Ježek at al., 2014</xref>
        ).
The labels belong to the System of Semantic
Types (see Figure 1 for an excerpt), a hierarchical
structure of semantic categories achieved by
performing the CPA procedure
        <xref ref-type="bibr" rid="ref8">(Hanks, 2004)</xref>
        , on the
evidence of 1200 Italian verbs (
        <xref ref-type="bibr" rid="ref10">Ježek, 2019</xref>
        ), i.e.
through the manual analysis of examples in
corpora of slots’s fillers and their co-occurrence
statistics. They characterize a group of lexical
elements with respect to their content, defining also
a criterion of similarity and dissimilarity on which
T-PAS clusters are created. STs are used here as a
reference for the comparison of the two clustering
models, for the verification of the clusters internal
semantic quality.
      </p>
      <sec id="sec-3-1">
        <title>Data and method</title>
        <p>
          The research has been developed through a
pipeline organized according to the following steps:
1. Data extraction: Data for both clusterings are
extracted from the web crawled corpus ItWac
reduced
          <xref ref-type="bibr" rid="ref1">(Baroni et al., 2009)</xref>
          . In this early stage the
clusters of Object fillers for each verb included in
T-PAS are extracted from the corpus annotated
lines, while for Sketch Engine, the clusters are
extracted for all verbs present in the ItWac corpus.
All lines in the corpus are then scanned and verbal
Objects are mapped through the condition: OBJ =
post verbal noun (PostV_N). Since T-PAS does
not annotate individual fillers as such but only
works at verb and sentence level, this function is
also used to retrieve its Objects.
        </p>
        <p>2. Data intersection: The obtained clusters are
intersected with each other in order to obtain a
database in which there are sets for the same verbs
and containing the same fillers, to focus on how
the two models carried out the partition.</p>
        <p>3. Data filtering: In this step the database is
cleared from:
a) verbs with structures recognized as complex
and non-compositional, i.e. idiomatic
constructions;
b) verbs with the ST [Anything] (top node in Fig.
1) in the object slot, as it does not entail selection
restrictions within the T-PAS clusters;
c) verbs with Object clusters with more than 29
internal elements.</p>
        <p>At the end of the filtering process the clusters of
the two models are aligned with respect to the
STs, i.e. all possible STs signaled in T-PAS for
the Object of a verb are treated as a single set of
semantic conditions, in order to analyze the
internal quality of SkE clusters through them. The
aligned structure of the verb acquisire (to acquire)
in Figure 2 is given here as an example.
tpas_acquisire
ske_acquisire
clust_1
clust_2
clust_3
[Institution], [Artifact]
The final database comes to a total of 397 verbs
and 3938 clusters, including both T-PAS and SkE
clusters. We provide an illustrative table (Table 1)
showing the first and last verb among those
analyzed and the information on their respective
clusters: abbagliare (to dazzle) and votare (to vote).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3.1 SkE clustering evaluation</title>
      <p>To verify the compatibility between the two
clusterings, the similarity between the two partitions
has been evaluated through different metrics able
to offer an external evaluation of the unsupervised
model. To account for both the presence of
common pairings, as well as the homogeneity and
completeness of the clustering, the following
metrics were considered: Fowlkes &amp; Mallow Index
(F&amp;M), Adjusted Rand Index (ARI),
Homogeneity, Completeness.</p>
      <p>
        F&amp;M, as the geometric mean between precision
and recall, was used to verify the similarity
between the two models from how many partition
pairings are in common. This index also allows to
better balance the possible noise or unrelatedness
between clustering
        <xref ref-type="bibr" rid="ref5">(Fowlkes &amp; Mallows, 1983)</xref>
        .
ARI
        <xref ref-type="bibr" rid="ref9">(Hubert &amp; Arabie, 1985)</xref>
        always gives
information on the overlapping of the two clusterings
in comparison but balances the very large number
of clustered elements in T-PAS
        <xref ref-type="bibr" rid="ref18">(Romano et al.,
2016)</xref>
        . Homogeneity and completeness metrics
        <xref ref-type="bibr" rid="ref19">(Rosenberg &amp; Hirschberg, 2007)</xref>
        are helpful to
better investigate the internal content of the SkE
clusters. They allow to highlight a possible
internal structure, hierarchically and semantically
coherent with the taxonomy identified for ST.
Homogeneity evaluates if all automatic clusters
created contain only elements that are members of a
single class in the manual reference.
Completeness, instead, evaluates if all the objects that are
members of a given cluster in SkE are elements of
the same cluster in T-PAS.
      </p>
      <p>
        As reported by their respective creators, all
metrics have an optimal result range between 0 and 1.
The possible results between these two limits can
be classified with respect to the greater or lesser
proximity to the optimal limit: the results closer to
1 denote greater similarity of output between the
two models, the results closer to 0 instead less
similarity
        <xref ref-type="bibr" rid="ref6">(Gan, Ma and Wu, 2007)</xref>
        .
      </p>
      <p>In this sense, we can define three bands of
possibilities, coherently with the approach the higher
the better generally used in cluster analysis: from
0.01 to 0.399, the clustering compared to the
golden standard is highly different, from 0.4 to
0.699 the result and the correspondence is
medium-good, while the results above 0.7 and up to
0.999 are the ideal ones, which indicate a marked
correspondence between the compared models.
However, since metrics such as F&amp;W and ARI
have shown the lack of partitions in higher ranges
(the first beyond 0.82, the latter beyond 0.7), we
choose to consider the whole group of medium
good results between 0.4 and 1. The absence of
the higher ranges stands for low compatibility
between the two models.</p>
      <p>Metric
Adjusted Rand Score
Fowlkes &amp; Mallow
Homogeneity
Completeness
Clusters in the
[0.4-1] range
11.08%
36.18%
94.96%
41.31%
As we see in Table 2, what we find in fact is a
situation of only limited correspondence between
the two clustering, with a rather low overlap and
similarity as indicated by the ARI and the F&amp;M,
even with the internal noise balance. At least two
reasons may be behind the scarce similarity: the
tendency of SkE to create small fine-grained
clusters populated by few elements that give more
weight to specificity than to generalization
capacity; the fact that in T-PAS for a given verb sense
the Object slot can be compatible with more STs
(see (1)), and such STs can also be hierarchically
distant in the general system of labels. This leads
to clusters containing fillers able to activate a
given verbal sense but which are quite
heterogeneous among themselves and semantically
dissimilar, with respect to the rest of the
distributional relations between the fillers. An example
can be the verb trasportare (to transport), which
has as T-PAS cluster for the first sense a set of 18
Objects (see (4)); such fillers belong to three
different STs: [Inanimate], [Animate] and [Energy].
The latter ST, [Energy], is hierarchically distant to
the others since it has a different parent node than
[Animate] and [Inanimate], which both pertain to
a lower level in the hierarchy.
(3) trasportare
1. [Human] | [Vehicle] | [Watercourse]
trasportare [Inanimate] | [Animate] | [Energy]
(4) trasportare_clust1: {acqua (water), alimento
(nourishment), animale (animal), arma (weapon),
bene (asset/good), bicicletta (bicicle), cadavere
(corpse), cibo (food), gas (gas), gommone
(inflatable raft), macchina (machine), oggetto (object),
peso (weight), student (student), terra (soil),
traffic (traffic), viaggiatore (traveller), visitatore
(visitor)}
It is clear that a fine-grained algorithm, not able to
generalize at a higher level as in T-PAS, will
divide fillers labelled with [Animate] or [Inanimate]
from those labelled with [Energy]. In fact, SkE for
the same verb creates 12 clusters.</p>
      <p>As shown in Table 2, the results of the
Completeness are in line with what has just been discussed
for ARI and F&amp;M: only in 40% of the cases all
members of a T-PAS cluster are members of a
single SkE cluster. These are generally small or
medium sized clusters with only one associated ST
or with hierarchically close alternative ST
structures. Homogeneity highlights the primary
characteristic of SkE clusters and the algorithm: it is
preferable to create smaller but internally purer
clusters, rather than larger sets with members of
other classes. This implies the creation in SkE of
semantically specific clusters, that privilege the
inter-relation between Object fillers but not the
higher semantic level between Object fillers and
verb.</p>
      <p>From a different perspective, we can say that the
noun-oriented criteria of clustering and the
verboriented ones tend to converge when we consider
small clusters, in which the elements belonging to
a set in SkE generally belong to the same set in
TPAS.</p>
      <p>As for wide clusters, they are particularly rare in
SkE and tend to be smaller in size than T-PAS
anyway. Their content also seems to be dependent
on various factors on which the linguistic analysis
has shed light.</p>
    </sec>
    <sec id="sec-5">
      <title>3.2 Linguistic analysis of the clusters</title>
      <p>To verify the nature of the diversity between the
two clusterings measured with the metrics
reported in 3.1, a detailed analysis of the
lexical-semantic phenomena visible internally to the
clusters was carried out considering:
- The consistency, for automatic clusters, with
one and only one of the aligned T-PAS STs, i.e.
the precision and purity at the semantic level of
clusters compared to the generalization of the ST;
- Internal homogeneity, i.e. whether the clusters
meet verb-sense oriented or noun-sense oriented
criteria and, if the latter, whether the cluster items
are linked by syntagmatic relationships and there
is some kind of affinity or implication between
them. Thus, the types of semantic relations
present between the words are identified;
- The overlap between clusters with respect to the
ARI, and in relation to cluster size and clustering
difficulty depending on several STs possible for
the same slot;
- The problem of incorrect mapping as Objects of
postverbal Subjects, subjects of inaccusative
verbs, structures with si particle (e.g. reflexive,
impersonal), i.e. the clusters' internal noise.
The research has shown that SkE clusters tend to
be small-medium sized, semantically
homogeneous, often able to isolate very specific semantic
relations. They are generally not consistent per se
with the verb sense identified by T-PAS but create
partitions: a) usually of medium size and
consistent with only one parallel ST, b) single
element groups that generally belong to a higher
level of specificity or to a different semantic
domain, and c) groups that are inconsistent with the
sense of the verb but cluster words on the basis of
the following criteria:
- Belonging to the same domain (e.g. informatics
for distribuire {software, applicazione});
- Being part of the same ST, but as very specific
instances, not separated by the T-PAS hierarchy
(e.g. {abbazia, monastero, santuario} for
saccheggiare and the type [Location]);
- The possibility of a conceptual association or
affinity (e.g. {seminario, incontro, seduta} for
organizzare);
- Purely distributional parameters and undefined
semantic relations (e.g. in gestire {contenuto,
caso});
- A relationship of synonymy or meronymy (e.g.
{spinta, propensione} for frenare or for fratturare
{dito, mano, braccio}); antonymy, hyponymy,
hyponymy are generally represented by different
clusters.</p>
      <p>The parameters of consistency, internal
homogeneity and overlapping between the models seem
to relate to the same factors: first, the size of the
clusters, i.e. how many clustered elements are part
of the set; second, the structure of STs possible for
the Object (see (5)), i.e. if for the same slot only
one ST is possible, if several alternatives are
available or, also, if a lexical set is signaled in the
T-PAS annotation - that is, if among the fillers a
set of lexical elements is present that has high
frequency or has the typical behaviour of a
collocation (e.g. {messaggio | ricordo} in (5)). This is
relevant since the computation of SkE starts
precisely from the frequency and collocational
behaviour of a word.
(5) cancellare (sense description: to eliminate, to
make inexistent):
1. [Human] | [Inanimate1] | [Abstract Entity1] |
[Eventuality1] cancellare [Inanimate2] | [Abstract
Entity2 {messaggio | ricordo}] | [Eventuality2]
The third relevant factor is hierarchical proximity,
i.e. if STs possible for a slot are sisters of the same
parent node between the types present in the
hierarchy (see (6)).
(6) no proximity: [Animate] vs. [Institution]
in proximity: [Command] vs [Request]
The clusters of SkE, even if not corresponding to
those of T-PAS, are rated totally consistent or at
70-80% consistent with one of the aligned STs
62% of the times; in the remaining 38% of cases,
there is a significant number of clusters that can
be generalized with a ST. Consistency is more
difficult to reach if a given sense is annotated in
T-PAS with several alternative STs or STs and
lexical sets co-presenting; SkE can produce new
combinations in which fillers corresponding to
different STs are included in the same cluster.
Very frequently SkE atomizes the set of fillers in
nuclear clusters, made up of only one or two
elements that are necessarily consistent with one ST
but are not of much help to the study of semantic
relations.</p>
      <p>As said, hierarchical proximity of STs and the size
of the cluster can influence its handling: if the
possible STs for the Obj-slot are hierarchically distant
and the T-PAS cluster is small, the SkE outcome
will tend to be heterogeneous and inconsistent.
Consider the verb affogare (to drown), which has
[Animate] and [Emotion] as possible STs for the
same slot but different senses. T-PAS clusters are
small, clust_1 counts 3 elements and clust_2
counts 7, in which the two possible types are
clearly distinguished. One would expect to find
the two senses separated in the SkE clusters as
well, since [Emotion] belongs to lower levels of
the hierarchy and has a different parent node of
[Animate]. However, for SkE we find clusters
such as:
(7) affogare_clust3: {figlio (son), bimbo (child),
pensiero (thought)}
If, on the contrary, we consider closer ST, the
clustering will be homogeneous, even if not
verbsense-oriented because too fine-grained.
Medium-sized clusters (more than 10 clustered
elements) seem to perform quite well both with
hierarchically close and distant STs. A useful
example can be smarrire (to lose) in T-PAS (8), for
which SkE presents the clusters in (9):
(8) smarrire
1. [Human] smarrire [Artifact]
2. [Human] | [Human group] smarrire [Concept] |
[Property]
(9) smarrire_clust1: {borsello (man bag)}
smarrire_clust2: {significato (meaning),
ragione (reason), memoria (memory), pensiero
(thought)}</p>
      <p>smarrire_clust3: {capacità (capacity),
consapevolezza (awareness)}</p>
      <p>smarrire_clust4: {nozione (notion), certezza
(certainty)}
smarrire_clust5: {senno (sense)}
smarrire_clust6: {fiducia (trust), voglia (will)}
smarrire_clust7: {documento (document)}
smarrire_clust8: {cellulare (mobile phone)}
Considering that for smarrire T-PAS creates only
two clusters (see (10)), the sets in (9) also
highlight the general tendency of SkE to create
smaller, semantically highly fine-grained clusters.
(10) smarrire_clust1: {borsello (man bag),
cellulare (mobile phone), documento (document)}
smarrire_clust2: {consapevolezza
(awareness), capacità (capacity), certezza (certainty),
fiducia (trust), memoria (memory), nozione
(notion), pensiero (thought), ragione (reason),
senno (sense), significato (meaning), voglia
(will)}
Large clusters are generally difficult to handle
because ideally portionable in more distributionally
cohesive groups. As regards the problem of
internal noise, due to the PostV_N relation, we can
note that the phenomenon is pervasive and
important, since it affects the results of internal
coherence and homogeneity. It is, however, a
phenomenon that can be curbed with a revision of the
extraction function. What emerges from the
analysis is a distance in the general structure of the two
clustering results but a good compatibility from
the internal semantic point of view. T-PAS
privileges rather more complex semantic groupings on
a level of co-composition between verb and
meaning, linked to conceptual operations of
generalization. On the contrary, SkE creates complex and
homogeneous structures of relations inside data,
even if sometimes this implies clusters that are too
fragmented and not always optimal also from a
noun-oriented perspective. T-PAS seems to
pertain to a higher level of granularity respect to SkE,
whose clusters can be considered as possible
subpartitions of the STs.
4</p>
      <sec id="sec-5-1">
        <title>Conclusion</title>
        <p>The paper presented the statistical and linguistic
results of a comparison between SkE
unsupervised clustering model and the manual and
verbsense oriented clustering of T-PAS. It highlighted
how the noun-oriented model and the
verb-oriented one are not overlapping if not partially. The
SkE clustering, even if not overlapping, can still
be considered as internally compatible with the
TPAS partition, since the homogeneity metric
reaches good results. The internal linguistic
analysis allowed to identify the semantic quality
through the consistency with a semantic type, the
internal homogeneity, the adherence with the
verb-oriented approach of T-PAS. The reasons
that regulate the fragmentation of clusters in SkE,
i.e. motivations that follow a fine-grained logic,
were then presented. The analysis made possible
to shed a light on the semantic compatibility
between the two approaches, which seem to pertain
to different levels of granularity.</p>
        <p>The difference in the partition output and the
parallel semantic compatibility allows us to claim
that the SkE automatic clustering is more useful
for the internal investigation of STs than to
investigate the verb-Object co-composition relation. It
would be interesting to conduct further
comparisons between other automatic clustering
techniques and that of T-PAS, to investigate additional
semantic implications of clustering through
nounbased and verb-based approaches.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernardini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferraresi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Zanchetta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>The WaCky wide web: a collection of very large linguistically processed web-crawled corpora</article-title>
          .
          <source>In Language resources and evaluation</source>
          ,
          <volume>43</volume>
          (
          <issue>3</issue>
          ):
          <fpage>209</fpage>
          -
          <lpage>226</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Baisa</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>El Maarouf</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rychlý</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Rambousek</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Software and Data for Corpus Pattern Analysis</article-title>
          .
          <source>In RASLAN</source>
          ,
          <fpage>75</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bishop</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>Neural networks for pattern recognition</article-title>
          .
          <source>Cambidge UK</source>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Brown</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Della Pietra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Della Pietra</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mercer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>1991</year>
          ).
          <article-title>Word sense disambiguation using statistical methods</article-title>
          .
          <source>In Proceedings of the 29th Meeting of the Association for Computational Linguistics (ACL-91)</source>
          ,
          <fpage>264</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Fowlkes</surname>
            ,
            <given-names>E. B.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Mallows</surname>
            ,
            <given-names>C. L.</given-names>
          </string-name>
          (
          <year>1983</year>
          ).
          <article-title>A method for comparing two hierarchical clusterings</article-title>
          .
          <source>In Journal of the American statistical association</source>
          ,
          <volume>78</volume>
          (
          <issue>383</issue>
          ):
          <fpage>553</fpage>
          -
          <lpage>569</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Gan</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , Ma,
          <string-name>
            <given-names>C.</given-names>
            , &amp;
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Data clustering: theory, algorithms, and applications</article-title>
          .
          <source>Society for Industrial and Applied Mathematics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Hanks</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>1996</year>
          ).
          <article-title>Contextual dependency and lexical sets</article-title>
          . In
          <source>International journal of corpus linguistics</source>
          ,
          <volume>1</volume>
          (
          <issue>19</issue>
          ):
          <fpage>75</fpage>
          -
          <lpage>98</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Hanks</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Corpus pattern analysis</article-title>
          .
          <source>In Euralex Proceedings</source>
          ,
          <volume>1</volume>
          :
          <fpage>87</fpage>
          -
          <lpage>98</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Hubert</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Arabie</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>1985</year>
          ).
          <article-title>Comparing partitions</article-title>
          .
          <source>In Journal of classification</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>193</fpage>
          -
          <lpage>218</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Ježek</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Sweetening Ontologies Cont'd: Aligning bottom-up with top-down ontologies</article-title>
          .
          <source>In JOWO.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Ježek</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magnini</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feltracco</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bianchini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>T-PAS: A resource of corpusderived Types Predicate-Argument Structures for linguistic analysis and semantic processing</article-title>
          .
          <source>In Proceedings of LREC</source>
          ,
          <fpage>890</fpage>
          -
          <lpage>895</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Kilgarriff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baisa</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bušta</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakubíček</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovář</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michelfeit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rychlý</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suchomel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>The Sketch Engine: ten years on</article-title>
          .
          <source>In Lexicography</source>
          ,
          <volume>1</volume>
          :
          <fpage>7</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Kilgarriff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baisa</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bušta</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakubíček</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovář</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michelfeit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rychlý</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suchomel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2015</year>
          ). Statistics used in Sketch Engine. (document available at: https://www.sketchengine.eu/wp-content/uploads/ske-statistics.pdf)
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Kilgarriff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rychly</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smrz</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Tugwell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>The Sketch Engine</article-title>
          . In Information Technology,
          <volume>105</volume>
          :
          <fpage>116</fpage>
          -
          <lpage>126</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Lavine</surname>
            ,
            <given-names>B. K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Mirjankar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Clustering and classification of analytical data</article-title>
          .
          <source>Encyclopedia of Analytical Chemistry: Applications</source>
          , Theory and Instrumentation.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Pustejovsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>The Generative Lexicon</article-title>
          . Cambridge MA. MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Pustejovsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Syntagmatic processes</article-title>
          . In Cruse, A. D.,
          <string-name>
            <surname>Hundsnurscher</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Job</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lutzeier</surname>
          </string-name>
          , P. (eds.)
          <article-title>Lexicology: A Handbook on the Nature and Structure of Words and Vocabularies</article-title>
          . de Gruyter.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Romano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinh</surname>
            ,
            <given-names>N. X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bailey</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Verspoor</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Adjusting for chance clustering comparison measures</article-title>
          . In
          <source>The Journal of Machine Learning Research</source>
          ,
          <volume>17</volume>
          (
          <issue>1</issue>
          ):
          <fpage>4635</fpage>
          -
          <lpage>4666</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Rosenberg</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Hirschberg</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>V-measure: A conditional entropy-based external cluster evaluation measure</article-title>
          .
          <source>In Proceedings of EMNLP-CoNLL</source>
          <year>2007</year>
          ,
          <fpage>410</fpage>
          -
          <lpage>420</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Wunsch</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2008</year>
          ). Clustering. Hoboken NJ. John Wiley &amp; Sons.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>