<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AHyDA: Automatic Hypernym Detection with feature Augmentation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ludovica Pannitto</string-name>
          <email>ellepannitto@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lavinia Salicchi</string-name>
          <email>lavinia.salicchi@libero.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Lenci</string-name>
          <email>alessandro.lenci@unipi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Pisa</institution>
        </aff>
      </contrib-group>
      <issue>0</issue>
      <abstract>
        <p>English. Several unsupervised methods for hypernym detection have been investigated in distributional semantics. Here we present a new approach based on a smoothed version of the distributional inclusion hypothesis. The new method is able to improve hypernym detection after testing on the BLESS dataset.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction and related works</title>
      <p>Within the Distributional Semantics framework,
semantic similarity between words is usually
expressed in terms of proximity in a semantic space,
where the dimensions of the space represent, at
some level of abstraction, the contexts in which
the words occur.</p>
      <p>Our intuitions about the meaning of words
allow inferences of the kind expressed in example
(1), and we expect Distributional Semantic
Models (DSMs) to support such inferences:
(1)
a. Wilbrand invented TNT ! Wilbrand
uncovered TNT
b. A horse ran ! An animal moved</p>
      <p>The type of relation between semantically
similar lexemes may differ significantly, but DSMs
only account for a generic notion of semantic
relatedness. Furthermore, not all lexical relations
are symmetrical (see example (2)), while most of
the similarity measures defined in distributional
semantics are, like the cosine.</p>
      <p>(2)
a. I saw a dog ! I saw an animal
b. I saw an animal 9 I saw a dog</p>
      <p>
        Hypernymy is an asymmetric relation.
Automatic hypernym identification is a very
wellknown task in literature, which has mostly been
addressed with semi-supervised, pattern-based
approaches
        <xref ref-type="bibr" rid="ref5 ref8">(Hearst, 1992; Pantel and Pennacchiotti,
2006)</xref>
        . Various unsupervised models have been
proposed
        <xref ref-type="bibr" rid="ref10 ref11 ref7 ref9">(Weeds and Weir, 2003; Weeds et al.,
2004; Clarke, 2009; Lenci and Benotto, 2012;
Santus et al., 2014)</xref>
        , based on the notion of
Distributional Generality
        <xref ref-type="bibr" rid="ref11">(Weeds et al., 2004)</xref>
        and on
the Distributional Inclusion Hypothesis (DIH)
        <xref ref-type="bibr" rid="ref4">(Geffet and Dagan, 2005)</xref>
        which has been derived
from it.
1.1
      </p>
      <sec id="sec-1-1">
        <title>The pitfalls of the DIH</title>
        <p>The DIH aims at providing a distributional
correlate of the extensional definition of hyponymy in
terms of set inclusion: x is a hyponym of y iff the
extension of x (i.e. the set of entities denoted by
x) is a subset of the extension of y. The DIH turns
this into the assumption that a significant number
of the most salient contexts of x should also
appear among the salient contexts of y. While this
is consistent with the logical inferences licensed
by hyponymy (cf. (2)), it does not take into
account the actual usage of hypernyms with respect
to hyponyms. Consider for instance the following
examples:
(3)
a. A horse gallops !? An animal gallops
b. A dog barks !? An animal barks
These inferences are truth-conditionally valid:
whenever the antecedent is true, the consequent is
also true. However, they are not equally
“pragmatically” sound. In fact, the fact that one uses
gallop
bark
a sentence like A dog barks does not entail that
in the same situation one would have also used
the sentence An animal barks. The latter
sentence would be pragmatically appropriate only in
cases in which one knows that something is
barking, without knowing which animal is producing
this sound. However, the latter condition hardly
applies, since barking is a very typical feature of
dogs: knowing that something is barking typically
entails knowing that it is a dog, since we know that
barking is something dogs do. The same argument
also applies to the case of horse and galloping.</p>
        <p>The problem of the DIH is that the assumption
it rests on, namely that the most typical contexts
of the hyponym are also typical contexts of the
hypernym, is not borne out in practical language
usage because of pragmatic constraints. The most
typical contexts of an hyponym are not
necessarily the typical contexts of its hypernym. This is
also proved by a simple inspection of corpus data,
as reported in Table 1. Despite animal (161; 107)
is more frequent than dog (128; 765) and horse
(90; 437), its co-occurrence with bark and gallop
is much lower than the ones of the hyponyms:
bark and gallop are not typical contexts of animal.</p>
        <p>If the inferences in (3) are pragmatically odd,
the following ones are instead fully acceptable:
a. A horse gallops ! An animal moves
b. A dog barks ! An animal calls
Salient features of the hypernym are indeed
supposed to be semantically more general than the
salient features of the hyponym. Santus et al.
(2014) tried to capture this fact by abandoning the
DIH and introducing an entropy-based measure to
estimate of informativeness of the hypernym and
hyponym contexts, under the assumption that the
former have a higher entropy, because they are
more general (e.g. move vs. gallop).</p>
        <p>In this paper, we address the same issue by
amending the DIH, to make it more consistent
with the actual distributional properties of
hyponyms and hypernyms. Therefore, we introduce
AHyDA (Automatic Hypernym Detection with
feature Augmentation), a smoothed version of the
DIH: given a context feature f that is salient for
a lexical item x, we expect co-hyponyms of x to
have some feature g that is similar to f , and an
hypernym of x to have a number of these clusters of
features. To remain in the animal sounds area, we
expect a dog to bark and a duck to quack and an
animal to produce either of those sounds or to
cooccur with a more general sound-emission verb.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>AHyDA: Smoothing the DIH</title>
      <p>All the measures implementing the DIH are based
on computing the (weighted) intersection of
the features of the hyponym and the hypernym.
This is then typically divided by the hyponym
features. AHyDA essentially proposes a new way
to compute the intersection of the hyponym and
hypernym contexts. Given a lexical item x, we
call Fx the set of its distributional features. Note
that features need not be pure lexical items. In
general, we define f as a pair (fw; fr) where fw
is typically a lexical item, and fr is any additional
contextual information, in the present case a
pattern occurring between x and fw, as explained
in section 3.1. The core novelty of AHyDA is to
use a smoothed version of Fx, called Fx0 .</p>
      <p>The idea is shown in figure 1, which provides
a simplified graphical example of the intersection
operation. Consider a case where the target
horse has some feature with gallop as a lexical
item, for example a feature f = (gallop; sbj)
meaning that horse is a possible subject of gallop.
Given what we have said in Section 1.1, we do
not expect animal to share this horse-specific
property. So, instead of looking for this
particular feature among the ones of animal, we
generate a new set Nhorse(gallop) of features
g = (gw; fr) such that gw is a neighbor of
gallop and is a feature (with the same syntactic
relation sbj) of some neighbor of horse.
Suppose that run, move, and cycle are neighbors of
gallop. As run and move are also features of
some neighbor of horse (e.g., lion), we would
have Nhorse(gallop) = fgallop; run; moveg.
Conversely, since cycle is not a feature of a close
neighbor of horse, it would not be included in the
expanded feature set.</p>
      <p>Mathematically, we define the expanded feature
set Fx0 as follows:</p>
      <p>Fx0 = f(f; Nx (f )) 8f 2 Fxg
Nx (f ) = fgjg = (gw; fr)g
(1)
(2)
where the following conditions hold for g:
d (fw; gw) &lt; k ^ 9yjd (x; y) &lt; h ^ g 2 Fy (3)
where d(x; y) is any distance measure in the
semantic space, k and h are empirically set
threshold values.</p>
      <p>Nx (f ) is generated by looking for features g that
are similar to fw, We then check whether this new
feature is shared by some neighbor of the target x,
and eventually include g in Nx (f ). This allows us
to redefine the intersection operation between Fx0
and Fy as:</p>
      <p>Fx0 \^Fy = ff jf 2 Fx ^ Nx (f ) \ Fy 6= ;g
(4)</p>
      <p>When expanding a feature f into Nx(f ), we
expect to find in Nx(f ) features that express
the same “property” in different ways. We
expect these features to be shared by hypernyms
more than co-hyponyms, because hypernyms are
supposed to collect features from all their
hyponyms, while co-hyponyms lack those of other
co-hyponyms (e.g. lions run but do not gallop).
AHyDA is thus defined as follows:</p>
      <p>AHyDA (x; y) =</p>
      <p>Pf2Fx jFx0 \ Fyj
jFxj
(5)</p>
      <p>
        Importantly, AHyDA only considers the
average cardinality of the intersections, without
looking at the feature weights. Moreover, the formula
is asymmetric (like the others implementing the
DIH), and therefore it is suitable to capture the
asymmetric nature of hypernymy.
Each lexical item u is represented with
distributional features extracted from the TypeDM
tensor
        <xref ref-type="bibr" rid="ref1 ref6">(Baroni and Lenci, 2010)</xref>
        . In TypeDM,
distributional co-occurrences are represented as a
weighted tuple structure, a set of ((u; l; v); ),
such that u and v are lexical items, l is a
syntagmatic co-occurrence link between u and v and is
the Local Mutual Information
        <xref ref-type="bibr" rid="ref3">(Evert, 2005)</xref>
        computed on link type frequency. Hence, each lexical
item u is represented in terms of features of the
kind (l; v).
      </p>
      <p>
        In addition to the sparse space, we also
produced a dense space of 300 dimensions
reducing the matrix with Singular Value Decomposition
(SVD). This additional space was used to retrieve
neighbors during the smoothing operation, as it
allowed us to perform faster and more accurate
calculations for cosines. The sparse space was
instead employed to retrieve features and get their
weights.
Evaluation was carried on a subset of the BLESS
dataset
        <xref ref-type="bibr" rid="ref2">(Baroni and Lenci, 2011)</xref>
        , consisting of
tuples expressing a relation between nouns.
      </p>
      <p>BLESS includes 200 English concrete nouns as
target concepts, equally divided between living
and non-living entities. For each concept noun,
BLESS includes several relatum words, linked to
the concept by one of the following 5 relations:
COORD (i.e. co-hyponyms), HYPER (i.e.
hypernyms), MERO (i.e. meronyms), ATTRI (i.e.
attributes), EVENT (i.e. verbs that define events
related to the target). BLESS also includes the
relations RANDOM-N, RANDOM-J, RANDOM-V,
which relate the targets to control tuples with
random noun, adjective and verb relata, respectively.</p>
      <p>
        By restricting to noun-noun tuples, we got
a subset containing these relations: COORD,
HYPER, MERO, RANDOM-N. We preprocessed the
dataset in order to exclude lexical items that are
not included in TypeDM. As reported in table 2,
the distribution (minimum, mean and maximum)
of the relata of all BLESS concepts is not even,
and therefore we took this into account while
relation
coord
hyper
mero
ran-n
avg
17.1
6.7
14.7
32.9
We compared AHyDA with a number of
directional similarity measures tested on BLESS, with
the goal of evaluating their ability to discriminate
hypernyms from other semantic relations, in
particular co-hyponyms. Given a lexical item x, Fx
is the set of its distributional features, wx(f ) is the
weight of the feature f for the term x:
WeedsPrec - quantifies the weighted inclusion of
the features of a term x within the features of a
term y
        <xref ref-type="bibr" rid="ref10 ref11 ref6">(Weeds and Weir, 2003; Weeds et al., 2004;
Kotlerman et al., 2010)</xref>
        the minimum and maximum number of items
holding a relation with x, and performed mmainxiimmuumm
random samples where each relation is presented with
minimum relata, and then averaged the results.
For example, consider the situation where x has
3 hypernyms, 6 co-hyponyms, 6 meronyms and
12 random nouns. In this situation, the minimum
number of relata for x would be 3, while the
maximum would be 12. Therefore, we would perform 4
random sampling for each relation, averaging the
results in order to obtain a singular measurement
for each relation in the end.
      </p>
      <p>We adopted the same evaluation methods
described in Lenci and Benotto (2012): plotting the
distribution of scores per relation across the BLESS
concepts, and calculating Average Precision (AP).
3.4</p>
      <sec id="sec-2-1">
        <title>Results</title>
        <p>Table 3 summarizes the Average Precision
obtained by AHyDA, the other DIH-based measures,
and the cosine. Although AHyDA’s improvement
is not big in hypernym detection, co-hyponyms get
lower values of AP, thus showing that smoothing
the intersection allows a better discrimination
between the two classes. It is worth remarking that
PPf2fF2xF\xFwywx(xf()f ) (6) t(hh2ieg0h1ve2ar)l,uthbeaesncaftouhrsoestheoefretohpteohreetrveadmlubeayatisoLunreenoscniatarhenedgbeBanleaennroaclteltyod
random samples of relations we have adopted. We
also reported, in table 4, the AP values obtained
through the standard measures, without
employPf2Fx\Fy min(wx(f ); wy(f )) ing the feature augementation procedure. Altough
Pf2Fx wx(f ) vmaaluinesdiffofrerhenycpeesrnayrme sindothenoctocohrdanvgaelumesu,chw,hitchhe
are generally higher without feature augmentation.</p>
        <p>As mentioned in section 3.1, the results for all the
measures are obtained using the sparse space. The
reduced space was employed to compute the
Cosine baseline.</p>
        <p>As regards the AP values for hypernyms, we
CD(x; y)) (8) must notice that not all hypernyms in BLESS share
the same status: some of them are what we would
consider logic entailments (e.g. eagle ! bird),
others depict taxonomic relations (e.g. alligator
! chordate), some are not true logic entailments
(e.g. hawk !? predator)</p>
        <p>Figure 2 shows the average score produced with
the new measure. Here hypernyms are neatly
set apart from co-hyponyms, whereas the distance
with meronyms and with the control group,
randoms, is less significative.</p>
        <p>
          WeedsPrec(x; y) =
ClarkeDE - a variation of WeedsPrec, proposed in
(Clarke, 2009)
ClarkeDE(x; y) =
(7)
invCL - a new measure introduced in
          <xref ref-type="bibr" rid="ref7">(Lenci and
Benotto, 2012)</xref>
          , to take into account not only the
inclusion of x in y but also the non-inclusion of
y in x. The measure is defined as a function of
ClarkeDE (CD).
        </p>
        <p>invCL(x; y) = pCD(x; y)(1</p>
        <p>We used the cosine as a baseline, since it is
a symmetric similarity measure and is commonly
used to evaluate semantic similarity/relatedness in
DSMs. In the definition of Nx(f ), the target and
feature neighbors are identified with the cosine,
setting the k and h parameters to 0.8 and 0.9
respectively.</p>
        <p>To avoid biases due to the relata distribution
among concepts, for each target x, we computed</p>
        <p>Figure 3 shows the average scores produced by
AHyDA when applied to the reverse hypernym
pair. It is interesting to notice that in this case
AHyDA produces basically the same results as
random pairs. This suggests that AHYDA
correctly predicts that hyponyms entail hypernyms,
but not vice versa, thereby capturing the
asymmetric nature of hypernymy.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>The Distributional inclusion hypothesis has
proven to be a viable approach to hypernym
detection. However, its original formulation
rests on an assumption that does not take into
consideration the actual usage of hypernyms in
texts. In this paper we have shown that, by adding
some further pragmatically inspired constraints,
a better discrimination can be achieved between
co-hyponyms and hypernyms. Our ongoing work
focuses on refining the way in which the
smoothing is performed, and testing its performance on
other datasets of semantic relations.
ceedings of the GEMS 2011 Workshop on
GEometrical Models of Natural Language Semantics, pages
1–10. Association for Computational Linguistics.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Distributional memory: A general framework for corpus-based semantics</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>36</volume>
          (
          <issue>4</issue>
          ):
          <fpage>673</fpage>
          -
          <lpage>721</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>How we blessed distributional semantic evaluation</article-title>
          .
          <source>In ProDaoud Clarke</source>
          .
          <year>2009</year>
          .
          <article-title>Context-theoretic semantics for natural language: an overview</article-title>
          .
          <source>In Proceedings of the workshop on geometrical models of natural language semantics</source>
          , pages
          <fpage>112</fpage>
          -
          <lpage>119</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Evert</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>The statistics of word cooccurrences: word pairs and collocations</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Maayan</given-names>
            <surname>Geffet</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ido</given-names>
            <surname>Dagan</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>The distributional inclusion hypotheses and lexical entailment</article-title>
          .
          <source>In Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics</source>
          , pages
          <fpage>107</fpage>
          -
          <lpage>114</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Marti A</given-names>
            <surname>Hearst</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>Automatic acquisition of hyponyms from large text corpora</article-title>
          .
          <source>In Proceedings of the 14th conference on Computational linguisticsVolume 2</source>
          , pages
          <fpage>539</fpage>
          -
          <lpage>545</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Lili</given-names>
            <surname>Kotlerman</surname>
          </string-name>
          , Ido Dagan, Idan Szpektor, and
          <string-name>
            <surname>Maayan</surname>
          </string-name>
          Zhitomirsky-Geffet.
          <year>2010</year>
          .
          <article-title>Directional distributional similarity for lexical inference</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>16</volume>
          (
          <issue>4</issue>
          ):
          <fpage>359</fpage>
          -
          <lpage>389</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Lenci</surname>
          </string-name>
          and
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Benotto</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Identifying hypernyms in distributional semantic spaces</article-title>
          .
          <source>In Proceedings of the First Joint Conference on Lexical and Computational Semantics-Volume 1: Proceedings of the main conference and the shared task, and Volume 2: Proceedings of the Sixth International Workshop on Semantic Evaluation</source>
          , pages
          <fpage>75</fpage>
          -
          <lpage>79</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Patrick</given-names>
            <surname>Pantel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Pennacchiotti</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Espresso: Leveraging generic patterns for automatically harvesting semantic relations</article-title>
          .
          <source>In Proceedings of the 21st International Conference on Computational Linguistics and the 44th annual meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>113</fpage>
          -
          <lpage>120</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Enrico</given-names>
            <surname>Santus</surname>
          </string-name>
          , Alessandro Lenci, Qin Lu, and Sabine Schulte Im Walde.
          <year>2014</year>
          .
          <article-title>Chasing hypernyms in vector spaces with entropy</article-title>
          .
          <source>In EACL</source>
          , pages
          <fpage>38</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Julie</given-names>
            <surname>Weeds</surname>
          </string-name>
          and David Weir.
          <year>2003</year>
          .
          <article-title>A general framework for distributional similarity</article-title>
          .
          <source>In Proceedings of the 2003 conference on Empirical methods in natural language processing</source>
          , pages
          <fpage>81</fpage>
          -
          <lpage>88</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Julie</given-names>
            <surname>Weeds</surname>
          </string-name>
          , David Weir,
          <string-name>
            <given-names>and Diana</given-names>
            <surname>McCarthy</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Characterising measures of lexical distributional similarity</article-title>
          .
          <source>In Proceedings of the 20th international conference on Computational Linguistics</source>
          , page 1015.
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>