<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>WebIsAGraph: A Very Large Hypernymy Graph from a Web Corpus</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefano Faralli</string-name>
          <email>stefano.faralli@unitelmasapienza.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Irene Finocchi</string-name>
          <email>irene.finocchi@uniroma1.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simone Paolo Ponzetto</string-name>
          <email>simone@informatik.uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paola Velardi</string-name>
          <email>velardi@di.uniroma1.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Mannheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Rome Sapienza</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Rome Unitelma Sapienza</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we present WebIsAGraph, a very large hypernymy graph compiled from a dataset of is-a relationships extracted from the CommonCrawl. We provide the resource together with a Neo4j plugin to enable efficient searching and querying over such large graph. We use WebIsAGraph to study the problem of detecting polysemous terms in a noisy terminological knowledge graph, thus quantifying the degree of polysemy of terms found in is-a extractions from Web text.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Acquiring concept hierarchies, i.e., taxonomies
from text, is a long-standing problem in
Natural Language Processing (NLP). Much previous
work leveraged lexico-syntactic patterns, which
can be either manually defined
        <xref ref-type="bibr" rid="ref8">(Hearst, 1992)</xref>
        or automatically learned
        <xref ref-type="bibr" rid="ref19">(Shwartz et al., 2016)</xref>
        .
Pattern-based methods were shown by
        <xref ref-type="bibr" rid="ref17">(Roller et
al., 2018)</xref>
        to outperform distributional methods,
and can be complemented with state-of-the-art
meaning representations such as hyperbolic
embeddings
        <xref ref-type="bibr" rid="ref14 ref5 ref9">(Nickel and Kiela, 2017)</xref>
        to infer
missing is-a relations and filter wrong extractions
        <xref ref-type="bibr" rid="ref12">(Le
et al., 2019)</xref>
        . Complementary to these efforts,
researchers looked at ways to scale hypernymy
detection to very large, i.e., Web-scale corpora
        <xref ref-type="bibr" rid="ref22">(Wu
et al., 2012)</xref>
        . Recently,
        <xref ref-type="bibr" rid="ref18">(Seitner et al., 2016)</xref>
        applied Hearst patterns to the CommonCrawl1 to
produce the WebIsaDb. Using Web corpora makes
it possible to produce hundreds of millions of
isa triples: the extractions, however, include many
false positives and cycles
        <xref ref-type="bibr" rid="ref16">(Ristoski et al., 2017)</xref>
        .
      </p>
      <p>Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).</p>
      <p>1http://commoncrawl.org</p>
      <p>
        Methods for hypernym detection like, e.g.,
pattern-based approaches, have a limitation in
that they do not necessarily produce proper
taxonomies
        <xref ref-type="bibr" rid="ref3">(Camacho-Collados, 2017)</xref>
        :
automatically detected is-a relationships, on the other hand,
can be used as input to taxonomy induction
algorithms
        <xref ref-type="bibr" rid="ref16 ref21 ref5 ref6">(Velardi et al., 2013; Faralli et al., 2017;
Faralli et al., 2018, inter alia)</xref>
        . These
algorithms rely on the topology of the input graph,
and, therefore, cannot be applied ‘as-is’ to
Webscale resources like WebIsaDb, since this resource
merely consists of a set of triples. Moreover,
WebIsADb does not contain fully semantified triples,
i.e., subjects and objects of the is-a relationships
consist of potentially ambiguous terminological
nodes. This is because, due to their large size,
source input corpora like the CommonCrawl
cannot be semantified upfront. Linking to the
semantic vocabulary of a reference resource like
DBpedia
        <xref ref-type="bibr" rid="ref14 ref5 ref9">(Hertling and Paulheim, 2017)</xref>
        also barely
mitigate this problem, since Wikipedia-centric
knowledge bases have not, and cannot be expected to
have, complete coverage over Web data
        <xref ref-type="bibr" rid="ref13">(Lin et al.,
2012)</xref>
        .
      </p>
      <p>In this paper, we present an initial solution to
these problems by building the first very large
hypernymy graph, dubbed WebIsAGraph, built from
is-a relationships extracted from a Web-scale
corpus. This is a relevant task: although
WordNet (and other thesauri) already provides a
catalog of ambiguous terms, many nodes of
WebIsAGraph are not covered in available lexicographic
resources, because they are proper names,
technical terms, or polysemantic words. Our graph –
which we make freely available to the research
community to foster further work on Web-scale
knowledge acquisition – is built from the
WebIsADb on top of state-of-the-art graph mining
tools2: thanks to an accompanying plugin, it can
be easily searched, queried, and explored.
We2Neo4j: https://neo4j.com/
bIsAGraph may represent an opportunity to
researchers for investigating approaches to a variety
of tasks on large automatically acquired term
tuples. As an example, we use our resource to
investigate the problem of identifying ambiguous
terminological nodes. To automatically detect whether
a lexicographic node is ambiguous or not, we use
information from both the graph (topological
features) and textual labels (word embeddings) as
features to train a model using supervised
learning. Our results provide a first estimate of the
degree of polysemy that can be found among is-a
relationships from the Web.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Creating WebIsAGraph</title>
      <p>
        We created a directed hypernymy graph from the
WebIsADb
        <xref ref-type="bibr" rid="ref18">(Seitner et al., 2016)</xref>
        . WebIsADb is
a Web-scale collection of noisy hypernymy
relations harvested with 58 extraction patterns and
consisting of 607,621,170 tuples. Since the aim of
WebIsADb was to study the behaviour (on a large
scale) of Hearst-like extraction patterns, rather
than collecting relations with high precision, in
order to reduce noise (false positives) we
preselected the top-20 more precise extraction
patterns in (2016) from the original 58 and identified
385,459,302 tuples.
      </p>
      <p>After removing matches with a frequency lower
than 3 and isolated nodes, i.e., nodes with degree
equal to 0, we obtained a directed graph
consisting of 33,030,457 nodes and 65,681,899 directed
edges (see Table 1). The generation of such a large
graph required several weeks of computation on a
quad-core machine with 32 GB of RAM, using a
state-of-the art graph-db system, like Neo4j. Note
that the inherent sequential nature of the task of
indexing tuples, nodes and edges does not benefit
from the use of parallel computation. Next, we
developed efficient tools for graph querying,
which are released to the community, and
described in https://sites.google.com/
unitelmasapienza.it/webisagraph/,
where we also include examples of queries.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Measuring the polysemy of</title>
    </sec>
    <sec id="sec-4">
      <title>WebIsAGraph</title>
      <p>Let pSI (n) be the function that predicts if a
terminological node n corresponds to a monosemous or
a polysemous concept. We leverage a companion
sense inventory as a ground truth, and we train
different classifiers with a combination of topological
WebIsAGraph
nodes
edges
weakly connected components
nodes of largest component
Avg. node Degree
and textual features, described hereafter.
Topological features. Our conjecture is that in a
taxonomy-like terminological graph (even a noisy
one) there is a correlation between the mutual
connectivity of a node neighborhoods and its
polysemy. For example, consider the polysemous
word machine – which, according to WordNet,
has at least six heterogeneous meanings, ranging
from the ‘any mechanical or electrical device’ to
‘a group that controls the activities of a political
party’ – and the monosemous word floppy disk.
We expect to observe a different degree of
mutual connectivity across the corresponding
incoming and outgoing nodes. In particular, for
monosemous words, we expect a higher mutual
connectivity. With reference to Figure 1, left side, the
two hypernyms of ”floppy disk”: ”memory” and
”data storage”, have also ”RAM” as a common
hyponym. In contrast, nodes in the direct
neighborhood of ”machine” (leftmost graph in Figure
1) do not have mutual connections.</p>
      <p>Our aim is thus to identify topological features
that may help quantifying the previously described
connectivity properties. To cope with
scalability, we consider topological features built on top
of 1-hop/2-hop sub-graphs of a node n. Hence,
we identify two induced sub-graphs G +(n) and
G+ (n), induced on V +(n) = In(n) [v2In(n)
Out(v) and V + (n) = Out(n) [v2Out(n) In(v)
respectively, where In(x) and Out(x) are the sets
of incoming and outgoing nodes of x (including
x). Next, we remove from these sub-graphs the
node n, and compute the following features:
ccG + (n) and ccG+ (n): the resulting number
of weakly connected components;
vG + (n) and vG+ (n): the resulting number of
nodes;
eG + (n) and eG+ (n): the resulting number of
edges.</p>
      <p>With reference to the example of Figure 1, the
light gray sub-graph (a) is G +(n), the dark
subdevice
group
memory
data
storage
keyboard
system
machine</p>
      <p>vehicle
clan
party</p>
      <p>ROM</p>
      <p>
        RAM
floppy
disk
abandoned
hardware
(a)
(b)
computer
ship
car
zip disk
induced sub-graphs for ”machine” and ”floppy
disk”, (a) G +(n) in gray and (b) G+ (n) in dark
gray. Dashed edges connect each n with its
hypernyms and hyponyms.
graph (b) is G+ (n), and furthermore for n =
”machine”: ccG + (n) = 2, ccG+ (n) = 2,
vG + (n) = 5, vG+ (n) = 5, eG + (n) = 3, and
eG+ (n) = 3, while for the n =”floppy disk”:
ccG + (n) = 1, ccG+ (n) = 1, vG + (n) = 4,
vG+ (n) = 2, eG + (n) = 3, and eG+ (n) = 1.
Textual features. Similarly to topological
features, our hypothesis is that textual features of the
neighborhood nodes should exhibit a lower
average similarity when n is polysemous. We extract
textual features on top of pre-trained word
embeddings, widely adopted in many NLP-related
tasks
        <xref ref-type="bibr" rid="ref17 ref2">(Camacho-Collados and Pilehvar, 2018)</xref>
        .
      </p>
      <sec id="sec-4-1">
        <title>Formally, given a node n:</title>
        <p>#
W (n) is the word embedding vector of n
computed as follows:
#
W (n) =</p>
        <p>P
t2tokens(n)
jtokens(n)j
#
we(t)
where tokens(n) is the function that retrieves
the set of tokens composing the word n (e.g., if
n = hot dog, tokens(n) = fhot , dogg), and
#
we(t) is a pre-trained word embedding vector;
in(n)# and</p>
        <p>out(n): the cosine similarity
between W (n) and the average word embeddings
vector of incoming and outgoing nodes of n
respectively;
(1)
#
in(n) = CosSim(W (n);</p>
        <p>#
out(n) = CosSim(W (n);</p>
        <p>P
m2In(n)</p>
        <p>P
m2Out(n)
#
W (m)
#</p>
        <p>W (m)
jIn(n)j
jOut(n)j
) (2)
) (3)</p>
        <p>Algo.</p>
        <p>Rnd</p>
        <p>NN
t
eN ABC
d
r
o
WGBC</p>
        <p>Rnd
ia NN
d
e
pB ABC
D</p>
        <p>GBC</p>
        <p>
          Rnd
a
i
d
ep NN
B
D
te[ABC
N
d
ro GBC
W
Computing features. Topological features are
efficiently extracted using the query tool
mentioned in Section 2. To compute textual features
(see Section 3) we use the Glove pre-trained word
embedding vector
          <xref ref-type="bibr" rid="ref15">(Pennington et al., 2014)</xref>
          of
length 300 from the CommonCrawl.3
        </p>
        <p>By combining these two types of features
(topological and textual) we obtained three different
vector input representations consisting of 6 (only
topological features), 303 (only textual features)
and 309 (textual and topological) dimensions
respectively.</p>
        <p>Finally, we created three ”ground truth” sets of
nodes in the graph for which pSI (n) is known. We
selected a balanced number of monosemous and
polysemous nouns, using the following sense
inventories: i) WordNet (14,659 examples); ii)
DBpedia (17,041 examples); iii) WordNet and
DBpedia (31,701 examples).</p>
        <p>Algorithms.</p>
      </sec>
      <sec id="sec-4-2">
        <title>We compared four algorithms:</title>
        <p>
          Random (Rnd): a random baseline which
randomly classifies the ambiguity of a node;
Neural Network (NN): a neural network with
Softmax activation function in the output layer
and dropout
          <xref ref-type="bibr" rid="ref20">(Srivastava et al., 2014)</xref>
          ;
3https://nlp.stanford.edu/projects/glove/.
Two ensemble-based learning algorithms,
namely AdaBoost (ABC)
          <xref ref-type="bibr" rid="ref23">(Zhu et al., 2009)</xref>
          and
Gradient Boosting (GBC)
          <xref ref-type="bibr" rid="ref7">(Friedman, 2001)</xref>
          :
both have been shown to have high predictive
accuracy
          <xref ref-type="bibr" rid="ref11">(Kotsiantis et al., 2006)</xref>
          and are good
competitors of neural methods, especially with
very large datasets.
        </p>
        <p>
          Parameter selection. Based on the Area Under
Curve ROC (AUC) analysis
          <xref ref-type="bibr" rid="ref10">(Kim et al., 2017)</xref>
          ,
NN parameters have been empirically set as
follows: i) when testing only with topological
features (6 dimensions), we use 2 hidden layers with
4 and 2 neurons respectively and a dropout of 0.2
and 0.15; ii) when using only textual (303
dimensions), or combined textual and topological
features (309 dimensions), we use 4 hidden layers,
with 128, 64, 32 and 8 neurons respectively and a
dropout of 0.3,0.25,0.2 and 0.15.
        </p>
        <p>Results. We show in Table 2 the resulting
precision, recall and F1 of the five systems across the
ground truths datasets and for the combinations
of features (see Section 3). The metrics are
averaged on five classification experiments, with a
random split (85% train, 10% validation and 5%
test) of the ground truth sets. As shown in Table
2, NN outperforms the others ensemble methods,
obtaining a F1 score around 0.70. The comparison
of performances across the three combinations of
features reveals that topological features are not
enough to build a model for polysemy
classification but can slightly boost the overall already
compelling performances of word embeddings-based
features.</p>
        <p>
          In Table 3 we show the Person coefficient and
the distance correlation dCor4, with the aim of
analyzing how each feature correlates with the
polysemy observed in the three ground truth
dictionaries. We observed that the features with the
highest correlation with polysemy are eG+ , ccG +
and vG + (see Section 3). Additionally we
report the resulting weights of Permutation
Importance (PI) applied to the NN system with the
aim of measuring how the performance decreases
when a feature is perturbed, by shuffling its
values across training examples
          <xref ref-type="bibr" rid="ref1">(Breiman, 2001)</xref>
          .
We observed that the features which most
influenced the performances are in(n) (WordNet and
WordNet[DBpedia) and ccG + (DBpedia).
Furthermore, we found that although topological
features affect the performance only by a 1% in the
average, a number of topologically related
features, such as ccG + , vG + and eG+ are shown
to be indeed related with polysemy. In our
future work, we plan to create an ad-hoc
groundtruth sense dictionary, since especially WordNet
includes extremely fine-grained senses that do not
help validating our conjecture about reduced
mutual connectivity and contextual similarity of a
node’s neighborhood in case of monosemy.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The main contribution of this work is a new
resource obtained by converting a large dataset of
is-a (hypernymy) relations automatically extracted
from the Web (such as WebIsADb) into a graph
structure. This graph, along with its
accompanying search tools, enables descriptive and
predictive analytics of emerging properties of
termino4 and dCor are indexes to estimate how two distributions
are independent.
logical nodes. We used here our new resource to
investigate whether a node polysemy can be
predicted from its topological features (i.e.,
connectivity patterns) and textual features (meaning
representations from word embeddings). The results
of this preliminary study have shown that textual
features are good predictors of polysemy, while
topological features appear to be weaker
predictors even if they have a significant correlation with
the polysemy of the related node.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Leo</given-names>
            <surname>Breiman</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Random forests</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>45</volume>
          (
          <issue>1</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          , Oct.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>Jose´ Camacho-Collados and Mohammad Taher Pilehvar</article-title>
          .
          <year>2018</year>
          .
          <article-title>From word to sense embeddings: A survey on vector representations of meaning</article-title>
          .
          <source>J. Artif. Intell. Res.</source>
          ,
          <volume>63</volume>
          :
          <fpage>743</fpage>
          -
          <lpage>788</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Jose</surname>
          </string-name>
          Camacho-Collados.
          <year>2017</year>
          .
          <article-title>Why we have switched from building full-fledged taxonomies to simply detecting hypernymy relations</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>H. A.</given-names>
            <surname>David</surname>
          </string-name>
          .
          <year>1968</year>
          .
          <article-title>Gini's mean difference rediscovered</article-title>
          .
          <source>Biometrika</source>
          ,
          <volume>55</volume>
          (
          <issue>3</issue>
          ):
          <fpage>573</fpage>
          -
          <lpage>575</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Faralli</surname>
          </string-name>
          , Alexander Panchenko, Chris Biemann, and Simone Paolo Ponzetto.
          <year>2017</year>
          .
          <article-title>The contrastmedium algorithm: Taxonomy induction from noisy knowledge graphs with just a few links</article-title>
          .
          <source>In Proc. of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>1</volume>
          ,
          <string-name>
            <surname>Long</surname>
            <given-names>Papers</given-names>
          </string-name>
          , pages
          <fpage>590</fpage>
          -
          <lpage>600</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Faralli</surname>
          </string-name>
          , Irene Finocchi, Simone Paolo Ponzetto, and
          <string-name>
            <given-names>Paola</given-names>
            <surname>Velardi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Efficient pruning of large knowledge graphs</article-title>
          .
          <source>In Proc. of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19</source>
          ,
          <year>2018</year>
          , Stockholm, Sweden., pages
          <fpage>4055</fpage>
          -
          <lpage>4063</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Jerome H.</given-names>
            <surname>Friedman</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Greedy function approximation: A gradient boosting machine</article-title>
          .
          <source>The Annals of Statistics</source>
          ,
          <volume>29</volume>
          (
          <issue>5</issue>
          ):
          <fpage>1189</fpage>
          -
          <lpage>1232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Marti A.</given-names>
            <surname>Hearst</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>Automatic acquisition of hyponyms from large text corpora</article-title>
          .
          <source>In Proc. of COLING</source>
          , pages
          <fpage>539</fpage>
          -
          <lpage>545</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Sven</given-names>
            <surname>Hertling</surname>
          </string-name>
          and
          <string-name>
            <given-names>Heiko</given-names>
            <surname>Paulheim</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Webisalod: Providing hypernymy relations extracted from the web as linked open data</article-title>
          .
          <source>In The Semantic Web - ISWC 2017 - 16th International Semantic Web Conference</source>
          , Vienna, Austria,
          <source>October 21-25</source>
          ,
          <year>2017</year>
          ,
          <string-name>
            <given-names>Proc.</given-names>
            ,
            <surname>Part</surname>
          </string-name>
          <string-name>
            <surname>II</surname>
          </string-name>
          , pages
          <fpage>111</fpage>
          -
          <lpage>119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Chulwoo</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sung-Hyuk</surname>
            <given-names>Cha</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Yoo</given-names>
            <surname>An</surname>
          </string-name>
          , and Ned Wilson.
          <year>2017</year>
          .
          <article-title>On roc curve analysis of artificial neural network classifiers</article-title>
          .
          <source>In Florida Artificial Intelligence Research Society Conference.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>S. B. Kotsiantis</surname>
            ,
            <given-names>I. D.</given-names>
          </string-name>
          <string-name>
            <surname>Zaharakis</surname>
            , and
            <given-names>P. E.</given-names>
          </string-name>
          <string-name>
            <surname>Pintelas</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Machine learning: a review of classification and combining techniques</article-title>
          .
          <source>Artificial Intelligence Review</source>
          ,
          <volume>26</volume>
          (
          <issue>3</issue>
          ):
          <fpage>159</fpage>
          -
          <lpage>190</lpage>
          , Nov.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Matt</given-names>
            <surname>Le</surname>
          </string-name>
          , Stephen Roller, Laetitia Papaxanthos, Douwe Kiela, and
          <string-name>
            <given-names>Maximilian</given-names>
            <surname>Nickel</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Inferring concept hierarchies from text corpora via hyperbolic embeddings</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Lin</surname>
          </string-name>
          , Mausam, and
          <string-name>
            <given-names>Oren</given-names>
            <surname>Etzioni</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Entity linking at web scale</article-title>
          .
          <source>In Proc. of the Joint Workshop on Automatic Knowledge Base Construction and Web-scale Knowledge Extraction (AKBCWEKEX)</source>
          , pages
          <fpage>84</fpage>
          -
          <lpage>88</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Maximilian</given-names>
            <surname>Nickel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Douwe</given-names>
            <surname>Kiela</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Poincare´ embeddings for learning hierarchical representations</article-title>
          . In I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , and R. Garnett, editors,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , pages
          <fpage>6341</fpage>
          -
          <lpage>6350</lpage>
          . Curran Associates, Inc.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In EMNLP</source>
          , volume
          <volume>14</volume>
          , pages
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Petar</given-names>
            <surname>Ristoski</surname>
          </string-name>
          , Stefano Faralli, Simone Paolo Ponzetto, and
          <string-name>
            <given-names>Heiko</given-names>
            <surname>Paulheim</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Large-scale taxonomy induction using entity and word embeddings</article-title>
          .
          <source>In Proc. of the International Conference on Web Intelligence</source>
          , WI '
          <volume>17</volume>
          , pages
          <fpage>81</fpage>
          -
          <lpage>87</lpage>
          , New York, NY, USA. ACM.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Roller</surname>
          </string-name>
          , Douwe Kiela, and
          <string-name>
            <given-names>Maximilian</given-names>
            <surname>Nickel</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Hearst patterns revisited: Automatic hypernym detection from large text corpora</article-title>
          .
          <source>In Proc. of the 56th ACL (Volume 2: Short Papers)</source>
          , pages
          <fpage>358</fpage>
          -
          <lpage>363</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Julian</given-names>
            <surname>Seitner</surname>
          </string-name>
          , Christian Bizer, Kai Eckert, Stefano Faralli, Robert Meusel, Heiko Paulheim, and Simone Paolo Ponzetto.
          <year>2016</year>
          .
          <article-title>A large database of hypernymy relations extracted from the web</article-title>
          .
          <source>In Proc. of the Tenth International Conference on Language Resources and Evaluation LREC</source>
          <year>2016</year>
          ,
          <article-title>Portorozˇ</article-title>
          , Slovenia, May
          <volume>23</volume>
          -28,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Vered</given-names>
            <surname>Shwartz</surname>
          </string-name>
          , Yoav Goldberg, and
          <string-name>
            <given-names>Ido</given-names>
            <surname>Dagan</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Improving hypernymy detection with an integrated path-based and distributional method</article-title>
          .
          <source>In Proc. of the 54th ACL (Volume 1: Long Papers)</source>
          , pages
          <fpage>2389</fpage>
          -
          <lpage>2398</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Nitish</given-names>
            <surname>Srivastava</surname>
          </string-name>
          , Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Dropout: A simple way to prevent neural networks from overfitting</article-title>
          .
          <source>J. Mach. Learn. Res.</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1929</fpage>
          -
          <lpage>1958</lpage>
          , January.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Paola</given-names>
            <surname>Velardi</surname>
          </string-name>
          , Stefano Faralli, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Navigli</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Ontolearn reloaded: A graph-based algorithm for taxonomy induction</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>39</volume>
          (
          <issue>3</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Wentao</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Hongsong</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Haixun</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <surname>Kenny</surname>
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Probase: A probabilistic taxonomy for text understanding</article-title>
          .
          <source>In Proc. of the 2012 ACM SIGMOD International Conference on Management of Data, SIGMOD '12</source>
          , pages
          <fpage>481</fpage>
          -
          <lpage>492</lpage>
          , New York, NY, USA. ACM.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Ji</given-names>
            <surname>Zhu</surname>
          </string-name>
          , Hui Zou, Saharon Rosset, and
          <string-name>
            <given-names>Trevor</given-names>
            <surname>Hastie</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Multi-class adaboost</article-title>
          .
          <source>Statistics and Its Interface</source>
          ,
          <volume>2</volume>
          (
          <issue>3</issue>
          ):
          <fpage>349</fpage>
          -
          <lpage>360</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>