<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Learning Semantic Relatedness from Human Feedback Using Relative Relatedness Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thomas Niebler</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Becker</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Politz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Hotho</string-name>
          <email>hothog@informatik.uni-wuerzburg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data Mining and Information Retrieval Group, University of Wurzburg</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>L3S Research Center Hanover</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>An important topic in Semantic Web research is to learn ontologies from text. Here, assessing the degree of semantic relatedness between words is an important task. However, many existing relatedness measures only encode information contained in the underlying corpus and thus do not directly model human intuition. To solve this, we propose RRL (Relative Relatedness Learning) to improve existing semantic relatedness measures by learning from explicit human feedback. Human feedback about semantic relatedness is extracted from the publicly available MEN dataset. The core result is that we can generalize human intuition on datasets such as MEN using RRL. This way, we can signi cantly outperform semantic relatedness scores produced by current state-of-the-art methods.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>An important topic in Semantic Web research is to learn ontologies from text. In
this scenario, assessing the semantic relatedness of words as perceived by humans
is a crucial task. Often, the relatedness score of two words is approximated by
calculating the cosine of their vector representations.</p>
      <p>Problem Setting and Approach. While many methods using this approach come
close to human intuition, they can only encode information from the underlying
corpus and thus do not explicitly represent the actual notion of semantic relatedness as
employed by humans. A natural way to solve this is to incorporate explicit human
feedback, in order to account for the deviations of the respective semantic
relatedness measure from human intuition. This can be achieved using metric learning.
However, most metric learning algorithms use constraints such as \w and w0 are
similar " or \w is more similar to w0 than to w00". With such constraint
formulations, human relatedness scores, i.e., absolute information about the degree of
relatedness (e.g., \w and w0 are 54% related "), cannot be used. To address this
issue, we propose Relative Relatedness Learning (RRL), which exploits these scores
to learn a semantic relatedness measure which ts human intuition by formulating
relative constraints in the form \w1 is more related to w10 than w2 to w20".
Contribution. The core result is that we can generalize human intuition from
semantic relatedness datasets using RRL. We can signi cantly improve the measured
semantic relatedness scores beyond the current state-of-the-art. Using two large,
public word embedding datasets, we con rm this by learning from and
evaluating on the MEN collection, which contains relatedness information generated from
human feedback.</p>
      <p>loss(M ) = P 0:5</p>
      <p>C(H)
+ tr(M) log det(M)</p>
      <p>n2
M M l rloss(M )
M mMi0n fkM M 0kF jM 0 2 P SDg</p>
      <p>max n0; p1
Data: V Sn 1: word vectors; H: semantic relatedness dataset (e.g. MEN);
learning rate l
Result: a relatedness matrix M for Equation (1)
Let M := In
while M not converged do
cosM (w1; w10)
p1
cosM (w2; w20)o 2
end
return M
Algorithm 1: The RRL algorithm to learn a relatedness measure from relative
relatedness information. M is updated using projected gradient descent,
regularization is performed via log det divergence.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Relative Relatedness Learning (RRL)</title>
      <p>
        Given in Algorithm 1, we propose a supervised approach with a custom loss function
to learn a symmetric, positive semide nite (PSD) matrix M to parameterize the
cosine measure so that it better measures semantic relatedness:
xT M y
cosM (x; y) := pxT M xpyT M y
(1)
Our algorithm is inspired by a metric learning approach called LSML [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] which
uses relative distance comparisons to learn a linear metric characterized by the
matrix M . In contrast, the training constraints C(H) in our algorithm are relative
relatedness comparisons:
      </p>
      <p>C(H) := f(w1; w10; w2; w20) : rel(w1; w10) &gt; rel(w2; w20)g
and are collected from semantic relatedness datasets H := f(wi; wi0; rel(wi; wi0))g
such as MEN, which contain word pairs (wi; w0) together with relatedness scores
i
rel(wi; wi0) collected from human feedback. Each word is represented by a normalized
vector, e.g., from a set of vector embeddings.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Datasets</title>
      <p>
        We use two word embedding datasets and a semantic relatedness dataset with
relatedness scores collected through human feedback to evaluate RRL.
WikiGloVe [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This dataset was trained on 6 billion tokens from Wikipedia
articles from a 2014 dump and the Gigaword 5 corpus using the GloVe embedding
algorithm and consists of 400,000 vectors with dimension 300.3
3 https://nlp.stanford.edu/projects/glove/
ConceptNet Numberbatch [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Speer et al. combined Word2Vec and GloVe
embeddings with relations from the semantic network ConceptNet to receive 426,572
300-dimensional word vectors currently posing the state-of-the-art on MEN.4
The MEN collection [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The MEN dataset contains 3,000 word pairs together
with human-generated scores about their perceived semantic relatedness.5 These
scores re ect human feedback, which we use both to train our relatedness measure
as well as for evaluation.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <p>
        In this section, we perform two experiments in order to demonstrate the usefulness
of RRL for learning semantic relatedness. First we train several metrics on both
vector datasets considering di erent amounts of user feedback and secondly assess
the robustness of the learned measures by training on false information. We publish
our code to enable reproducibility of our experiments.6
Experiment Setup. For both experiments, we randomly split MEN into a 80%
training and a 20% test set. In the second experiment, we replace the
relatedness scores in the training set by new random scores completely uncorrelated ( &lt;
0:0005) to the original training scores, while the test scores stay the same. From
the training data, we then sampled subsets of di erent sizes (10% - 100%) on which
we train a metric each. The metric is evaluated on the previously sampled 20%
test data by applying a standard approach of comparing arti cial relatedness scores
produced by the metric with human-collected ones using the Spearman correlation
coe cient (cf. [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]). We repeat sampling training sets and training a metric 25
times. Then, for each training sample size, we take the mean of the scores produced
by the 25 trained metrics. In all training cases, the standard deviation was negligible
so we do not report it here. As a baseline, we also report the Spearman correlation
using the standard cosine measure on the 20% test data.
      </p>
      <p>
        Integrating Di erent Levels of User Intentions. We rst investigate how
the amount of user feedback used for training in uences the quality of the learned
semantic relatedness measure. Figure 1 shows that we can inject user feedback
information about semantic relatedness into our measure (dashed line, diamond markers)
and in doing so, improve the t of our measure to human intuition signi cantly. On
the ConceptNet embeddings, it appears that we have reached a maximum boundary
of achievable correlation on unseen data of 0.88. This is very close to the inter
annotator agreement reported in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Furthermore, it is important to note that although
the correlation improvements seem very small, i) correlation scores are nonlinear,
i.e., improving a high correlation score is much more di cult than improving a
low correlation score, ii) with increasing amount of training data, the number of
constraints grows roughly quadratically and iii) all di erences are signi cant at
p &lt; 0:05 with at least 50% training data when comparing mean correlation scores
with a Fisher transformation.
4 https://github.com/commonsense/conceptnet-numberbatch/tree/16.09
5 http://clic.cimec.unitn.it/~elia.bruni/MEN
6 http://dmir.org/semmele
cos
metric true
metric false
cos
metric true
      </p>
      <p>metric false
20%
80%</p>
      <p>100%
20%
40% Sample Size</p>
      <p>60%
(a) WikiGloVe</p>
      <p>Robustness of the Learned Semantic Relatedness Measure. Now we inject
false user feedback into RRL to see if we can in uence the score in not only a positive,
but also a negative direction. Figure 1 shows that false user feedback ( &lt; 0:0005 to
the original scores) exhibits a large negative in uence on the learned metric (dotted
line, star markers), as expected. Nevertheless, we assume that the score decrease is
mitigated by the inherent semantic content of the embeddings. Overall, this shows
that although we can improve our measure's t to human intuition, we need a high
semantic quality of both the word embeddings and the latent collected relatedness
scores through human feedback. Furthermore, we need a certain minimum amount
of training data to produce signi cantly improved results.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this work, we presented an approach to learn semantic relatedness from
human intuition, using a relative constraint formulation. The core result is that we
can inject this intuition into a relatedness measure with which we can produce
signi cantly improved results compared to the standard cosine measure and more
realistically assess human intuition of semantic relatedness. A noteworthy result is
that we can even outperform the current state-of-the-art correlation with MEN on
the ConceptNet embeddings, thus de ning a new state-of-the-art result on MEN.
Acknowledgements. This work has been partially funded by the DFG grant \Posts II"
and the BMBF funded junior research group \CLiGS" (grant identi er FKZ 01UG1408).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Elia</given-names>
            <surname>Bruni</surname>
          </string-name>
          ,
          <string-name>
            <surname>Nam-Khanh Tran</surname>
          </string-name>
          , and Marco Baroni. \Multimodal Distributional Semantics.
          <article-title>" In: JAIR (</article-title>
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E. Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          et al. \
          <article-title>Metric Learning from Relative Comparisons by Minimizing Squared Residual."</article-title>
          <source>In: ICDM. Dec</source>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Je</surname>
            <given-names>rey Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          . \
          <source>Glove: Global Vectors for Word Representation." In: EMNLP</source>
          . Vol.
          <volume>14</volume>
          .
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Robert</given-names>
            <surname>Speer</surname>
          </string-name>
          , Joshua Chin, and Catherine Havasi.
          <source>\ConceptNet 5</source>
          .
          <article-title>5: An Open Multilingual Graph of General Knowledge."</article-title>
          <source>In: AAAI</source>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>