<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>August</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards Analogy-based Recommendation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benchmarking of Perceived Analogy Semantics</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Analogy-Enabled Recommendation</institution>
          ,
          <addr-line>Relational Similarity, Analogy Benchmarking</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Christoph Lo Web Information Systems - TU Del Mekelweg 4 Del</institution>
          ,
          <addr-line>Netherlands 2628CD</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Nava Tintarev Web Information Systems - TU Del Mekelweg 4 Del</institution>
          ,
          <addr-line>Netherlands 2628CD</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <volume>31</volume>
      <issue>2017</issue>
      <abstract>
        <p>Requests for recommendation can be seen as a form of query for candidate items, ranked by relevance. Users are however oen unable to crisply dene what they are looking for. One of the core concepts of natural communication for describing and explaining complex information needs in an intuitive fashion are analogies: e.g., “What is to Christopher Nolan as is 2001: A Space Odyssey to Stanley Kubrick?”. Analogies allow users to explore the item space by formulating queries in terms of items rather than explicitly specifying the properties that they nd aractive. One of the core challenges which hamper research on analogy-enabled queries is that analogy semantics rely on consensus on human perception, which is not well represented in current benchmark data sets. erefore, in this paper we introduce a new benchmark dataset focusing on the human aspects for analogy semantics. Furthermore, we evaluate a popular technique for analogy semantics (word2vec neuronal embeddings) using our dataset. e results show that current word embedding approaches are still not not suitable to suciently deal with deeper analogy semantics. We discuss future directions including hybrid algorithms also incorporating structural or crowd-based approaches, and the potential for analogy-based explanations.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>In this paper we explore one method of eciently communicating
information needs or for explaining results of an inquiry used in
natural human interaction, namely analogies. Using analogies in
natural speech allows communicating dense information easily and
naturally by implying that the “essence” of two concepts is similar
or at least perceived similarly. us, analogies can be used to map
factual and behavioural properties from one (usually beer known
concept, the source) to another (usually less well known, the target)
concept. is is particularly eective for natural querying and
explaining when only vague domain knowledge is available.</p>
      <p>In this paper we consider two types of analogy queries, 4-term
analogy completion queries (like “What is to Christopher Nolan as
is 2001: A Space Odyssey to Stanley Kubrick?”), and 4-term
analogon ranking (like “What is similar to ’2001: A Space Odyssey to
ComplexRec 2017, Como, Italy.
2017. Copyright for the individual papers remains with the authors. Copying permied
for private and academic purposes. is volume is published and copyrighted by its
editors. Published on CEUR-WS, Volume 1892..
Stanley Kubrick’? a) e Fih Element to Christopher Nolan, b)
Memento to Christopher Nolan, c) Dunkirk to Christopher Nolan).
Analogy completion queries might be seen an extension on classical
critiquing in recommender systems which can be formulated in
terms of “like x, but with properties y modied”. In critiquing the
feature (price) and the modication (cheaper) needs to be explicit,
whereas in an analogy, the semantics are implicitly given by seing
terms in relation, which is interpreted based on both
communication partners’ conceptualization of that domain (e.g., “e Fih
Element is like 2001: A Space Odyssey but by Scorsese” carries a
lot of implicit information).</p>
      <p>Analogies typically are represented as rhetorical gures of speech
which need to be interpreted and reasoned on using the receiver’s
background knowledge - which is dicult for information systems
to mimic. is complex semantic task is further complicated by the
lack of useful benchmark datasets. Most current benchmarks are
restricted in scope, usually focusing on syntactic features instead
of relevant properties of analogies as for example their suitability
for transferring information (which is central when one wants to
use analogies for querying, or explaining recommendations).</p>
      <p>erefore, one of our core contributions is an improved
benchmark dataset for analogy semantics which focuses specically on
the usefulness of an analogy with respect to querying and
explaining, and providing it as a tool for guiding future research into
analogy queries in recommender systems. In this work, we are
focusing on general domain analogies. is allows us the choose
the right technologies for future adaption in a domain-specic
recommender system. In detail, our contributions are as follows:
Discuss dierent properties of analogy semantics, and highlight
their importance for querying and explaining recommendations.
Introduce a new benchmark dataset systematically built on top
of existing sets, rectifying many current limitations. Especially,
we focus on perceived diculty of analogies, and the quality and
usefulness of analogies.</p>
      <p>Showcase and discuss the performance of an established word
embedding-based algorithm on our test set.
2</p>
      <p>
        DISCUSSIONS ON ANALOGY SEMANTICS
e semantics of analogies have been researched in depth in the
elds of philosophy, linguistics, and in the cognitive sciences, such
as [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], or [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. ere have been several models for analogical
reasoning from philosophy (like works by Plato, Aristotle, or Kant
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]), while other approaches see analogies as a variant of induction
in formal logic [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], or mapping the structures of relationships and
properties of entity pairs (structure mapping theory [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]). However,
those analogy denitions are rather complex and hard to grasp
computationally, and thus most recent works on computational analogy
processing rely on the simple 4-term analogy model which is given
by two sets of word pairs (the so-called analogons), with one pair
being the source and one pair being the target. A 4-term analogy
holds true if there is a high degree of relational similarity between
those two pairs. is is denoted by »a1; a2¼ :: »b1; b2¼, where the
relationship between a1 and a2 is similar to the relation between b1 and
b2, as for example in »StarW ars; Sci f i¼ :: »ForrestGump; Comedy¼
(both are dening movies within their genres). is model has
several limitations, as is discussed by Lo in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]: the semantics of
“a high degree of relational similarity” from an ontological point
of view is unclear, and the model ignores human perception and
abstractions (e.g., analogies are usually not right or wrong, but
beer or worse based on how well humans understand them).
      </p>
      <p>
        erefore, in this paper we promote an improved
interpretation of the 4-term analogy model [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], and assume that there
can be multiple relationships between the concepts of an
analogon, some of them being relevant for the semantics of an analogy
(the dening relationships), and some of them not. An analogy
holds true if the sets of dening relationships of both analogons
show a high degree of relational similarity. For illustrating the
dierence and importance of this change in semantics, consider
the analogy statement: »StanleyKubrick; 2001 : ASpaceOdyssey¼ ::
»T axiDriver ; MartinScorsese¼. Kubrick is the director of 2001: A
Space Odyssey, and Scorsese is the director of Taxi Driver. Both
analogons contain the same “movie is directed by person” relationship,
and this could be considered a valid analogy with respect to the
simple 4-term analogy model. Still, this is a poor analogy statement
from a human communication point of view because 2001: A Space
Odyssey is not like Taxi Driver at all. erefore, this statement
does neither describe the essence of 2001: A Space Odyssey nor
the essence of Taxi Driver particularly well: both movies are iconic
for their respective directors, but they do not for example describe
that one movie is in the science ction genre, and the other could
be classied as a (violent) crime movie. Understanding which
relationships actually dene the essence of an analogon from the
viewpoint of human perception is a very challenging problem, but
this understanding is crucial for judging the usefulness and value of
an analogy statement. Furthermore, the degree to which
relationships are dening an analogon may vary with dierent contexts.
In short, there can be beer or worse analogies based on two core
factors (we will later encode the combined overall quality with an
analogy rating): the degree of how well the relationships shared
by both analogons are dening them (i.e., are the relationships
shared between both analogons indeed the dening relationships
which describe the intended semantic essence), and the relational
similarity of the shared relationships. Based on these observations,
we dene two basic types of analogy queries (loosely adopted from
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]) which can be used for analogy-enabled recommender systems:
(1) Analogy completion ? : »a1; a2¼ :: »b1; ?¼
is query can be used to nd the missing concept in a 4-term
analogy. is is therefore the most useful query type in a future
analogy-enabled information systems [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Solving this query
requires identifying the set of dening relationships between
a1 and a2, and then nding a b2 such that the set of dening
relationships between b1 and b2 is similar.
(2) Analogon ranking multiple-choice
? : »a1; a2¼ ::?f»b1; b2¼; »c1; c2¼ :::g
A simpler version of the general analogon ranking query are
multiple choice ranking queries as they are for example used in
the SAT benchmark dataset (discussed below). Here, the set of
potential result analogons is restricted, and an algorithm would
simply need to rank the provided choices (e.g.,»b1; b2¼; »c1; c2¼)
instead of freely discovering the missing analogon.
      </p>
      <p>
        Previous approaches to processing analogies algorithmically
cover prototype systems operating on Linked Open Data (LOD), as
for example [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], but also approaches which mine analogies from
natural text [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. A very popular recent trend in Natural Language
Processing (NLP) is training neuronal word embeddings [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. e
popular word2vec implementation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] learns relational similarity
between word pairs from large natural language corpora by
exploiting the distributional hypothesis [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] using neuronal networks.
Word-embeddings are particularly interesting for use in
analogyenabled recommender systems as they can be easily trained on
big text corpora (like for example user reviews), and do not
require structured ontologies and taxonomies like other approaches
(e.g., [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]). In this paper, we evaluate the analogy reasoning
performance of word embeddings using our new benchmark dataset
which represents analogy semantics more naturally.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>BENCHMARK DATASETS</title>
      <p>For benchmarking the eectiveness of analogy algorithms, there are
several established Gold standard datasets for general-knowledge
analogies (i.e. not tailored to a specic domain). However, all of
them are lacking in some respect, as discussed in the following
sections.
3.1</p>
      <p>
        Mikolov Benchmark Dataset &amp; Wordrep
e Mikolov Benchmark set [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] is one of the most popular
benchmark sets for testing the analogy reasoning capabilities of neuronal
word embeddings covering 19,558 4-term analogies with 14 distinct
relationship types, focusing exclusively on analogy completion
queries »a; b¼ :: »c; ?¼. Nine of these relationships focus on
grammatical properties (e.g., the relationship “is plural for a noun”), while
ve relationships are of a semantic nature (e.g., “is capital city of a
country” like »Athens; Greece¼ :: »Oslo; N orway¼ ). e benchmark
set is generated by collecting pairs of entities which are members
of the selected relationship type from Wikipedia and DBpedia, and
then combining all these pairs into 4-term analogy tuples. e
Wordrep dataset [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] extends the Mikolov set by adding more challenges,
and expanding to 25 dierent relationship types.
      </p>
      <p>A core weakness of this type of test set is that it does not include
any human judgment with respect to the dening relationship types
and the usefulness as an analogy for querying or explaining, but
instead focuses only on “correct” relationships which do usually
not carry any of the subtleties of rhetorical analogies, as e.g., “is
city in” or ”is plural of”. In short, the Mikolov test set does not
benchmark if algorithms can capture analogy semantics from a
human perspective, but instead focuses purely on relational similarity
for a very limited selection of relationship types. us, we feel that
this benchmark dataset does not represent the challenge of real
world analogy queries well.</p>
      <p>
        SAT Analogy Challenges
e SAT analogy challenge [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] plays an important role in
realworld college admission tests in the United States to assess a
prospective student’s vocabulary depth and general analytical skills by
focusing on analogon ranking queries. As the challenge’s original
intent is to assess the vocabulary and reasoning skills of prospective
students, it contains many rare words. In order to be able to
evaluate the test without dispute, there is only a single correct answer
while all other answers are denitely wrong (and cannot also be
argued for). e SAT dataset contains 374 challenges like this one:
“legend is to map as is: a) subtitle to translation, b) bar to graph, c)
gure to blue-print, d) key to chart, e) footnote to information.” Here,
the correct answer is d) as a key helps to interpret the symbols in a
chart as does the legend with the symbols of a map.
      </p>
      <p>
        While it is easy to see that this answer is correct when the
solution is provided, solving these challenges is a dicult task for
aspiring high school students as the correctness rates of the
analogy section of SAT tests is usually reported to be around 57% [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
An interesting aspect of the SAT analogy test set is that a large
variety of dierent relationship types are covered, and that the
contained challenges have a high degree of variance of diculty
from a humans’ perspective.
      </p>
      <p>Some analogy challenges are harder than others for the average
person, usually based on the rareness of the used vocabulary and the
complexity of the reasoning process required in order to grasp the
intended semantics. Understanding the diculty of challenges is thus
important when judging the eectiveness of an algorithmic solution,
as this provides us with a sense for ”human-level performance”.</p>
      <p>
        However, while the SAT dataset can be obtained easily for
research purposes, there is no publicly available assessment of the
dicultly of dierent challenges. Lo et al examined the
performance of crowd workers recruited from Amazon Mechanical Turk
when faced with SAT challenges [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. ey recruited 99 workers
from English-speaking countries, and each challenge was solved
by 8 crowd workers each (in average, each worker solved 29
challenges.) It turned out that workers are rarely in agreement, and
that most challenges (&gt; 67%) received 3 or 4 dierent answers from
just 8 crowd workers. Only 3:6% challenges are easy enough to
garner unequivocal agreement from all 8 crowd workers, while
most challenges only had an agreement of 4 to 5 workers.
      </p>
      <p>
        In this paper, we use those results to classify each SAT challenge
by their crowd diculty, and make this data publicly available.
is allows us to discuss the performance of analogy algorithms
in comparison to human performance (see section 4). To this end,
we classied each challenge into one of four diculty groups, easy
for challenges which could be handled by most crowd workers
(7-8 correct votes), medium (5-6 correct votes), dicult (3-4 correct
votes), and advanced (0-2 correct votes). Our SAT diculty ratings
can be downloaded on our companion page [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and the resulting
diculty distribution is shown in Figure 1.
3.3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Analogy Gold Standard AGS-1</title>
      <p>
        Lo et al. introduced a benchmark dataset which aims at rectifying
the shortcomings of aforementioned sets, the Analogy Gold
Standard AGS dataset [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In this dataset, each challenge can be used
to benchmark both completion queries, but also ranking queries.
AGS was systematically build using dierent seed datasets like the
SAT dataset or the WordSim-353 dataset [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. One of the core
contribution of AGS was to include human judgments of how “good” an
analogy is from a subjective point of view (as compared to previous
test sets where analogy statements are either correct or wrong).
e quality of an analogy is inuenced by the relational similarity
of the analogons, combined with how well the shared relationships
represent the essence of the analogy (see section 2 for a discussion).
      </p>
      <p>As previous experiments showed that human workers had
problems to individually quantify those aspects, for AGS, a new analogy
rating was introduced which implicitly encodes both relational
similarity and representativeness, i.e. the analogy rating represents
how “useful” and “good” the analogy is perceived by humans on
a scale of 1 to 5. e judgments of at least 8 humans recruited via
crowd-sourcing platforms was used for the AGS analogy ratings of
each analogon pair in a challenge.</p>
      <p>An example AGS challenge is given in Table 1: in this example,
while sombreros are typically considered to come from Mexico,
and in the source analogon sushi typically comes from Japan (i.e.
there is similar relationship between both analogons), the resulting
analogy »sushi; Japan¼ :: »sombrero; Mexico¼ still has a low analogy
rating because the the dening relationship in the source analogon
»sushi; Japan¼ is usually understood by people as “stereotypical
food from a country” - and thus the analogy is deemed not useful
by most humans.
3.4</p>
    </sec>
    <sec id="sec-4">
      <title>Improved Analogy Gold Standard AGS-2</title>
      <p>
        In this paper we introduce the AGS-2 dataset which signicantly
improves on AGS-1 with respect to several aspects. e AGS-2
dataset can be obtained at our companion page [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], thus providing
tangible benets to ongoing analogy-enabled information system
research. Notably, the core improvements of AGS-2 are as follows:
Added diculty rating for each challenge based on the crowd
feedback of 5 workers. e scale is similar to our diculty rating
for SAT challenges introduced in section 3.2: advanced, hard,
      </p>
      <p>medium, easy. e resulting diculty distribution of AGS-2 is
also shown in Figure 1.</p>
      <p>
        Extended size and scope: e initial seed analogons for creating
AGS-1 were extracted from the the Simlex [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and Wordsim
datasets [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], resulting in 93 challenges overall. For AGS-2, this
was extended using suggestions of potentially interesting analogy
pairs by a subset of trusted crowd workers. We manually selected
a subset of these suggestions, resulting in 168 challenges overall.
Improved balance of analogy ratings: In AGS-1, challenges could
be imbalanced with respect to the analogy ratings (i.e., a given
challenge could contain many analogons with high analogy
ratings, but only few with low ratings). For AGS-2 we ensure that
each challenge covers 1-2 analogons each for high, medium, and
low analogy ratings. Note that an analogy with high diculty and
high analogy rating might still not be understandable by many
people (due to being dicult by using rarely known concepts or
complex reasoning), but still will be accepted as a good analogy
by the same people aer the semantics are explained to them.
4
      </p>
    </sec>
    <sec id="sec-5">
      <title>EVALUATION</title>
      <p>
        In this section, we give a preview on how word-embeddings perform
on analogy challenges of varying diculty as previous work only
relied on the aforementioned Mikolov dataset where they showed a
comparably good accuracy of 53.3%. e embedding we evaluated
is Gensim’s skip-gram word2vec [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], trained on a Wikipedia dump.
We use the SAT and AGS-2 datasets, and split the results by the
new diculty measurements introduced in section 3.2.
      </p>
      <p>
        As the SAT dataset only supports analogy ranking queries, for
brevity we only report ranking query results in the following. We
followed the test protocol outlined in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]: we rank each analogon
of a challenge by their relational similarity score with respect to
the source as computed by word2vec, and consider a challenge
successfully solved if the top-ranked analogon is the correct one
(SAT) or has an analogy score higher than a chosen threshold (4.0
in our case for AGS-2).
      </p>
      <p>
        e results are summarized in Figure 2, and are rather
disappointing: for SAT, the average accuracy is 23:4% (random
guessing achieves 20% as each challenge has 5 answer options, average
human-level performance is 57% [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], while one of the current best
performing analogy systems based on both distributional semantics
and structural reasoning using ontologies (like DBpedia or
WordNet) performs at 56.1% [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]). For AGS-2 the word2vec accuracy is
25% (random guessing achieves 16:7%). Interestingly, the
performance of word2vec is consistent with respect to diculty levels,
and it performs worse than humans for easy challenges, but
comparably well for advanced challenges. From this preliminary result we
conclude that the analogy reasoning capabilities of neuronal
embeddings, despite some of their advantages like ease of use and easy
training just relying on text collections, are inferior than current
anecdotal and empirical evidence suggests. However, specialized
analogy reasoning algorithms have been shown to achieve up to
56.1% on SAT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and this has mostly been realized by also
incorporating ontological knowledge (as suggested by [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]). Unfortunately,
this could be challenging for some domain-specic recommender
systems where such ontologies are not easily available, thus
promoting future research to overcome this issue.
      </p>
    </sec>
    <sec id="sec-6">
      <title>SUMMARY AND OUTLOOK</title>
      <p>In this paper, we introduced the challenge of analogy semantics
for recommender systems. Analogies are a natural communication
paradigm to eciently state an information need or to explain a
complex concept. is relies on exploiting the background
knowledge and reasoning capabilities of both communication partners, a
task challenging even for human judges. We introduced the AGS-2
benchmark data set overcoming many shortcomings of previous
datasets, such as being usable for both analogy ranking and analogy
completion queries, and is based on human perception for rating
the quality and diculty of a benchmark analogy. We evaluated the
performance of neuronal word embeddings (using word2vec as a
representative) on previous datasets and our new benchmark
(AGS2). While the method worked worse than human-level performance
for simple queries, it was comparable for more complex queries
with rare vocabulary and requiring extensive common knowledge.</p>
      <p>In our next steps we plan to investigate the performance of hybrid
methods, using both embeddings as well as structural reasoning to
enable analogy queries for recommender systems. is particularly
involves adopting analogy semantics to specic domains like music,
books, or movies - while current analogy systems and benchmark
datasets (including AGS-2) focus on common-knowledge analogies
(which is slightly easier due to the ready availability of both large
text corpora and ontologies).</p>
      <p>In this context, we also plan to evaluate the performance of
analogy based explanations for supporting people in making
decisions about recommended items. is is a particularly interesting
challenge as approaches based only on distributed semantics do
implicitly encode analogy semantics, but have now explicit knowledge
on the type of relationships which would be required to explain
the semantics to a human user, thus again underlining the need for
explicit information on the relationships used in an analogy.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <year>2013</year>
          .
          <article-title>Word2Vec</article-title>
          . hps://code.google.com/archive/p/word2vec/. (
          <year>2013</year>
          ). Accessed:
          <fpage>2017</fpage>
          -06-01.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <fpage>2014</fpage>
          .
          <article-title>Gensim Word2Vec</article-title>
          . hps://radimrehurek.com/gensim/models/word2vec. html. (
          <year>2014</year>
          ). Accessed:
          <fpage>2017</fpage>
          -06-01.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <fpage>2017</fpage>
          .
          <article-title>Analogy Semantics Companion Page</article-title>
          . hps://github.com/WISDel/ analogy semantics. (
          <year>2017</year>
          ). Accessed:
          <fpage>2017</fpage>
          -08-01.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <year>2017</year>
          .
          <article-title>SAT Analogy estions: State of the Art</article-title>
          . hps://www.aclweb.org/aclwiki/ index.php
          <article-title>?title=SAT Analogy estions (State of the art)</article-title>
          .
          <source>(</source>
          <year>2017</year>
          ). Accessed:
          <fpage>2017</fpage>
          -06-01.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Danushka</given-names>
            <surname>Bollegala</surname>
          </string-name>
          , Tomokazu Goto, Nguyen Tuan Duc, and
          <string-name>
            <given-names>Mitsuru</given-names>
            <surname>Ishizuka</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Improving Relational Similarity Measurement using Symmetries in Proportional Word Analogies</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>49</volume>
          ,
          <issue>1</issue>
          (
          <year>2012</year>
          ),
          <fpage>355</fpage>
          -
          <lpage>369</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Lev</given-names>
            <surname>Finkelstein</surname>
          </string-name>
          , Evgeniy Gabrilovich, Yossi Matias, Ehud Rivlin, Zach Solan, Gadi Wolfman, and
          <string-name>
            <given-names>Eytan</given-names>
            <surname>Ruppin</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Placing search in context: the concept revisited</article-title>
          .
          <source>In Int. Conf. on World Wide Web (WWW)</source>
          .
          <source>Hong Kong</source>
          , China.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Bin</given-names>
            <surname>Gao</surname>
          </string-name>
          , Jiang Bian, and
          <string-name>
            <surname>Tie-Yan Liu</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>WordRep: A Benchmark for Research on Learning Word Representations</article-title>
          .
          <source>In ICML Workshop on Knowledge-Powered Deep Learning for Text Mining</source>
          . Beijing, China.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D</given-names>
            <surname>Gentner</surname>
          </string-name>
          .
          <year>1983</year>
          .
          <article-title>Structure-mapping: A theoretical framework for analogy</article-title>
          .
          <source>Cognitive science 7</source>
          (
          <year>1983</year>
          ),
          <fpage>155</fpage>
          -
          <lpage>170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Harris</surname>
          </string-name>
          .
          <year>1954</year>
          .
          <string-name>
            <given-names>Distributional</given-names>
            <surname>Structure</surname>
          </string-name>
          .
          <source>Word</source>
          <volume>10</volume>
          (
          <year>1954</year>
          ),
          <fpage>146</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Reichart</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Korhonen</surname>
          </string-name>
          .
          <year>2014</year>
          . SimLex-999:
          <article-title>Evaluating Semantic Models with (Genuine) Similarity Estimation</article-title>
          .
          <source>Preprint published on arXiv. arXiv:1408:3456</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Douglas</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hofstadter</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Analogy as the Core of Cognition</article-title>
          . In e Analogical Mind.
          <fpage>499</fpage>
          -
          <lpage>538</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Esa</given-names>
            <surname>Itkonen</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Analogy as structure and process: Approaches in linguistics, cognitive psychology and philosophy of science</article-title>
          .
          <source>John Benjamins Pub Co.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Immanuel</given-names>
            <surname>Kant</surname>
          </string-name>
          .
          <volume>1790</volume>
          . Critique of Judgement.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Li</surname>
          </string-name>
          <article-title>man</article-title>
          and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Turney</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>SAT Aanalogy Challange Dataset</article-title>
          . (
          <year>2016</year>
          ). hps://www.aclweb.org/aclwiki/index.php
          <article-title>?title=Analogy (State of the art)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Christoph</surname>
            <given-names>Lo.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Analogy eries in Information Systems: A New Challenge</article-title>
          .
          <source>Journal of Information &amp; Knowledge Management (JIKM) 12</source>
          ,
          <issue>3</issue>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Christoph</surname>
            <given-names>Lo.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Just ask a human? - Controlling ality in Relational Similarity and Analogy Processing using the Crowd</article-title>
          .
          <source>In CDIM Workshop at Database Systems for Business Technology and Web (BTW)</source>
          . Magdeburg, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Christoph</surname>
            <given-names>Lo</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Athiq</surname>
            <given-names>Ahamed</given-names>
          </string-name>
          , Pratima Kulkarni, and Ravi akkar.
          <year>2016</year>
          .
          <article-title>Benchmarking Semantic Capabilities of Analogy erying Algorithms</article-title>
          .
          <source>In Int. Conf. on Database Systems for Advanced Applications (DASFAA)</source>
          . Dallas, TX, USA.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Nieke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Collier</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Discriminating Rhetorical Analogies in Social Media</article-title>
          .
          <source>In Conf. of the Europ</source>
          .
          <article-title>Chapter of the Association for Computational Linguistics (EACL)</article-title>
          . Gothenburg, Sweden.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Kai Chen,
          <source>Greg Corrado, and Jerey Dean</source>
          .
          <year>2013</year>
          .
          <article-title>Ecient estimation of word representations in vector space</article-title>
          .
          <source>International Conference on Learning Representations (ICLR)</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S. Corrado, and Je Dean.
          <year>2013</year>
          .
          <article-title>Distributed Representations of Words and Phrases and their Compositionality</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          <volume>21</volume>
          (
          <year>2013</year>
          ),
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Cameron</given-names>
            <surname>Shelley</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Multiple Analogies In Science And Philosophy</article-title>
          . John Benjamins Pub.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Robert</surname>
            <given-names>Speer</given-names>
          </string-name>
          , Joshua Chin, and
          <string-name>
            <given-names>Catherine</given-names>
            <surname>Havasi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>ConceptNet 5.5: An Open Multilingual Graph of General Knowledge</article-title>
          .
          <source>AAAI Conference on Articial Intelligence</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>P.</given-names>
            <surname>Turney</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          <article-title>man</article-title>
          .
          <year>2005</year>
          .
          <article-title>Corpus-based learning of analogies and semantic relations</article-title>
          .
          <source>Machine Learning</source>
          <volume>60</volume>
          (
          <year>2005</year>
          ),
          <fpage>251</fpage>
          -
          <lpage>278</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>