<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SAFIR: a Semantic-Aware Neural Framework for IR</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Discussion Paper</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maristella Agosti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Marchesin</string-name>
          <email>stefano.marchesin@unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianmaria Silvello</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering, University of Padua</institution>
          ,
          <addr-line>Via Giovanni Gradenigo 6/b, 35131, Padova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The semantic mismatch between query and document terms - i.e., the semantic gap - is a long-standing problem in Information Retrieval (IR). Two main linguistic features related to the semantic gap that can be exploited to improve retrieval are synonymy and polysemy. Recent works integrate knowledge from curated external resources into the learning process of neural language models to reduce the efect of the semantic gap. However, these knowledge-enhanced language models have been used in IR mostly for re-ranking. We propose the Semantic-Aware Neural Framework for IR (SAFIR), an unsupervised knowledge-enhanced neural framework explicitly tailored for IR. SAFIR jointly learns word, concept, and document representations from scratch. The learned representations encode both polysemy and synonymy to address the semantic gap. We investigate SAFIR application in the medical domain, where the semantic gap is prominent and there are many specialized and manually curated knowledge resources. The evaluation on shared test collections for medical retrieval shows the efectiveness of SAFIR to address the semantic gap.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge-enhanced retrieval</kwd>
        <kwd>representation learning</kwd>
        <kwd>semantic gap</kwd>
        <kwd>medical literature</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        This paper addresses the semantic gap, a long-standing problem in Information Retrieval (IR).
The semantic gap refers to the diference between the machine-level description of
document/query contents and the human-level interpretation of their meanings [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It can be
described also as the mismatch between users’ queries and the way retrieval models answer
to such queries. We focus on two linguistic features related to the semantic gap: synonymy
and polysemy. Synonymy occurs when diferent words convey the same meaning, whereas
polysemy occurs when the same word has diferent meanings depending on the context.
      </p>
      <p>
        In the past years, two main lines of work have emerged to bridge the semantic gap between
queries and documents: (i) the use of external knowledge resources to enhance query and
document bag-of-words representations, and (ii) the use of semantic models to perform matching
between the latent representations of queries and documents. Semantic models, which are
based on the Distributional Hypothesis, have been revived by the advent of neural language
models [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Neural language models learn distributed representations of words, also known as
word embeddings, based on the context surrounding words. However, their learning process
      </p>
      <p>
        The full paper has been originally published in ACM Transactions on Information Systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
relies exclusively on text corpora and does not consider any external resources, which encode
factual knowledge that can help to reduce the semantic gap.
      </p>
      <p>To this end, recent works integrate external knowledge into the learning process of neural
language models to reduce the efect of the semantic gap between queries and documents [ 4, 5].
However, even though knowledge-enhanced language models have been proven efective in
many Natural Language Processing (NLP) tasks, their efectiveness is limited in IR [5]. We
identify two reasons causing this performance gap. First, knowledge-enhanced language models
have been used in IR mostly for re-ranking [4, 5]. Secondly, IR tasks are diferent from NLP tasks.
IR requires to match a given query to a set of relevant documents, whereas NLP mostly deals
with the discovery of semantic and linguistic regularities. Therefore, (knowledge-enhanced)
neural language models do not encode relevance signals or discriminative aspects between
queries and documents – which are fundamental to efectively address IR tasks.</p>
      <p>In this work, we investigate which feature between synonymy and polysemy can be exploited
to reduce the semantic gap and improve retrieval, and how external knowledge resources can
help to bridge the semantic gap between queries and documents. To this end, we propose
the Semantic-Aware Neural Framework for IR (SAFIR), an unsupervised knowledge-enhanced
neural framework for IR. SAFIR jointly learns word, concept, and document representations
from scratch. The learned representations are optimized for IR and encode both polysemy and
synonymy to address the semantic gap between queries and documents.</p>
      <p>We conduct an experimental evaluation to compare SAFIR with other (knowledge-enhanced)
neural models on a specific task of Clinical Decision Support ( CDS): medical literature retrieval.</p>
      <p>The rest of the paper is organized as follows: Section 2 presents SAFIR, Section 3 describes
the experimental evaluation, and Section 4 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. The Semantic-Aware Neural Framework for IR</title>
      <p>SAFIR consists of three main components: semantic indexing, representation learning, and
semantic matching. Below, we give an overview of the framework and its main components.
For each component, we outline the required inputs and the provided outputs and we describe
its high-level functioning.</p>
      <p>The semantic indexing component takes as input a corpus and a knowledge resource
and applies Named Entity Recognition (NER) and Entity Linking (EL) techniques to produce a
knowledge-enhanced corpus. For each word, NER detects a list of candidate concepts, if any, and
ranks them from the most to the least likely. Then, EL disambiguates candidate concepts against
the knowledge resource relying on the context of the concept mentions (e.g., the document). The
disambiguated mention-concept pair forms the atomic constituent of each knowledge-enhanced
document.</p>
      <p>The representation learning component consists of a shallow neural network that relies
on the outputs of the the semantic indexing component to learn word, concept, and document
representations. The network models polysemy and synonymy while optimizing representations
for document retrieval via multi-task learning. For polysemy, word and concept representations
are combined to form contextual representations for word-concept pairs, thus conveying a
unique meaning in the vector space. For synonymy, the distance between the representations
CDS15
of words presenting a synonymy relation within the knowledge resource is minimized in the
vector space. Regarding retrieval, contextual representations are learned to be close to the
representations of documents that contain the corresponding word-concept pairs. This entails
a matching relation specific to IR between word, concept, and document representations.</p>
      <p>The semantic matching component uses the learned representations to perform semantic
matching between knowledge-enhanced query and documents. Documents are ranked in
decreasing order of the similarity score computed between query and document representations.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Evaluation</title>
      <p>Experimental Setup. As test collections, we consider TREC Clinical Decision Support 2014
(CDS14), 2015 (CDS15), and 2016 (CDS16).1 As knowledge resource, we adopt the 2018AA
release of the UMLS Metathesaurus2.</p>
      <p>
        We use P@10 and Recall@1000 to evaluate systems and we consider three categories of
retrieval models: bag-of-words models, corpus-driven models, and knowledge-enhanced models.
As Bag-of-Words (BoW) models, we consider BM25 and BM25/RM3. As Corpus-Driven (CD)
models, we consider word2vec [
        <xref ref-type="bibr" rid="ref3">3, 6</xref>
        ] and the Neural Vector Space Model (NVSM) [7]. As
Knowledge-Enhanced (KE) models, we consider retrofitted word2vec ( rword2vec) [4] and three
variants of the Semantic-Aware Neural Framework for IR (SAFIR): SAFIRsp, which integrates
both synonymy and polysemy; SAFIRs which integrates synonymy but not polysemy; and
SAFIRp which integrates polysemy but not synonymy.
      </p>
      <p>Experimental Results We present the experimental results below. Table 1 shows model
performances for medical literature retrieval on the considered collections.</p>
      <p>The experimental results for document retrieval show that all SAFIR variants provide efective
results in the considered collections. This indicates that SAFIR efectively encodes the text
1http://www.trec-cds.org/
2https://www.nlm.nih.gov/research/umls/knowledge_sources/metathesaurus/
matching signals required to perform retrieval regardless of the linguistic feature(s) modeled.
Among the three variants, SAFIRp provides the best results in most cases. Regarding SAFIRs
and SAFIRsp, they exhibit performances close to or slightly lower than those of NVSM and
SAFIRp. This suggests that the impact of synonymy in CDS collections might be limited or
even detrimental. In particular, we expect that modeling polysemy helps to order relevant
documents in top positions of the ranking list, while modeling synonymy helps to retrieve a
higher number of relevant documents which contain synonyms of the query terms. While the
results confirm this trend for polysemy, they do not for synonymy. The negative results of
rword2vec – which models only synonymy – compared to those of word2vec further support
this intuition. Therefore, the results suggest that polysemy impacts more than synonymy on
retrieval performances for CDS collections.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions and Future Work</title>
      <p>We introduced the Semantic-Aware Neural Framework for IR (SAFIR), an unsupervised
knowledge-enhanced neural framework for IR. The evaluation showed that the integration of
knowledge resources into the learning process of neural IR models is efective and helps to bridge
the semantic gap between queries and documents. The learned representations encode text
matching signals, necessary for IR tasks, and linguistic features to retrieve relevant documents
that are most afected by the semantic gap. In particular, the results showed that modeling
polysemy is efective, whereas, on average performance, the impact of synonymy is marginal.</p>
      <p>As future work, we plan to integrate deeper neural architectures into SAFIR representation
learning component to better model linguistic features and their interactions with IR-oriented
objective functions. Two other directions are the extension of SAFIR to phrase-concept
associations and the sensitivity of the learned representations to the NER and EL components.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments References</title>
      <p>The work was partially supported by the ExaMode project, as part of the European Union H2020
program under Grant Agreement no. 825292.
[4] X. Liu, J. Y. Nie, A. Sordoni, Constraining Word Embeddings by Prior Knowledge -
Application to Medical Information Retrieval, in: Proc. of AIRS 2016, Springer, 2016, pp.
155–167.
[5] L. Tamine, L. Soulier, G. H. Nguyen, N. Souf, Ofline Versus Online Representation Learning
of Documents Using External Knowledge, ACM Trans. Inf. Syst. 37 (2019) 42:1–42:34.
[6] I. Vulić, M. F. Moens, Monolingual and Cross-Lingual Information Retrieval Models Based
on (Bilingual) Word Embeddings, in: Proc. of SIGIR 2015, ACM, 2015, pp. 363–372.
[7] C. Van Gysel, M. de Rijke, E. Kanoulas, Neural Vector Spaces for Unsupervised Information
Retrieval, ACM Trans. Inf. Syst. 36 (2018) 38:1–38:25.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Agosti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marchesin</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Silvello, Learning Unsupervised Knowledge-Enhanced Representations to Reduce the Semantic Gap in Information Retrieval</article-title>
          ,
          <source>ACM Trans. Inf. Syst</source>
          .
          <volume>38</volume>
          (
          <year>2020</year>
          )
          <volume>38</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          :
          <fpage>48</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          , G. Zuccon,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bruza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sitbon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lawley</surname>
          </string-name>
          ,
          <article-title>Information retrieval as semantic inference: a Graph Inference model applied to medical search</article-title>
          ,
          <source>Inf. Retr. Journal</source>
          <volume>19</volume>
          (
          <year>2016</year>
          )
          <fpage>6</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Eficient estimation of word representations in vector space</article-title>
          ,
          <source>CoRR abs/1301</source>
          .3781 (
          <year>2013</year>
          ).
          <article-title>a r X i v : 1 3 0 1 . 3 7 8 1</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>