<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Facts That Matter: Dynamic Fact Retrieval for Entity-Centric Search Queries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Atharva Prabhat Paranjpe</string-name>
          <email>atharva.paranjpe@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rajarshi Bhowmik</string-name>
          <email>rajarshi.bhowmik@rutgers.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gerard de Melo</string-name>
          <email>gdm@demelo.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hasso Plattner Institute, University of Potsdam</institution>
          ,
          <addr-line>Potsdam</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>New Brunswick</institution>
          ,
          <addr-line>Piscataway, NJ</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Rutgers University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Entity-centric queries constitute a signi cant proportion of all search queries processed by the popular search engines. Answering such queries often involves selecting facts pertaining to an entity from an underlying knowledge graph. Prior work on this draws on hand-crafted features that require scanning the entire knowledge graph beforehand. Instead, we propose a neural method that exploits the linguistic and semi-linguistic nature of the entity search queries and the facts, and can hence be applied dynamically to entirely new sets of candidate facts. We optimize our model using a pairwise loss function to correctly predict the relevance and importance scores for each fact for a given query entity, while the overall fact ranking is based on a linear combination of these scores. We show that our simple approach outperforms previous work, ensuring better fact retrieval for entity-centric search queries.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In recent years, knowledge cards have become an integral part of popular Web
search engines. Knowledge cards are information boxes that appear on the search
engine result pages when a user searches for entity-related information [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Such
cards provide a series of facts taken from a knowledge graph [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and enable the
user get a brief overview of pertinent key facts about the entity without the need
to navigate to various individual web pages.
      </p>
      <p>
        In practice, di erent entity-related queries may pertain to quite di erent
aspects of an entity. A search engine query such as \einstein education" ought
to give preference to other facts than a query such as \einstein family ". To
address this task of dynamic query-speci c fact ranking, Hasibi et al. proposed
a model called DynES [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] that performs fact retrieval and entity summarization
based on a linear combination of two measures: importance and relevance, and
compared the results to human judgments.
      </p>
      <p>However, DyNES is based on hand-crafted features that are cumbersome to
compute, as they need to be extracted beforehand from the set of all facts in
the large-scale knowledge graph, rendering this method unsuitable for ad hoc
1 Copyright c 2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
settings. Moreover, DynES performs a simple pointwise ranking of the facts,
where each fact is considered in isolation, using Gradient Boosted Regression
Trees, which learn an ensemble of weak prediction models.</p>
      <p>
        In contrast, we propose a novel model that obviates the need for a
cumbersome process of extracting hand-crafted features from the large knowledge graph.
Our key contributions are as follows. (1) We propose a deep neural model with
a pairwise loss function to address the task of query-dependent fact retrieval for
entity-centric search queries. (2) Rather than depending on a large knowledge
graph for feature extraction, our model draws on recent advances in
Transformers with self attention [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] to better model the linguistic connection between the
query and the candidate facts, and thus can be applied even to entirely novel
sets of candidate facts. (3) We conduct a set of experimental evaluations showing
that our approach outperforms previous work.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Preliminaries</title>
      <p>In the following, we de ne relevant terminology that is used in the remainder
of the paper. We consider a fact f as a predicate{object pair returned when a
query is made with regard to an entity, with that entity serving as the subject.
De nition 1. (Importance) Importance is an attribute of a fact f that
determines its relation to the subject entity s in absolute terms, irrespective of the
provided query. It is denoted as is(f ).</p>
      <p>De nition 2. (Relevance) Relevance, in turn, describes to what extent a given
candidate fact f is pertinent with regard to a given natural language search query
q issued by the user along with the entity s as the subject. It is denoted as rs;q(f ).
De nition 3. (Utility) The overall utility of a fact f with respect to a query q
and entity s is de ned as a weighted sum of the importance and relevance scores
of the fact with respect to query and entity. It is denoted as us;q(f ) and computed
as us;q(f ) = is(f ) + rs;q(f )</p>
      <p>The weights , may be adjusted freely to account for application
scenariospeci c considerations. Thus, utility relates the fact to the query in a more
comprehensive manner than the importance and relevance scores alone can.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Model</title>
      <p>
        Given the natural language input query Q as well as a candidate fact fi =
hp; oi 2 F , where F is the set of all candidate facts for Q, our model accepts the
query along with the natural language labels of p and o and invokes BERT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], a
deep neural Transformer encoder, to encode bidirectional contextual information
for the given sequence of input tokens. Since we simultaneously supply both the
query and the candidate fact to the Transformer, the self-attention layers are
able to establish connections between (parts of) these two inputs.
      </p>
      <p>Before feeding the tokenized sequences Q, P , O to the model, special tokens
[CLS] and [SEP] are inserted into the input sequence. The [CLS] token signi es
the start of each sequence, while the [SEP] token serves as a demarcation point
separating the query segment from the fact segment in the input sequence. The
resulting sequence of input token identi ers now becomes:
[CLS], q1; : : : ; ql, [SEP], p1; : : : ; pm, o1; : : : ; on, [SEP]</p>
      <p>The most relevant component for our task lies in the encoded representation
of the [CLS] token, which serves as a representation of the entire input sequence.
This representation is passed through a fully-connected layer followed by a
sigmoid activation function to yield a ranking score, which is then compared to
the ground truth. Formally, g(fi) = (Whp + b), where hp 2 Rd is the [CLS]
representation from the nal hidden layer of the BERT encoder, W 2 R1 d and
b are trainable parameters, and (x) = 1+1e x .</p>
      <p>The model is trained to minimize a pairwise ranking loss that considers pairs
of facts (fs and fi) and encourages the model to predict scores for the two
involved facts that re ect the correct relative ordering between them. Ideally,
the di erence between the two predicted scores (g(fs) g(fi)) should equal the
di erence (r(s) r(i)) between the corresponding ground truth ranking scores.
These di erences are computed as signed values rather than absolute values, so
the ordering is crucial. Based on this intuition, we de ne the loss function as a
pairwise mean squared error as follows:</p>
      <p>h r(s) r(i) g(fs) g(fi) i2
L(g; F ; R) = Psn=11 Prn(i)i=&lt;1r(s)</p>
      <p>The nal ranking is created by ordering the candidate facts fi 2 F by g(fi)
in descending order, breaking ties arbitrarily. Thus, if g(fi) &gt; g(fj ), then fi
should be ranked higher than fj .
4</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>
        We perform our experiments on two dataset variants put forth by Hasibi et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
For the rst variant, Complete Dataset, the entire collected data is considered.
This data consists of 100 English language queries, 4,069 facts, and 41 facts
per query on average. Their second variant, URI-only Dataset, keeps only the
subset of facts for which the objects are genuine entities identi ed by a URI,
while facts with literal values are omitted. It contains the same 100 queries,
1,309 facts, and 14 facts per query on average.
      </p>
      <p>
        We compare several di erent models, including DynES [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], a BiLSTM Dual
Encoder, a BERTBASE variant of our model that does not ne-tune the BERT
encoder, a pointwise score prediction variant of our model, and our pairwise
model. For reproducibility and future research, we release the source code of
our model2. We use the standard NDCG metric for evaluation with ranked lists
of length 5 (NDCG@5) and 10 (NDCG@10), and report the evaluation scores
obtained using 5-fold cross validation.
      </p>
      <p>
        Model
RELIN [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] y 0.6300 0.7066 0.6368 0.7130 N/A N/A
LinkSum [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] y 0.6504 0.6648 0.7018 0.7031 N/A N/A
SUMMARUM [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] y 0.6719 0.7111 0.7181 0.7412 N/A N/A
DynES [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] y 0.8164 0.8569 0.8291 0.8652 N/A N/A
Bi-LSTM Dual Encoder 0.6416 0.7225 0.6821 0.7508 N/A N/A
BERTBASE 0.7055 0.7675 0.6521 0.7274 0.4563 0.5498
Our Model (pointwise) 0.7850 0.8285 0.8635 0.8821 0.6165 0.6741
Our Model (pairwise) 0.8515 0.8761 0.8454 0.8743 0.6621 0.7269
Table 2. 5-fold cross-validation results on the URI-only Dataset. y: results taken from
Hasibi et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], who did not report separate relevance prediction results apart from
the overall utility prediction results.
      </p>
      <p>The results of our experiments on the Complete Dataset and URI-only
Dataset are given in Tables 1 and 2, respectively. Our model outperforms the
DynES model with absolute gains of 4:9% and 13:2% in terms of the NDCG@10
metric on the Complete Dataset for the utility and importance-based
rankings (as de ned in Section 2), respectively. We observe a similar trend for
2 https://github.com/AtharvaParanjpe/Dynamic-Fact-Ranking-For-Entity-Centric-Queries
the URI-only Dataset, where our model consistently outperforms the DynES
model, with respective absolute gains of 4:3% and 2:0% in the NDCG@5 metric
for utility and importance rankings.</p>
      <p>The moderate performance of the pre-trained BERTBASE model suggests
that BERTBASE already has su cient linguistic information embedded in it to
be able to rank the facts to a certain degree. In fact, without any ne-tuning,
BERTBASE outperforms the BiLSTM Dual Encoder baseline in most of the cases.</p>
      <p>Our model outperforms the pointwise ranking variant with as high as 8:5%
and 5:7% absolute gain in NDCG@5 and NDCG@10 metrics for the utility scores
on the URI-only Dataset. We conjecture that this is because a pairwise loss
function allows the model to better assess the di erences between di erent facts
and because this training regime better exploits the available training data.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we propose a new neural method to learn entity-centric fact
rankings, accounting for both the saliency and the relevance of facts with regard
to the query. Our method adopts a pairwise ranking approach while drawing
on state-of-the-art deep neural modeling techniques to analyze the semantics of
queries and candidate facts along with their semantic connections. Unlike
previous work, it can dynamically be applied to entirely new candidate facts without
the need to compile knowledge graph statistics. In our experimental evaluation,
we observe substantial improvements over previous work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bhowmik</surname>
          </string-name>
          , R., de Melo, G.:
          <article-title>Generating ne-grained open vocabulary entity type descriptions</article-title>
          .
          <source>In: Proceedings of ACL</source>
          <year>2018</year>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bhowmik</surname>
          </string-name>
          , R., de Melo, G.:
          <article-title>Be concise and precise: Synthesizing open-domain entity descriptions from facts</article-title>
          .
          <source>In: Proceedings of The Web Conference</source>
          <year>2019</year>
          . ACM (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Cheng, G.,
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>RELIN: Relatedness and informativeness-based centrality for entity summarization</article-title>
          .
          <source>In: Proceedings of ISWC 2011</source>
          . Springer (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          : BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In: Proc. of NAACL</source>
          <year>2019</year>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hasibi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balog</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bratsberg</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          :
          <article-title>Dynamic factual summaries for entity cards</article-title>
          .
          <source>In: Proceedings of SIGIR 2017</source>
          . pp.
          <volume>773</volume>
          {
          <fpage>782</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hogan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomqvist</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cochez</surname>
          </string-name>
          , M.,
          <string-name>
            <surname>d'Amato</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Melo</surname>
          </string-name>
          , G.,
          <string-name>
            <surname>Gutierrez</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Labra</surname>
            <given-names>Gayo</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.E.</given-names>
            ,
            <surname>Kirrane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Neumaier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Polleres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Navigli</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ngonga</given-names>
            <surname>Ngomo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.C.</given-names>
            ,
            <surname>Rashid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.M.</given-names>
            ,
            <surname>Rula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Schmelzeisen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Sequeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Staab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Zimmermann</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Knowledge graphs</article-title>
          .
          <source>ArXiv</source>
          <year>2003</year>
          .
          <volume>02320</volume>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Thalhammer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lasierra</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rettinger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>LinkSUM: Using Link Analysis to Summarize Entity Data</article-title>
          .
          <source>In: Proceedings of ICWE 2016</source>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Thalhammer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rettinger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Browsing DBpedia entities with summaries</article-title>
          .
          <source>In: ESWC (Satellite Events)</source>
          .
          <source>LNCS</source>
          , vol.
          <volume>8798</volume>
          , pp.
          <volume>511</volume>
          {
          <fpage>515</fpage>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          , pp.
          <volume>5998</volume>
          {
          <fpage>6008</fpage>
          . Curran Associates (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>