<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Do Social Information Help Book Search?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ludovic Bonnefoy</string-name>
          <email>ludovic.bonnefoy@etd.univ-avignon.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Romain Deveaud</string-name>
          <email>romain.deveaud@univ-avignon.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrice Bellot</string-name>
          <email>patrice.bellot@lsis.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LIA - University of Avignon</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LSIS - Aix-Marseille University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe our participation in the INEX 2012 Book Track. The collection enters its second year of age and is composed of Amazon and LibraryThing entries for real books, and their associated user reviews, ratings and tags. Like in 2011, we tried a simple yet e ective approach of reranking books using a social component that takes into account both popularity and ratings. We did experiments using tags as well.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Previous editions of the INEX Book Track focused on the retrieval of real
outof-copyright books [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These books were written almost a century ago and the
collection consisted of the OCR content of over 50; 000 books. It was a hard
track because of vocabulary and writing style mismatches between the topics
and the books themselves. Information Retrieval systems had di culties to found
relevant information, and assessors had di culties judging the documents.
      </p>
      <p>In 2011, for the books search task, the document collection changed and is
now composed of the Amazon pages of real books. IR systems must now search
through editorial data and user reviews and ratings for each book, instead of
searching through the whole content of the book. The topics were extracted
from the LibraryThing1 forums and represent real requests from real users.</p>
      <p>Like we already did last year, we used a Language Modeling approach to
retrieval. For our recommendation runs, we used the reviews and the ratings
attributed to books by Amazon users. We computed a "social score" for each
book, considering the amount of reviews and the ratings. This score is then used
to modify the initial ranking obtained by a Markov Random Field baseline that
proved to be highly e ective last year. We also used tags to build a pro le for
both a query and the books of the collection which we compared to rank the
books.</p>
      <p>The rest of the paper is organized as follows. The following Section gives an
insight into the document collection whereas Section 2 describes the our retrieval
framework. Finally, we describe our runs in Section 3.</p>
      <sec id="sec-1-1">
        <title>1 http://www.librarything.com/</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Retrieval Model</title>
      <sec id="sec-2-1">
        <title>Sequential Dependence Model</title>
        <p>
          We used a language modeling approach to retrieval [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. We use Metzler and
Croft's Markov Random Field (MRF) model [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to integrate multiword phrases in
the query. Speci cally, we use the Sequential Dependance Model (SDM), which is
a special case of the MRF. In this model three features are considered: single term
features (standard unigram language model features, fT ), exact phrase features
(words appearing in sequence, fO) and unordered window features (require words
to be close together, but not necessarily in an exact sequence order, fU ).
        </p>
        <p>Finally, documents are ranked according to the following scoring function:
scoreSDM (Q; D) = T</p>
        <p>X fT (q; D)
q2Q
jQj 1
+ O X fO(qi; qi+1; D)
i=1
jQj 1
+ U X fU (qi; qi+1; D)
i=1
where the features weights are set according to the author's recommendation
( T = 0:85, O = 0:1, U = 0:05). fT , fO and fU are the log maximum likelihood
estimates of query terms in document D, computed over the target collection
with a Dirichlet smoothing.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Modeling book likeliness</title>
        <p>The basic idea behind this likeliness is that if a book has a lot of reviews and if
its ratings are generally good, then it must be a very good book.</p>
        <p>L(D) = log(#reviews(D))</p>
        <p>Pr2RD r
#reviews(D)
where RD is the set of all ratings given by the users for the book D, and
#reviews(D) is the number of reviews.</p>
        <p>We further rerank the books by weighting the previously computed SDM
with the likeliness score. The scoring function of a book D given a query Q is
thus de ned as follows:</p>
        <p>s(Q; D) = L(D) scoreSDM (Q; D)
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Modeling book thematic relatedness</title>
        <p>We want to represent each query Q by a thematic pro le and rank books
according to their relatedness to it. For this rst attempt at using thematic (or
topic) relatedness we choosed to rely exclusively on user tags associated with the
books in the collection. We consider as a thematic pro le a set of tags weighted
according to their signi cance for Q and we call it a tag pro le (TP). As a
preprocessing step, a tag pro le is associated to each book in the collection. Tags
are weighted according to a classic tf.idf measure (where the tf is the number of
users who associated the tag to the book).</p>
        <p>The main issue is to estimate a tag pro le for a query. To construct it, inspired
by the pseudo relevance feedback method, we summed the pro le of the x top
ranked books retrieved by mean of a information retrieval model (more details
in runs section). Once the query's tag pro le is build, we can compare book's
tag pro le to it with a vector similarity measure like the cosinus.</p>
        <p>Finally, books of the collection are ranked according to the similarity of their
pro le to the query's one.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Runs</title>
      <p>We submitted 4 runs for the Social Search for Best Books task. We used Indri2 for
indexing and searching. We did not remove any stopword and used the standard
Krovetz stemmer.
mrf-booklike This run is the implementation of the SDM model described in
Section 2.1 with the likeliness score.</p>
      <p>IOT30 and IT30 Those two runs are based on the tag pro le approach
presented in Section 2.3. In this approach four parameters have to be xed : The
number x of top ranked books used to build the query's tag pro le, the weight
given to each tag in query's pro le, the information retrieval model used to
retrieved books and the similarity measure to compare pro les. For both runs,
x is xed to 30, Indri's language modeling approach is used and the similarity
measure is the cosinus angle between vectors.</p>
      <p>The last parameter is the weight given to each tag of the query pro le. For
the IOT30 run, the ti tag's weight is compute as the sum of its tf.idf weight in
each of the top x books returned by Indri:
w(ti) =</p>
      <p>X tf:idf (ti; b)
b2T opx
where b is one the T opx books retrieved.</p>
      <p>However, we had the intution that all selected books can not contribute
equally to the weight of a tag. So, for the IT30 run, we combine the tf.idf of a
tag in a book with the relevance of this book according to the retrieval model
used in order to penalize contribution of less relevant books:
w(ti) =</p>
      <p>X tf:idf (ti; b) score(b; Q)
b2T opx</p>
      <sec id="sec-3-1">
        <title>2 http://www.lemurproject.org</title>
        <p>where score(b; Q) is the measure of relevance of the book b according to Indri.</p>
        <p>deduce
B IT30 30 For the last run we wanted to take advantage of both particularities
of mrf-booklike run and a tag pro le based one. We combine the mrf-booklike
run to the IT30 run by mean of a logistic regression. We trained a model with
two classes (relevant or not) and the book scores predicted by both runs as
features. Training instances were the 30 top ranked books returned by each run
along with their relevance judgment deduce from 2011 qrels.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>Recip rank
Recall@10
0.2282
0.2071
0.2105
0.2081
0.1890
0.3069
0.3410
0.3584
0.2933
0.2999
0.5398
0.4811
0.4626
0.4524
0.4337
0.1527
0.1659
0.1514
0.1503
0.1426</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this paper we presented our contributions for the INEX 2012 Book Track.
We proposed a simple method for reranking books based on their likeliness and
an e ective way to take into account user tags. Finally a combination of both
methods with a logistic regression approach gives the best results. Results does
not allow us to answer on the usefulness of social information for book search
despite quite good results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Gabriella</given-names>
            <surname>Kazai</surname>
          </string-name>
          , Marijn Koolen, Antoine Doucet, and
          <string-name>
            <given-names>Monica</given-names>
            <surname>Landoni</surname>
          </string-name>
          .
          <article-title>Overview of the INEX 2010 Book Track: At the Mercy of Crowdsourcing</article-title>
          . In Shlomo Geva, Jaap Kamps, Ralf Schenkel, and Andrew Trotman, editors,
          <source>Comparative Evaluation of Focused Retrieval</source>
          , pages
          <volume>98</volume>
          {
          <fpage>117</fpage>
          . Springer Berlin / Heidelberg,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D.</given-names>
            <surname>Metzler</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>Combining the language model and inference network approaches to retrieval</article-title>
          . Inf. Process. Manage.,
          <volume>40</volume>
          :
          <fpage>735</fpage>
          {
          <fpage>750</fpage>
          ,
          <year>September 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Donald</given-names>
            <surname>Metzler</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. Bruce</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>A markov random eld model for term dependencies</article-title>
          .
          <source>In Proceedings of the 28th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          ,
          <source>SIGIR '05</source>
          , pages
          <fpage>472</fpage>
          {
          <fpage>479</fpage>
          , New York, NY, USA,
          <year>2005</year>
          . ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>