<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SBS 2016 : Combining Query Expansion Result and Books Information Score for Book Recommendation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Amal Htait</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sebastien Fournier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrice Bellot</string-name>
          <email>patrice.bellotg@univ-amu.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aix Marseille Universite</institution>
          ,
          <addr-line>CNRS, ENSAM, Toulon Universite, LSIS UMR 7296,13397, Marseille</addr-line>
          ,
          <country country="FR">France.</country>
          <institution>Aix-Marseille Universite</institution>
          ,
          <addr-line>CNRS, CLEO OpenEdition UMS 3287, 13451, Marseille</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we present our contribution in Suggestion Track at the Social Book Search Lab. This track aims to develop test collections for evaluating ranking e ectiveness of book retrieval and recommender systems. In our experiments, we combine the results of Sequential Dependence Model (SDM) and the books information that includes the price, the number Of P ages and the publication Date. We also expand topics' queries by the similar books information to improve the recommendation performance.</p>
      </abstract>
      <kwd-group>
        <kwd>Social Information Retrieval</kwd>
        <kwd>Recommendation</kwd>
        <kwd>Sequential Dependence Model</kwd>
        <kwd>Expand Query</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The Social Book Search (SBS) Tracks [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] were introduced by INEX in 2010 with
evaluation purposes for supporting users in searching collections of books based
on book metadata and associated user-generated content.
      </p>
      <p>Social Book Search Lab includes the following tracks: Suggestion Track,
Interactive Track and Mining Track. Our work is on Suggestion Track, which suggests
a list of the most relevant books according to the request provided by the user.
Since 2011, for the social books search task, the document provided is a
collection of 2.8 million records containing professional metadata (Amazon1) extended
with user-generated content and social metadata (LibraryThing2). In addition,
a set of 113,490 anonymous users pro les is provided from LibraryThing (LT).
Therefore, Information Retrieval (IR) Systems must search through editorial
data, user reviews and ratings for each book, instead of searching through the
whole content of the book. The topics provided each year are extracted from the
LibraryThing forums and by represent real requests from real users.</p>
    </sec>
    <sec id="sec-2">
      <title>1 http://www.amazon.com/</title>
    </sec>
    <sec id="sec-3">
      <title>2 www.librarything.com</title>
      <p>
        Our participation in 2011 and 2012 was based on re-ranking books using
social component such as popularity and ratings [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. On 2014, we were able
to achieve the second best run using InL2 model implemented in Terrier3[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
And for 2015 participation, we combined results of InL2 and Sequential
Dependence Model (SDM). Also, we integrated tools from natural language processing
(NLP) and approaches based on graph analysis to improve the recommendation
performance[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>This year's participation is through an IR system based on 3 main steps:
{ We expand the topic queries using the similar books information, since the
topics contain books titles mentioned by the user as similar or example books
to those he seeks.
{ We apply a re-ranking method using a score calculated of books information
including the price, the number Of P ages and the publication Date.
{ We apply these methods on Amazon book collection and on the users pro les
collection.</p>
      <p>For our participation in SBS 2016, we submitted 4 runs in which we applied
the previously mentioned steps. The rest of this paper is organized as follows.
The following section describes the data processing and indexing. In section 3,
we have the description of our retrieval framework. In section 4, we describe the
submitted runs. Finally, we present the obtained results in section 5.
2</p>
      <sec id="sec-3-1">
        <title>Data processing and indexing</title>
        <p>We use, in addition to the Amazon book Collection, the users pro les
Collection provided by SBS Lab track which contains the cataloguing transactions
of 113,490 users. The cataloguing transactions of a user is a list of
information concerning the books read by the user. Each transaction is represented by
a row, where each row contains eight columns; user, book, author, book title,
publication year, month in which the user added that book, rating and a set of
tags assigned by this user to this book. From the users pro les, we create for
each book an XML le with all its information. An example is illustrated in the
following XML code of Figure 1.</p>
        <p>For indexing the Amazon book collection, we take all the tags of the XML
les identi ed by the ISBNs. And for indexing users pro les collection, we take
all the tags of the created XML les identi ed by the LibraryThingID. Also, we
use the following Indri4 indexing parameters: P orter Stemmer and Stop W ords
Removal.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3 hhtp://terrier.org</title>
    </sec>
    <sec id="sec-5">
      <title>4 http://www.lemurproject.org/indri/</title>
      <p>3</p>
      <sec id="sec-5-1">
        <title>Retrieval Model</title>
        <p>Query Expansion by example books information
To build our queries we use mainly the title of the query and the information
of similar example books mentioned by the user in the topic. Also, we use the
tags of these similar books extracted from the users pro les collection for query
expansion. The XML code in Figure 2 illustrates an example of adding similar
book tags for query expansion.
3.2</p>
        <p>
          Sequential Dependence Model
SDM relies on the idea of integrating multi word phrases by considering a
combination of query terms with proximity constraints such as: single term features
(standard unigram language model features, fT ), exact phrase features (words
appearing in sequence, fO) and unordered window features (require words to be
close together, but not necessarily in an exact sequence order, fU ) [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. In Table 1,
more details about the term weighting functions are shown, where tfe;D is the
number of times term e matches in document D, cfe;D is the number of times
term e matches in the entire collection, jDj is the length of document D, and jCj
is the size of the collection. Finally, is a weighting function hyperparameter
that is set in our work to 2500 [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
fO(qi; qi+1; D) = log tf#1(qi;qi+1);jDD+j+
fU (qi; qi+1; D) = log tf#uw8(qi;qi+1);jDD+j+
cf#1(qi;qi+1)
        </p>
        <p>jCj
cf#uw8(qi;qi+1)
jCj</p>
        <sec id="sec-5-1-1">
          <title>Description</title>
        </sec>
        <sec id="sec-5-1-2">
          <title>Weight of unigram qi</title>
          <p>in document D.</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>Weight of exact phrase</title>
          <p>'qi qi+1' in document D.</p>
        </sec>
        <sec id="sec-5-1-4">
          <title>Weight of unordered</title>
          <p>window 'qi qi+1'
(span=8) in document D.</p>
          <p>And the documents are ranked according to the below scoring equation,
Equation 1:</p>
          <p>SDM (Q; D) =</p>
          <p>T Pq2Q fT (q; D)
+ O PjiQ=j1 1 fO(qi; qi + 1; D)
+ U PjiQ=j1 1 fU (qi; qi + 1; D)
(1)</p>
          <p>We used the Equation 1 with feature weights set to T = 0.85, O = 0.1 and
U = 0.05, like previous participation years. We applied this model to the queries
using Indri 5.4 4 Query Language 5. An example of Indri Query Language is in
Figure 3.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5 http://www.lemurproject.org/lemur/IndriQueryLanguage.php</title>
      <p>
        Combination of Retrieval System output and books' information
We combine the results of SDM model with a sum of normalized scores, which
we calculate from the book's price, publication Date and number Of P ages.
And since the combined values are of di erent weighting, we use the maximum
and minimum scores according to Lees formula [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] as followed in Equation 2.
normalizedScore =
oldScore
maxScore
minScore
minScore
(2)
      </p>
      <p>The scores of SDM model and books information have di erent levels of
retrieval e ectiveness, thus it is necessary to weigh scores depending on their
overall performance. We used an interpolation parameter ( ) that varies in
testing for the goal of achieving the best interpolation that provides better retrieval
e ectiveness, as shown in the Equation 3.</p>
      <p>SDM bookInf o = :(SDM (Q; D)) + (1
):(bookInf o(D))
(3)</p>
      <p>After several testings on 2015 SBS topics 6, is set to 0.55 with the best
result. bookInf o(D) is calculated by a normalized score of the values of price
only, since the price alone obtains the best result on 2015 SBS topics compared
to the values of price, publication Date and number Of P ages combined. In
Table 2, an example of our tests showing a modest but still an increase in the
results when combining books prices to the equation with = 0.55.</p>
      <sec id="sec-6-1">
        <title>Runs</title>
        <p>We submit 4 runs for the SBS Suggestion Track:
Run1 ExeOrNarrativeNSW Collection: We concatenate the title of the
topic and the similar books elds (title, author and tags), then perform a
retrieval using the SDM model. But since not all topics have example books, in this
case we concatenate the title and the narrative elds of the topic after removing
the Stop W ords from the narrative eld. This run is applied on Amazon book
collection.</p>
        <p>Run2 ExeOrNarrativeNSW UserPro le : This run is same as Run1 but it
is applied on users pro les collection.</p>
        <p>Run3 ExeOrNarrativeNSW Collection AddData : In this run, we
combine the books price normalized score to the results of Run1.</p>
        <p>Run4 ExeOrNarrativeNSW UserPro le AddData : Also in this run, we
combine the books price normalized score to the results of Run2.
5</p>
      </sec>
      <sec id="sec-6-2">
        <title>Results</title>
      </sec>
      <sec id="sec-6-3">
        <title>Conclusion</title>
        <p>In this paper, we present our contribution for the Suggestion Track of Social
Book Search Lab. In the 4 submit runs, we use SDM retrieval model and we
extend the query by the similar books information (title, author and tags). We
apply the retrieval on Amazon book collection, and on users pro les collection.
We combine the results of the retrieval system (SDM) with the normalized score
of the books prices. The best result is achieved by using SDM retrieval model
with the extended query on Amazon book Collection. We should note that the
topics of SBS 2015 had a eld named mediated query, which contained the key
words of the user's request (from eld narrative). The mediatedquery eld is
used in our testing on SBS 2015 topics and helped to increase the results. But
since this eld is not in the topics of SBS 2016, we had to use the narrative eld
which contains many useless information that e ect negatively the information
research. Thus, to increase the results for future participation, we must work on
extracting only the key words from the narrative eld to be used in the query,
and eliminate any noise information.</p>
        <p>Acknowledgments. This work was supported by the French program
Investissements dAvenir Equipex "A digital library for open humanities" of
OpenEdition.org.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Gabriella</given-names>
            <surname>Kazai</surname>
          </string-name>
          , Marijn Koolen, Jaap Kamps, Antoine Doucet, and
          <string-name>
            <given-names>Monica</given-names>
            <surname>Landoni</surname>
          </string-name>
          .
          <article-title>Overview of the inex 2010 book track: Scaling up the evaluation using crowdsourcing</article-title>
          .
          <source>In Shlomo Geva</source>
          , Jaap Kamps, Ralf Schenkel, and Andrew Trotman, editors,
          <source>INEX</source>
          , volume
          <volume>6932</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>98117</fpage>
          . Springer,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Deveaud</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , SanJuan,
          <string-name>
            <given-names>E.</given-names>
            , &amp;
            <surname>Bellot</surname>
          </string-name>
          ,
          <string-name>
            <surname>P..</surname>
          </string-name>
          <article-title>Social recommendation and external resources for book search</article-title>
          .
          <source>Working Notes for CLEF 2011 Conference, 7424 LNCS</source>
          ,
          <year>6879</year>
          .
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Ludovic</given-names>
            <surname>Bonnefoy</surname>
          </string-name>
          , Romain Deveaud, and
          <string-name>
            <given-names>Patrice</given-names>
            <surname>Bellot</surname>
          </string-name>
          .
          <article-title>Do social information help book search</article-title>
          ? In Pamela Forner, Jussi Karlgren, and
          <string-name>
            <surname>Christa</surname>
          </string-name>
          Womser-Hacker, editors,
          <source>CLEF (Online Working Notes/Labs/Workshop)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Benkoussas</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamdan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albitar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ollagnier</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bellot</surname>
            ,
            <given-names>P. .</given-names>
          </string-name>
          <article-title>Collaborative Filtering for Book Recommendation</article-title>
          .
          <source>Working Notes for CLEF 2014 Conference</source>
          ,
          <volume>501507</volume>
          . (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Benkoussas</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ollagnier</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bellot</surname>
            ,
            <given-names>P..</given-names>
          </string-name>
          <article-title>Book Recommendation Using Information Retrieval Methods and Graph Analysis</article-title>
          .
          <source>Working Notes for CLEF 2015 Conference</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Donald</given-names>
            <surname>Metzler</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. Bruce</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>A markov random eld model for term dependencies</article-title>
          . In Ricardo A.
          <string-name>
            <surname>Baeza-Yates</surname>
          </string-name>
          , Nivio Ziviani, Gary Marchionini, Alistair Mo at, and John Tait, editors,
          <source>SIGIR</source>
          , pages
          <fpage>472479</fpage>
          . ACM,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Joon</given-names>
            <surname>Ho</surname>
          </string-name>
          <article-title>Lee</article-title>
          .
          <article-title>Combining multiple evidence from di erent properties of weighting schemes</article-title>
          .
          <source>In Proceedings of the 18th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 95</source>
          , pages
          <fpage>180188</fpage>
          , New York, NY, USA,
          <year>1995</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Benkoussas</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ollagnier</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bellot</surname>
            ,
            <given-names>P..</given-names>
          </string-name>
          <article-title>Book Recommendation Using Information Retrieval Methods and Graph Analysis</article-title>
          ,
          <source>CLEF 2015 Conference and Labs of the Evaluation Forum</source>
          , pp.
          <volume>8</volume>
          p.,
          <string-name>
            <surname>Toulouse</surname>
          </string-name>
          (France), sep
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>