<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LIG at CLEF 2015 SBS Lab</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nawal Ould-Amer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philippe Mulhem</string-name>
          <email>Philippe.Mulhemg@imag.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mathias Gery</string-name>
          <email>mathias.gery@univ-st-etienne.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Nawal.Ould-Amer</institution>
          ,
          <addr-line>Philippe.Mulhem</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Univ. Grenoble Alpes, LIG</institution>
          ,
          <addr-line>F-38000 Grenoble</addr-line>
          ,
          <country country="FR">France</country>
          <addr-line>CNRS, LIG, F-38000 Grenoble</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universite de Lyon</institution>
          ,
          <addr-line>F-42023, Saint-Etienne</addr-line>
          ,
          <country country="FR">France</country>
          ,
          <institution>CNRS, UMR 5516, Laboratoire Hubert Curien</institution>
          ,
          <addr-line>F-42000, Saint-Etienne</addr-line>
          ,
          <institution>France Universite de Saint-Etienne</institution>
          ,
          <addr-line>Jean-Monnet, F-42000, Saint-Etienne</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the work achieved by the MRIM research group of Grenoble, using some data from the LaHC of SaintEtienne, in a way to test personalized retrieval of books for the Social Book Search Lab of CLEF 2015. Our proposal rely on a biased fusion of content-only retrieval, using BM25F and LGD retrieval models, user non-social pro le based on the catalog of the requester, and social proles using user/user links generated from their catalogs and ratings on books. The o cial results obtained show a clear positive impact of user pro le, and a small positive impact of the social elements we used. Post o cial results that present non biased fusion scores are also presented.</p>
      </abstract>
      <kwd-group>
        <kwd>Fusion of scores</kwd>
        <kwd>user pro le</kwd>
        <kwd>social links</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This paper describes our participation to INEX Social Book Search Suggestion
Track challenge. The goal of this challenge is to evaluate approaches for
supporting users in searching collections of books based on book metadata and
associated user-generated content [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The work described here focuses on
several aspects of personalized information retrieval that integrates social networks
information. Our objectives during the participation was twofold: a) to rely as
much as possible on Information Retrieval Systems to handle non-social and
social pro les, and b) to provide a simple integration of the three elements (i.e.,
content, non-social pro les, social pro les) according to linear combination of
scores. Relying heavily on existing tested IR tools allows us to focus on
experimenting ideas. Proposing simple score fusions allows us to analyze more easily
how con gurations behave. More precisely, our experiments conducted for SBS
2015 emphasizes on:
{ Studying the impact of using a simple user pro le as query extension
(nonsocial pro le);
{ Studying the impact of generated friend relations on the quality of the results
(social pro le).
      </p>
      <p>From the data provided by the SBS 2015 dataset, the following elements were
used at one time or another:
{ the elds title, summary, content and tags from the documents: all
concatenated for unstructured retrieval, and separated for eld-based retrieval using
BM25F;
{ the elds title, mediated query and narrative from the topics;
{ the documents and ratings from the \topic users" (a topic user is the
description of the user that asks a query): used to compute \friendship"
relationships between users;
{ the documents and ratings from the pro les of the non-topics users: used to
compute \friendship" relationships between users.</p>
      <p>
        The IR processes were achieved on the Terrier system3 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>The section 2 focuses on the description of the fusion that was exploited: one
original point relates to the biases that we propose. Section 3 tackles multiple
content-only matching for the documents, as we found out that such integration
is bene cial. Then, we introduce in section 4 the use of non-social pro les, and we
detail how we de ned friendship relations between users, using their catalogs and
ratings, as well a the way we used the pro les of such friends when processing
queries. Additional processing must be achieved on SBS data to get results.
We discuss in section 5 some of these elements before depicting in section 6
the o cial, as well as some non-o cial, results obtained, before concluding in
section 7.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Biased linear fusions of scores</title>
      <p>
        Fusing scores of several IR systems is nontrivial problem. In our case, as described
in the introduction, we propose to use biased linear fusions of scores, as an
extension of the \Zero-one" normalization used by Wu, Crestani and Bi in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>Let us focus on a fusion of two results lists L1 and L2, composed of couples
(doc,rsv). To be realistic, L1 and L2 are limited to the top n results. Assume
that a document d has a score value of score(d; Li) in Li, with i 2 f1; 2g, and d
is at rank rank(d; Li) in Li; that bi is the bias of Li; and that hi is the horizon
(a rank position) above which we do not look at the in results list Li. The
normalized score of d in Li is then:
f (d; Li) =
( (1
vvmmaaxx(L(Li)i) svcomrien(d(L;Li)i) ) + bi if rank(d; Li)
hi
0
otherwise
3 Terrier: http://www.terrier.org</p>
      <p>with vmin and vmax the minimal and maximal value of scores in a result
list.</p>
      <p>
        Compared to the tting used by [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], our idea is that we allow di erent search
results to t into di erent intervals. This is independent of the way di erent
result lists are combined, but a kind of \boosts" that forces the nal score values
for a result list to be in f0g [ [bi; bi + 1] (the value 0 denotes that the document
does not occur in the top hi elements of the list). This boost is independent from
the way the scores are fused afterward. If we make a parallel with the general
tting proposed by [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], our proposal allows an independent scaling for each list
fused.
      </p>
      <p>
        Then, the overall fusion computes a weighted average of the normalized scores
(COMB-sum from [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]) using a parameter that denotes the relative importance
of L1 over L2, and rerank the results according to the new fused scores.
Compared to a usual weighted average, the di erence here comes mainly from the bi.
Assigning 0 to all bis leads to a usual weighted average COMB-sum.
      </p>
      <p>We discuss the impact of such biases in the section dedicated to the
experiments.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Fusion of content-only scores (run LIG 1)</title>
      <p>
        On experiments conducted over the SBS 2014 dataset, we noticed that fusing
several content-only runs had a positive impact with a relative nDCG@10
improvement larger than 10%. That is why we propose to fuse one result coming
from BM25F [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] run (parameters values taken from [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]) and one result coming
from a Log logistic model [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Grid optimization of parameters on SBS 2014
data led to the parameters used for the o cial run tagged LIG 1, described in
table 1. The fusion score is computed as follows:
      </p>
      <p>RSVLIG 1(Q; d) =</p>
      <p>BM25F (ScoreBM25F (Q; d) + bBM25F )
+
b
0.5
0.4</p>
    </sec>
    <sec id="sec-4">
      <title>Personalized IR exploiting pro les</title>
      <sec id="sec-4-1">
        <title>Non-social user pro le (run LIG 2)</title>
        <p>
          What we depict here as \non-social" corresponds to the individual user data. In
our case, these data refer to the catalog of the user. We assume that Catu denotes
the catalog (list of books) of a given user u (from the corpus U of users). To
construct a user pro le, we take inspiration from Cai and Li who consider each
user pro le as a vector of tags and use a L1 normalized term frequency (NTF)
to denote the preference degree of user on a tag [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Similarly, we describe the
pro le P rofu of a user u as a weighted vector based on Catu, where each term
is weighted by its NTF. In a way to keep only the major interests of a user, we
consider only the top n terms according to their values. In our runs, we keep the
top n = 100 terms in a pro le.
        </p>
        <p>Such pro le is used as an expansion of initial query. In a way to re ect the
relative importance of a term in a pro le, we de ne a function Exp(Q; u) that
expands the query by a xed number of terms, and the relative importance of
each term in the pro le is re ected as corresponding number of occurrences of
this term in the expanded query. For instance, if the value corresponding to
the term t in a user pro le accounts for 40% of the occurrences of terms in the
pro le, and suppose that we x the number of terms added in the query to 100,
then the query will be expanded by 0:4 100 = 40 occurrences of the term t.
Then a BM25 retrieval is achieved on the documents corpus.</p>
        <p>We noticed in SBS 2014 data that such expansion does not provides good
results, but that the fusion of the results of such expanded queries and the
usual content-only queries lead to better results. That is why we experimented
such fusion with the parameters de ned in table 2, where N SP rof denotes the
parameters related to the non social pro le fusion. The overall score for LIG 2
is:</p>
        <p>RSVLIG 2(Q; d; u) =</p>
        <p>BM25F (ScoreBM25F (Q; d) + bBM25F )
+
+
NSP rof (ScoreBM25(Exp(Q; u); d) + bNSP rof )
(2)
where BM25F , LGD; NSP rof are the relative importance of BM25
result list, LGD results list, the non social pro l list respectively. Also, bBM25F ,
bLGD; bNSP rof are the respective bias of each results list. Moreover, the
normalized scores use a horizon h of 1000, as presented in the table 2.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Friendship link generation</title>
        <p>For the social user pro le, we choose to generate \friendship" links between topic
users and the non-topic users provided by SBS. To achieve that, we assume that
what makes (topic or non-topic) users similar to others is their catalog and
the ratings they provide. We represent then all the non-topics users as a text
document corresponding to concatenation of the document ids from the user
catalog. We include the ratings (integer values) by using the ratings as the tf
values for the number of occurrences of the documents ids.</p>
        <p>To be able to nd the non-topic users similar to topic users, we describe the
users topics in the same way as the non-topic users as described above. Then we
used the topic-users descriptions as queries on the corpus of non-topic users using
a classical BM25 matching. For rst experimentation, we lter the relationships
to the top 2 most similar non-topic users for each topic user, and we plan to
experiment the top k similar users in future works.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Usage of \friends"</title>
        <p>Once 2 closer friends of a topic user are obtained, we apply a process similar to
section 4 to generate the non-social pro les of the friends, and then we match
the topic query with the friends pro les to get documents that match the query.
The matching is computed as follows:</p>
        <p>RSVLIG 3(Q; d; u) =</p>
        <p>The fusion parameters used for the o cially submitted run LIG 3 are given
in table 3, with F ri1 and F ri2 the two friends of u.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Documents given as \examples" (runs LIG 4, LIG 5 and LIG 6)</title>
      <p>One important point to notice is that the post processing of the obtained
results have a dramatic impact on the results. For instance, as the initial corpus
ids (isbn) are not the ones on which the results are evaluated (LibraryThing
ids), and because of potential duplicates generated, it is not obvious to handle
the translation. Our approach for such duplicate removal was the same that is
provided by the organizers of SBS.</p>
      <p>
        Additionally, for our runs for SBS 2015, we focused on integrating the users
examples to post-process the queries. Our idea was that a user might be
interested if he nds as answers documents that he read and that he appreciated, as
this would be an indicator that the system is providing relevant documents to
him. We declined this hypothesis in two ways:
Reranking: Achieve a reranking where the documents a user likes are boosted,
and the documents he dislikes are removed for the result. After a \Zero-one"
score normalization [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] between 0 and 1 of the overall score, we add 1 for the
documents that the user likes and set the score to 0 for the documents that
he dislikes. This process is then a post-processing that is run after the fusion,
and is the result of our o cial run LIG 4. It is worth noting that, if several
retrieved documents are liked by the user, their relative initial ranking is
preserved;
Relevance Feedback: De ne a relevance feedback, positive for the documents
that the topic user likes, and negative for the documents he does not like.
We achieved such relevance feedback on the Log logistic run LGD for our
o cial run LIG 5, and also on both content-runs, i.e., BM25F and Log
logistic, for our o cial run LIG 6. The relevance feedback uses all the positive
documents and selects the top 10 terms according to the default selection of
Terrier [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Results</title>
      <p>We present here two elements. First, we list the o cial results obtained for our
6 runs o cially submitted to SBS 2015. Second, we discuss additional results
generated after the release of the SBS 2015 qrels, presenting the impact of the
\biases" we used (see section 2) over \unbiased" results.
6.1</p>
      <p>O</p>
      <p>cial Results
We comment here mainly the lines of table 4 corresponding to boldfaced run
ids that use nor reranking neither relevance feedback. We notice then that the
impact of the user non-social pro le is clearly bene cial, however the p-value of
a bilateral paired Student t-test on LIG 1 versus LIG 2 equal 5.45%, thus with
a signi cance threshold of 5% this di erence is not statistically signi cant. Such
value is even larger between LIG 1 and LIG 3. We notice a slight improvement of
nDCG@10 results when integrating the 2 best \friends", however our generation
or usage of relationships between users does not seem to be e ective enough.
According to what we de ned for our fusion (see section 2), we also notice that
the weights assigned to the friends are very small, 0.05. With higher relative
values the results degrade. So we conclude for now that our proposal does not
outperform the integration of non-social user information.</p>
      <p>As we see on table 4, the results with reranking or relevance feedback lower
the quality of the results, but these elements are related to our interpretation of
the catalogs and examples that are incompatible with the interpretation of the
SBS organizers. In fact, our interpretation was somewhat the exact contrary of
what decided the SBS organizers (they choose that the catalog + examples must
not be part of the result), this explains why these additional runs behave worse
than our initial runs.
We describe in table 5 the impact of using the biases as de ned in section 2.
To be fair compared to the o cial results, we choose to only remove the bias
from the con gurations used for the o cial runs and to compare the relative
gain or loss (between parentheses) with respect to the biased respective runs
from table 4. In this table, the values use 3 digit precision numbers, where the
percentages are computed on 4 digit precision numbers. We notice that the e ect
of the bias are positive for all the measures for the runs LIG 2 and LIG 3, and
have almost no e ect of the content only LIG 1 run. So, the e ect of the biases
seem to be more positive as we fuse many lists.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>We presented in this paper the experiments that were conducted for the
participation of LIG to the SBS 2015 lab evaluation. Our main nding, according
to our integration of non-social and social pro les, is that the use of non-social
pro le has a clear positive impact on the quality of the retrieval, where the
integration of generated friendship relationships does not really increase the quality
of the system provided. One important conclusion that we draw from the SBS
experiments is that the post processing of results has a dramatic impact on the
quality of the results, and then must be carefully studied.</p>
      <p>The experiments reported here depict our rst steps to grasp the complexity
of personalized information retrieval in social context, and many e orts will focus
on re ning and characterizing the numerous elements involved in such retrieval
process.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgment References</title>
      <p>This work is supported by Region Rh^one-Alpes through the ReSPIr project.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          :
          <article-title>Personalized search by tag-based user pro le and resource pro le in collaborative tagging systems</article-title>
          .
          <source>In: Proceedings of the 19th ACM International Conference on Information and Knowledge Management</source>
          . pp.
          <volume>969</volume>
          {
          <fpage>978</fpage>
          . CIKM 10,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Clinchant</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaussier</surname>
          </string-name>
          , E.:
          <article-title>A Log-Logistic Model for Information Retrieval</article-title>
          .
          <source>In: 18th ACM Conference on Information and Knowledge Management. CIKM 10</source>
          , vol.
          <volume>14</volume>
          , pp.
          <volume>5</volume>
          {
          <fpage>25</fpage>
          .
          <string-name>
            <surname>Hong-Kong</surname>
          </string-name>
          ,
          <source>China</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hafsi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gery</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beigbeder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : LaHC at INEX 2014:
          <article-title>Social book search track</article-title>
          .
          <source>In: Working Notes for CLEF 2014 Conference</source>
          . pp.
          <volume>514</volume>
          {
          <issue>520</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Koolen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bogers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamps</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kazai</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Preminger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Overview of the INEX 2014 social book search track</article-title>
          . In: Working Notes for CLEF 2014 Conference,
          <article-title>She eld</article-title>
          ,
          <source>UK, September 15-18</source>
          ,
          <year>2014</year>
          . pp.
          <volume>462</volume>
          {
          <issue>479</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ounis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amati</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plachouras</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macdonald</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Terrier: A High Performance and Scalable Information Retrieval Platform</article-title>
          .
          <source>In: Proceedings of ACM SIGIR'06 Workshop on Open Source Information Retrieval (OSIR</source>
          <year>2006</year>
          )
          <article-title>(</article-title>
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaragoza</surname>
          </string-name>
          , H.:
          <article-title>The probabilistic relevance framework: BM25 and beyond</article-title>
          .
          <source>Found. Trends Inf. Retr</source>
          .
          <volume>3</volume>
          (
          <issue>4</issue>
          ),
          <volume>333</volume>
          {
          <fpage>389</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Shaw</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>E.A.</given-names>
          </string-name>
          :
          <article-title>Combination of multiple searches</article-title>
          .
          <source>In: The Second Text REtrieval Conference (TREC-2)</source>
          . pp.
          <volume>243</volume>
          {
          <issue>252</issue>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Evaluating score normalization methods in data fusion</article-title>
          . In: Ng,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Leong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.K.</given-names>
            ,
            <surname>Kan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.Y.</given-names>
            ,
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.) Information Retrieval Technology - Third
          <source>Asia Information Retrieval Symposium</source>
          ,
          <string-name>
            <surname>AIRS</surname>
          </string-name>
          <year>2006</year>
          . vol.
          <volume>4182</volume>
          , pp.
          <volume>642</volume>
          {
          <fpage>648</fpage>
          . Springer Berlin Heidelberg (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>