<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SOCIAL BOOK SEARCH TRACK: ISM@INEX'14 SUGGESTION TASK</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ritesh Kumar</string-name>
          <email>R@1000</email>
          <email>ritesh4rmrvs@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sukomal Pal</string-name>
          <email>sukomalpal@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering, Indian School of Mines Dhanbad</institution>
          ,
          <addr-line>826004</addr-line>
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <fpage>521</fpage>
      <lpage>524</lpage>
      <abstract>
        <p>This paper describes the work that we did at Indian School of Mines towards Social Book Search Track for INEX 2014. We submitted ve runs in its Suggestion Task. We investigated individual e ect of title, group, mediated query, and narrative elds of the topics in our runs. For all the runs we used language modelling technique with Dirichlet smoothing. The run using only mediated query eld was our best. Overall, our performance is not satisfactory. However, as new entrant to the eld, our scores are encouraging enough to work for better results in future.</p>
      </abstract>
      <kwd-group>
        <kwd>Book Search</kwd>
        <kwd>Social Book Search</kwd>
        <kwd>Language modelling</kwd>
        <kwd>Information Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        With growing numbers of online portals and book catalogues, our current time
sees a rapid evolution in the way we acquire, share and use books. In order to
enable users, Social Book Seach Track at INEX [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] provides a relevant experimental
platform to investigate techniques of searching and navigating professional
metadata provided by publishers/booksellers and user-generated content from social
media [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. At INEX 2014, they o ered two tasks: Suggestion Task and Interactive
Task. We participated in the rst where we were supposed to recommend books
based on user's request and her personal catalogue data (list of books with rating
and tags maintained for the user in the social cataloguing site). We were also
provided with a large set of anonymised user pro les from LibraryThing forum
members. Each user request is provided in the form of topics containing di erent
elds like title, mediated query, group, narrative and catalogue information.
      </p>
      <p>As a newcomer to this eld, our goal this year was to investigate the
contribution of di erent topic elds in book recommendation. We only considered
title, mediated query, group, narrative elds from each topic. We did not consider
topic-creator's catalogue information. Neither we consulted anonymous user
proles.</p>
      <p>
        We submitted ve runs (run-ids: ISMD-341, ISMD-342, ISMD-350,
ISMD354, ISMD-355) in the Suggestion Task. For all the runs, Language modelling
with Dirchlet smoothing was used in Lemur's Indri search system [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Our overall
performance was not satisfactory. The run with only mediated query was best
among our submissions.
      </p>
      <p>Organization of rest of the paper is as follows. We describe our approach in
Section 2. Section 3 describes dataset and Section 4 reports results. In Section
5 we analyse our results. Finally, we conclude in Section 5 with directions for
future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <p>This year we took a simple approach similar to standard adhoc retrieval. The
document collection provided was stopword-removed and then stemmed using
Krovetz stemmer. It was indexed with Lemur Indri search system for all the
elds having text within.</p>
      <p>During retrieval, we tried to see the e ect of di erent components of a topic in
turn. We therefore used only title (Run-id ISMD-341), only group(Run
ISMD342), only title with stopword removed (Run ISMD-350), only mediated query
(Run ISMD-354), and only narrative with stopword removed eld (Run
ISMD355) from each topic.</p>
      <p>On top of standard English stopwords we identi ed a set of a few more like
recommendation, hello, suggestion, reference, recent, hi, thank, etc. which we
removed in the run ISMD-355.</p>
      <p>We also removed punctuation marks manually from all the textual content
of these elds and used only free text queries in all the runs.</p>
      <p>We did not consider any other information like catalogue information and
user pro le during retrieval.</p>
      <p>For each topic, we submitted 1000 book suggestions in the form of ISBNs.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data</title>
      <p>
        Test collection provided by INEX 2014 SBS orgainzers for Suggestion Task had
a document collection and a topicset. The document collection consists of 2.8
million book description with metadata from Amazon and LibraryThing. From
Amazon there is formal metadata like booktitle, author, publisher, publication
year, library classi cation codes, Amazon categories and similar product
information, as well as user-generated content in the form of user ratings and reviews.
From LibraryThing, there are user tags and user-provided metadata on awards,
book characters and locations and blurbs. There are additional records from the
British Library and the Library of Congress. The entire collection was 7.1 GB
in size. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
      </p>
      <p>The topic-set contains 681 topics each describing a user's request for
suggestion of books. Each topic has a set of elds like title, mediated query, group,
narrative and user's personal catalogue at the time of topic creation. The catalogue
contains a list of book-entries with information like LibraryThing id of the book,
its entry-date, rating and tags.</p>
      <p>The organizers also supplied 94,000 anonymised user pro les from
LibraryThing.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>
        The scores obtained by our ve runs are given in Table 1. The o cial
evaluation measure by INEX'14 is nDCG@10 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The performance of our runs are in
decreasing order. Our best performance is by ISMD-354 where we use only
mediated query eld. We also show the best score in the task demonstrated by
runid USTB-run6.SimQuery1000.rerank all.L2R RandomForest(*), for the
sake of comparison.
ISMD-354 22
ISMD-341 24
ISMD-350 27
ISMD-355 29
ISMD-342 32
best* 1
0.123
0.106
0.090
0.089
0.018
0.464
Although our performance is not up to the mark, there are few take-home lessons.
As individual elds, mediated query is the most e ective, followed by title and
narrative. Removing stopwords from the title is actually detrimental (ISMD-341
and ISMD-350). We did not consider any combination of these elds. It would
be interesting to see the performance of di erent combinations of these elds.
6
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>This year we participated in the Suggestion Task of Social Book Search as
initial venture. We tried to see the individual e ect of di erent topic- elds on book
recommendation. We considered only a handful of elds like mediated query,
title, narrative etc from the topics. While there can be no denial of the fact that
our overall performance is dismal, initial results are suggestive as to what should
be done next. We need to consult other elds like book catalogue of the topic
creators, ratings of the books in the catalogue during retrieval. We also need to
take into account pro les of other users. It is also imperative to see the
performance of combination of di erent elds in the topics as well as other elds in
user catalogues and user pro les. We shall be exploring some of these tasks in
the coming days.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Marijn</given-names>
            <surname>Koolen</surname>
          </string-name>
          , Gabriella Kazai, Jaap Kamps, Michael Preminger,
          <article-title>Antoine Doucet and Monica Landoni, Overview of the INEX 2012 Social Book Search Track</article-title>
          . INEX'12 Workshop Pre-proceedings, Shlomo Geva, Jaap Kamps, Ralf Schenkel (editors),
          <source>September 17-20</source>
          ,
          <year>2012</year>
          , Rome , Italy.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. INEX,
          <article-title>Initiative for the Evaluation of XML Retrieval</article-title>
          . https://inex.mmci.unisaarland.de/data/documentcollection.jsp
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. INDRI:
          <article-title>Language modeling meets inference networks</article-title>
          , Available at http://www.lemurproject.org/indri/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jarvelin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kekalainen</surname>
          </string-name>
          , J.:
          <article-title>Cumulated Gain-based Evaluation of IR Techniques</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          <volume>20</volume>
          (
          <issue>4</issue>
          ) (
          <year>2002</year>
          )
          <fpage>422</fpage>
          -
          <lpage>446</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. INEX,
          <article-title>Initiative for the Evaluation of XML Retrieval</article-title>
          . https://inex.mmci.unisaarland.de/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>