<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>USTB at Social Book Search 2016 Suggestion Task: Active Books Set and Reranking</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shao-Hui Feng</string-name>
          <email>shaohui_feng@xs.ustb.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bo-Wen Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zan-Xia Jin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xu-Cheng Yin</string-name>
          <email>xuchengyin@ustb.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jian-Lin Jin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jian-Wei Wu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Le-Le Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hao-Jie Pan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fan Fang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fang Zhou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Technology, University of Science and Technology Beijing (USTB)</institution>
          ,
          <addr-line>Beijing 100083</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we describe our participation in the Social Book Search(SBS) Suggestion Task. We developed some new re-ranking models, based on what we applied in last year. We used Galago search from the index which was built with active books. The queries input to search engine were composed by key words from topics, and then we performed re-ranking models(popularity related) on Galago searching results on enriched XML index by 14 di erent elds. Experiments on these approaches shown that an enriched index and key query model improves the e ectiveness. As our approaches in INEX2014 [1],SBS2015 [2] and combined those re-ranking models according to a speci c order shown the best performance.</p>
      </abstract>
      <kwd-group>
        <kwd>XML retrieval</kwd>
        <kwd>key query</kwd>
        <kwd>re-ranking</kwd>
        <kwd>active books set</kwd>
        <kwd>popularity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In this paper, we describe our participation in the Social Book Search 2016
suggestion task. Our goals for this task were (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) to investigate the contribution of
key words from topics in searching; (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) to testify the active books set functioning
in searching (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )to testify the e ect of popularity related re-ranking approaches
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) to nd an e ective approach to combine the results of di erent re-ranking
models.
      </p>
      <p>The structure of this paper is as follows. We start Section 2 by describing
our methodology: pre-processing on the XML formatted documents, indexing,
searching by Galago, introduction of key query and active books set. In the
section 3,we describe the re-ranking models and the re-ranking models experiments.
In section 4, we describe the results of our enriched index, key query model,
active books set and re-ranking models. Section 5 describes about the runs which
we submitted,with the results of those runs presented. We discuss our results
and conclude in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <sec id="sec-2-1">
        <title>Data Pre-Processing</title>
        <p>We perform a process similar to [1], such as expand and enrich the documents
XML with replacing the numeric information with textual information. For
instance:
We change it to</p>
        <p>&lt;tag count="3"&gt; ction&lt;/tag&gt;
&lt;tag&gt; ction ction ction&lt;/tag&gt;.</p>
        <p>In addition, this time we expand to 14 elds when we clean the XML documents,
it is di erent from last year. They are
title; isbn; tags; review content; review summary; dewey; f irstwords;
lastwords; chracters;places; subjects; browsenodes; abstract; addcontent.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Indexing</title>
        <p>Galago 1 is an open-source search engine. In order to improve the search
effectiveness, we study two strategies to build the index. One indexing strategy
is the normal indexing approach described as follows. Experimentally, we nd
that the elds (etc. the title, tag, content and summary) are more relevant and
meaningful than others in the XML formatted documents. So we build our
basic index by removing the other elds which were not useful. Another strategy
is to enrich the basic index. Observing the book information from the Library
Thing, we nd out that a large proportion of books lack the content and
summary elds. Therefore, documents expansion technology is expected to utilized
to enrich the basic index. Firstly, we select two web sites which contain a large
amount of more useful metadata of books. The books we use are the literatures
written in English in douban.com 2 and all books in lookupbyisbn.com. Then
we crawl the brief introduction of douban.com and the book description eld of
lookupbyisbn.com. Both web sites are available by ISBN. With the content from
both web sites, we enrich six hundred thousand of books (see the examples of
book document which is used for index in XML 1 and XML 2). The enriched
index is based on the enriched information.</p>
        <p>XML 1: Book document
&lt;book&gt;
&lt;title&gt;Mister Monday&lt;/title&gt;
&lt;summary&gt;So good, you can't put it down!&lt;/summary&gt;
&lt;content&gt;Now, I had...&lt;/content&gt;
&lt;tag count="9"&gt;children's literature&lt;/tag&gt;
&lt;/book&gt;
1 http://www.galagosearch.org/
2 http://book.douban.com/</p>
        <p>XML 2: Enriched book document
&lt;book&gt;
&lt;title&gt;Mister Monday&lt;/title&gt;
&lt;summary&gt;So good, you can't put it down!&lt;/summary&gt;
&lt;content&gt;Now, I had...&lt;/content&gt;
&lt;tag count="9"&gt;children's literature&lt;/tag&gt;
&lt;brief introduction&gt;the content is from the douban.com&lt;/brief introduction&gt;
&lt;description&gt;the content is from the lookupbyisbn.com&lt;/description&gt;
&lt;/book&gt;
2.3</p>
        <p>Generate Keywords from queries</p>
        <p>We collect all the topics from 2011 to 2015. Our purpose is to get the best
evaluation results after a great quantity of attempts. So that we can get the best
queries for these topics, then we summed up a word list for topic queries. In
Social Book Search 2016, we used the word list to lter the request eld, then
combine it with the title eld compose the queries for Galago.
2.4</p>
        <p>Active subset in book collections</p>
        <p>Through the analysis of the pro le, we select a subset of high frequent books
which is called active book set. In this set, all the books are considered much
more should be recommended to the topic author. We lter out all the books
which are not in the active set. The operation is similar to lter the catalog and
example books of the topic.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Re-ranking Models</title>
      <p>Some of the re-ranking approaches were proposed and used by USTB at
INEX2014 [1] and proposed by Toine Bogers in 2012 [4], which proved to be
e ective. This year our re-rank approach can be roughly divided into two
categories 1) Tag-Rerank (T), similar product re-rank (S), Read by one re-rank (R)
2) popularity re-rank (P ), example recommended re-rank (E), Reader number
re-rank (R), Browsnodes re-ranking (B). The rst category was based on
computing the similarity of two books that from the original result from Galago. The
second category was based on attributes (number of readers, popularity ) of the
books or the example books of the topics.</p>
      <p>
        We use these models to re-rank by the following stages:
1)Similarity Calculation. Models like T , N focus on the eld &lt;tag&gt; and
&lt;BrowseNode&gt;. We can build a feature matrix for features for each eld.
Equation (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) is used to calculate the T , N , T N similarities of two documents.
      </p>
      <p>
        Features like I focus on the eld &lt;similar-product&gt;, the similarities of two
documents based on the feature I is calculated by the Equation (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ).
! !
simij(f ) = cos &lt; !fi ; !fj &gt;= !fi !fj
jfi jjfj j
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
simij (I) =
8&gt;1;
&gt;
&gt;
&gt;
&gt;
&gt;
&lt;
&gt;
&gt;
&gt;
&gt;
&gt;:&gt;0;
i is j's similar product or
j is i's similar product
0:5; i is j's similar product's similar product
or j is i's similar product's similar product
else
First category rerank We re-rank the top 1000 list of initial ranking for the
above-mentioned features by Equation (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ). For feature R, we use Equation (
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
[6] and for B, we use Equation (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ).
      </p>
      <p>
        score0(i) =
score(i) + (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
N
X simij score(j)(j 6= i)
j=1
      </p>
      <p>
        Pr2Ri r
jreviews(i)j
score0(i) =
score(i) + (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) log(jreviews(i)j)
score(i) (
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
where Ri is the set of all ratings given by users for the document i, and jreviews(i)j
is the number of reviews.
      </p>
      <p>
        score0(i) =
score(i) + (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
score(i)
      </p>
      <p>
        (
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
1 + BA(i)
1 + BAmax
where BA(i) is the Bayesian average rating of document i, which can be referred
to[7].
      </p>
      <p>In addition, Tag re-rank and read by one re-rank (we hold the opinion that if
two books are read by same reader, they are related) are very similar to similar
product re-rank.</p>
      <sec id="sec-3-1">
        <title>Second category reranking</title>
      </sec>
      <sec id="sec-3-2">
        <title>Popularity re-rank approach</title>
        <p>We got a data set from Library Thing like this:</p>
        <p>W ork idi;</p>
        <p>popularityi</p>
        <p>
          When we input initial search result to this model, we will change the works
score by equation (
          <xref ref-type="bibr" rid="ref6">6</xref>
          )
score0i =
scorei + (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) scorei (1
popularity=30; 0000)
        </p>
        <p>Just as what we see, popularity is the number from the set, score is from
Galago search result, score is the nal output score from this re-rank modes.</p>
        <p>Example recommended reranking approaches</p>
        <p>Through summarizing the information from the pro le, we got a example
recommended list for some topics like this format:</p>
        <p>
          T opic idi recommended1; recommended2; recommended3; etc:
When using the re-ranking approach, if a work of the result is contained in
the topics list. We would change it rate by the equation (
          <xref ref-type="bibr" rid="ref7">7</xref>
          )
score0i = scorei +
scorei
(
          <xref ref-type="bibr" rid="ref7">7</xref>
          )
if the work in the result was not contained in the topic id 's list,
= 0.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Reader number re-ranking approach</title>
        <p>We extracted information from the pro le, then we got a set which showing
the workss reader number.Like this:</p>
        <p>word idi; reader numberi
Through this approach, we optimize the search result by the below equation(8)
3) Combining. We applied these approaches on the key query search result
according to a speci c order, then got the nal result.</p>
        <p>As shown in Table 1, the best performance is obtained from Initial+keywords+
allReRank + f iltercatalogandexampl ,and active; readernumber make greater
contributions to the improvements.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Submitted Runs</title>
      <p>We selected six automatic runs for submission to SBS2016 based on our Key
Query and re-ranking Models. They are:</p>
      <p>run1. This run was made by a searching-re-ranking process where the
initial retrieval result was based on the selection of query keywords and a small
index of active books, the re-ranking results based on a combination of several
strategies (number of people who read the book from pro le, similar-product
from amazon.com, popularity from LT forum, etc.).</p>
      <p>run2. This run was made by a searching-reranking process where the initial
retrieval result was based on the selection of query keywords and a small index
of active books, the re-ranking results based on number of people who read the
book from pro le.</p>
      <p>run3. This run was made by a searching-reranking process where the initial
retrieval result was based on the selection of query keywords and a small index of
active books, the re-ranking results based on the books in a same user's pro le.</p>
      <p>Run4. This run was made by a searching-reranking process where the initial
retrieval result was based on the selection of query keywords and a small index
of active books, the re-ranking results based on similar products provided by
amazon.</p>
      <p>run5. This run was made by a searching-reranking process where the initial
retrieval result was based on the selection of query keywords and the full index
ltered by active books, the re-ranking results based on the books in a same
user's pro le.</p>
      <p>run6. This run was made by a searching-reranking process where the initial
retrieval result was based on the selection of query keywords and the full index
ltered by active books, the re-ranking results based on a combination of several
strategies (number of people who read the book from pro le, similar-product
from amazon.com, popularity from LT forum, etc.).
5</p>
    </sec>
    <sec id="sec-5">
      <title>Result</title>
      <p>The runs submitted to the Social Book Search 2016 were evaluated using
graded relevance judgments. The relevance value were labeled manually according to
the behaviors of topic creators, for example, if creator adds book to catalogue
after it's suggested, the book is treated as highly relevant. A decision tree was
built to help the labeling 3. All runs were evaluated using NDCG@10, MRR,
MAP, R@1000 with NDCG@10 as the main metric. Table 2 shows the o cial
evaluation results. Results of the six submitted runs on Social Book Search 2016,
evaluated using all 120 topics with relevance value calculated from the decision
tree. The best run scores are printed in bold, we got the rst place in the
competition. In addition, all the results we submitted are in the top six.</p>
      <p>So we got the rst place in the Social Book Search suggestion task 2016.</p>
      <p>It necessary to state that run 7, run 8 and run 9 are our additional
experiments, obvious we can nd that key query and active books set have a
great increase in results .Key query improve Initial+stopwords run by about 25
percent, then active books set improves the Initial+keyQuery run by about 24
percent.We see that the best-performing run on all 120 topics was run1 with an
NDCG@10 of 0.2157. Run 1 used Key Query, small index built by active
books and all re-ranking models combine. Also we see that re-ranking model does
improve over the initial results by Galago searching engine.</p>
      <p>All the runs from 1 to 6 were ltered by the topics catalog books set and
example books set.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion &amp; Conclusion</title>
      <p>All of the re-ranking approaches can improve the evaluation results, the best
results are from combined all re-ranking approaches. Of course, the use of key
queries and active book set make the greatest contributions to the e ectiveness
of our systems.</p>
      <p>This year, we used much information from pro le, and we got a better
performance, but we failed to make use of the random forest to combine the re-ranking
to improve the result. We keep the opinion that machine learning can get a
better results. So, it is worth discussing how to combining the re-ranking results
with machine learning algorithms.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bo-Wen</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Xu-Cheng Yin,
          <string-name>
            <surname>Xiao-Ping</surname>
            <given-names>Cui</given-names>
          </string-name>
          , Bin Geng, Jiao Qu,
          <string-name>
            <surname>Fang Zhou</surname>
          </string-name>
          ,
          <article-title>Li Song and Hong-Wei Hao</article-title>
          . USTB at INEX2014:
          <article-title>Social Book Search Track</article-title>
          . In INEX'13 Workshop Pre-proceedings. Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Working Notes of CLEF 2015 -
          <article-title>Conference and Labs of the Evaluation forum</article-title>
          , Toulouse, France, September 8-
          <issue>11</issue>
          ,
          <year>2015</year>
          . 2015
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>T.</given-names>
            <surname>Bogers</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Larsen</surname>
          </string-name>
          . Rslis at inex 2013:
          <article-title>Social book search track</article-title>
          .
          <source>In INEX'13 Workshop</source>
          Pre-proceedings. Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>T.</given-names>
            <surname>Bogers</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Larsen</surname>
          </string-name>
          . Rslis at inex 2012:
          <article-title>Social book search track</article-title>
          .
          <source>In INEX'12 Workshop</source>
          Pre-proceedings, pages
          <fpage>97</fpage>
          -
          <lpage>108</lpage>
          . Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bo-Wen</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Xu-Cheng Yin,
          <string-name>
            <surname>Xiao-Ping</surname>
            <given-names>Cui</given-names>
          </string-name>
          , Bin Geng, Jiao Qu,
          <string-name>
            <surname>Fang Zhou</surname>
            ,
            <given-names>Li</given-names>
          </string-name>
          <string-name>
            <surname>Song</surname>
          </string-name>
          and
          <string-name>
            <surname>Hong-Wei Hao</surname>
          </string-name>
          .
          <article-title>Social Book Search Reranking with Generalized ContentBased Filtering</article-title>
          . CIKM'
          <volume>14</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>R. D. Ludovic</surname>
            Bonnefoy and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Bellot</surname>
          </string-name>
          .
          <article-title>Do social information help book search</article-title>
          ? In INEX'12 Workshop Pre-proceedings, pages
          <fpage>109</fpage>
          -
          <lpage>113</lpage>
          . Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Marijn</given-names>
            <surname>Koolen</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          .
          <article-title>Comparing topic representations for social book search</article-title>
          .
          <source>In INEX'13 Workshop</source>
          Pre-proceedings. Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>