<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Genre Prediction to Inform the Recommendation Process</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>CCS Concepts</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Prediction</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Genre</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Books</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Time Sequence</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ARIMA</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Maria Soledad Pera Computer Science Department Boise State University Boise, ID</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Nevena Dragovic Computer Science Department Boise State University Boise, ID</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present a time-based genre prediction strategy that can inform the book recommendation process. To explicitly consider time in predicting genres of interest, we rely on a popular time series forecasting model as well as reading patterns of each individual reader or group of readers (in case of libraries or publishing companies). Based on a conducted initial assessment using the Amazon dataset, we demonstrate our strategy outperforms its baseline counterpart.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Books, which constitute a billion dollar industry1, are the
most popular reading material among all generations of
readers, both for leisure and educational purposes. Hundreds of
thousands of books of di erent types (e.g., paperback and
ebooks) and styles (e.g., ction and non- ction) are published
on a yearly basis, giving readers a variety of options to choose
from. Book recommendation systems, which are meant to
enhance the decision making process, can help users by
identifying, among the sometimes overwhelming number of diverse
books, the ones that best suit their interests and preferences.
These recommenders are not exclusively designed to aid
individuals in their quest for reading materials. They can also
improve the decision making process for libraries, by
suggesting what books to buy in order to maximize the use of
library resources by their patrons, and publishing companies,
by advising which books to publish in order to maximize
revenue. To better serve stakeholders, recommenders must
be able to predict interest and needs. However, given that
preferences may alter over time for di erent readers, the
1http://goo.gl/GMn8Nc
Copyright held by the author(s).
time component is important and crucial to consider in the
prediction process [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        There are many avenues that can be explored from a
timesensitive stand point in order to generate better
recommendations, including user-generated ratings, reviews or books
metadata. One of them, which is often overlooked and is the
focus of our paper, is genre. By its de nition, genre (e.g.,
drama, comedy) is a category of literary composition,
determined by literary technique, tone, content, or even length.
While genre has been studied as a part of the
recommendation process [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], the in uence of its distribution over time
on suggesting suitable books for individual or group of users
(as in the case of libraries and publishing companies) has
not been explored. Change of genre in time is a signi cant
dimension to improve the genre prediction process. This can
consequentially in uences the process performed by book
recommenders since it provides the likelihood of reader(s)
interest in each genre based on its occurrences at a speci c
point of time, not only the most recent or the most frequently
read one. As an answer to this need, we propose a genre
prediction strategy that examines genre distribution over
time and applies time series analysis models. The goal of
our strategy is to discover di erent areas of users' interests,
not only the most dominant ones.
      </p>
      <p>From the users' point of view, explicitly including time based
analysis to inform the recommendation process will lead to
relevant suggestions that satisfy their speci c reading needs.
Finally, from the commercial point of view, the bene t would
be in understanding the in uence of reading patterns on
decisions about what genre should be published or acquired
in a given point of time.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        A considerable number of studies examine the importance
of book genre on readers' activity [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. However, to the best
of our knowledge, research based on past genre distribution
coupled with time series analysis to in uence the
recommendation process has not been conducted. As presented in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
genre is used as a data point to inform a cross-domain
collaborative ltering approach that recommends books based on
users' genre preferences. You et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] propose a clustering
method based on users' ratings and genre interests extracted
from social networks to solve the cold-start problem a
ecting collaborative ltering approaches. Unlike the proposed
methods, we use time series to predict the genres most likely
currently of interest to each individual user, which can further
enhance the book recommendation process.
3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>METHOD</title>
      <p>Predicting genre to inform the recommendation process,
regardless of the major stakeholder (a reader, library or
publishing company), involves examining genres considered
in the past, either read by speci c users or purchased by
customers. While a simple genre distribution analysis yields
probabilities or weights that determine the most favored
genres, it lacks the ability to consider genre preference
evolution over time. To overcome this drawback, we propose a
time-based genre examination which requires information on
reading activities among readers. As a rst step to our
proposed strategy, we explore reading activity of a user to obtain
the distribution of his/her genre interest during continuous
periods of time. IN every period of time, for each genre we
calculate a signi cance score that captures its importance
by considering a number of books read of that genre in that
period of time. Thereafter, to explicitly consider the change
of genre preference distribution over time, our genre
prediction strategy takes advantage of Auto-Regressive Integrated
Moving Average2 (ARIMA). We selected ARIMA since it
is one of the most popular models that uses time series for
prediction purposes.</p>
      <p>By using ARIMA we are able to determine a model tailored
to each genre distribution to predict its importance for the
corresponding user in real time based on its previous
occurrences. Note that each predicted genre importance score is
based on: its occurrences in the past, a speci c time when
it occurred and its importance for a speci c user. To de ne
length of time periods used by ARIMA, we used information
from a recent study done by Pew3 on reading habits in the
USA, we establish one month long \windows" of time in
which each user is expected to read at least one book, so our
strategy uses 1 month time frames from the the rst book
log (either bookmarked or rated book) to last.</p>
    </sec>
    <sec id="sec-4">
      <title>4. INITIAL EVALUATION</title>
      <p>
        Framework. To validate the performance of our proposed
time-based genre prediction strategy, we selected a subset of
the Amazon/LibraryThing4 book dataset. Since the dataset
does not always include genre as a part of the provided
metadata, we extended it by including genre information from
the Library of Congress5. We used 1214 users6 along with
the books they rated or reviewed. To quantify the
assessment, we applied Mean Average Error (MAE), Accuracy and
Kullback-Leibler (KL) divergence [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. MAE estimates the
di erence between the predicted genre importance and the
ground truth, i.e., genre distribution for a user at a given
time, whereas Accuracy applies a binary strategy that re ects
if the predicted genres correspond to the ones read by a user
in a given period of time. KL divergence measures how well
a distribution q generated by a prediction strategy
approximates to distribution p, the ground truth. In establishing
the ground truth for each user considered for evaluation
purposes, we adopted the well-known N-1 strategy, such that
the genre of the books rated by a given user U in the N time
frame are treated as \relevant" genres for U, and the genre
of the books rated in the previous N-1 windows are used for
training U 's genre prediction model. As a baseline of our
initial assessment, we use a traditional prediction strategy
that considers the proportion of occurrences of each genre
based on data collected over N-1 periods of time to estimate
the importance of each genre for a given user on the current,
i.e., N, time period.
      </p>
      <p>Results. As shown in Table 1, for N=117 outperform the
baseline. KL divergence scores showcase that genre
distribution predicted using time-series approach better
approx2http://goo.gl/Dhzcg7
3http://goo.gl/BAUQK4
4http://goo.gl/drH0yF
5https://www.loc.gov/
6In our initial assessment, we considered Amazon users who
provided ratings for at least 35 books.
7We empirically veri ed that for 6&lt;N&lt;11 the results are
comparable to the ones for N=11
With Time Series
Without Time Series
With Time Series (3+ genre)
Without Time Series (3+ genre)
imates to the ground truth. Furthermore, the probability
of occurrence of each considered genre is closer to the real
values when the time component is included in the prediction
process. As a further assessment, we observed the di erences
in genre predictions among users who read di erent number
of distinct genres. For users who read only one to two genres,
the time-based prediction strategy does not perform better
than the baseline. However, if a user reads three or more
genres, our time-based genre prediction strategy outperforms
the baseline in all three metrics. This is not surprising,
given that it is not hard to determine area(s) of interest for
a user who constantly reads only one or two book genres,
which is why the baseline performs as good as time-based
prediction strategy. Given that users that read 3 or more
genres represent 91% of the users in our sampled dataset,
the proposed strategy provides signi cant improvements in
predicting preferred genre for the vast majority of readers.</p>
    </sec>
    <sec id="sec-5">
      <title>5. CONCLUSIONS</title>
      <p>In this paper, we described our e orts in developing a
time-based genre prediction strategy that can better inform
the recommendation process. The novelty of our approach
consists of incorporating an explicit time component to
generate genre distribution. To the best of our knowledge, this is
the rst time that the well-known time series ARIMA model
is used to predict book genre of readers' interests. The
described strategy provides successful predictions and
outperforms the baseline for 77% of users based on the presented
initial evaluation, while for the remaining users it provides
predictions comparable to the baseline. Because of the scope
of this paper, the conducted evaluation showcases the genre
prediction performance for a single user, while we still need to
conduct further assessments in terms of quantitatively
determining the degree to which the proposed strategy (i)provides
successful genre predictions for libraries and (ii)publishing
companies and in uences the recommendation process to
assist all three stakeholders.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>P.</surname>
          </string-name>
          <article-title>A erbach. The in uence of prior knowledge and text genre on readers' prediction strategies</article-title>
          .
          <source>Journal of Literacy Research</source>
          ,
          <volume>22</volume>
          (
          <issue>2</issue>
          ):
          <volume>131</volume>
          {
          <fpage>148</fpage>
          ,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lynch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Shapiro</surname>
          </string-name>
          .
          <article-title>Examining the e ects of gender and genre on interactions in shared book reading</article-title>
          .
          <source>Literacy Research and Instruction</source>
          ,
          <volume>43</volume>
          (
          <issue>4</issue>
          ):1{
          <fpage>20</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Schu</surname>
          </string-name>
          <article-title>tze. Foundations of statistical natural language processing</article-title>
          , volume
          <volume>999</volume>
          . MIT Press,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sarawagi</surname>
          </string-name>
          and
          <string-name>
            <given-names>S. H.</given-names>
            <surname>Nagaralu</surname>
          </string-name>
          .
          <article-title>Data mining models as services on the internet</article-title>
          .
          <source>ACM SIGKDD Explorations Newsletter</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <volume>24</volume>
          {
          <fpage>28</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <surname>Y. Zhang.</surname>
          </string-name>
          <article-title>Opportunity model for e-commerce recommendation: right product; right time</article-title>
          .
          <source>In ACM SIGIR</source>
          , pages
          <volume>303</volume>
          {
          <fpage>312</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>You</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Rosli</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Ha</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.-S.</given-names>
            <surname>Jo</surname>
          </string-name>
          .
          <article-title>Clustering method based on genre interest for cold-start problem in movie recommendation</article-title>
          .
          <source>JIIS</source>
          ,
          <volume>19</volume>
          (
          <issue>1</issue>
          ):
          <volume>57</volume>
          {
          <fpage>77</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>