<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Algorithms Requests Clicks CTR(%)
Recency</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>The Degree of Randomness in a Live Recommender Systems Evaluation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gebrekirstos G. Gebremeskel</string-name>
          <email>gebre@cwi.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arjen P. de Vries</string-name>
          <email>arjen@acm.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Access, CWI</institution>
          ,
          <addr-line>Amsterdam, Science Park 123, 1098 XG Amsterdam</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <volume>56</volume>
      <issue>350</issue>
      <abstract>
        <p>This report describes our participation in the CLEF NEWSREEL News Recommendation Evaluation Lab. We report and analyze experiments conducted to study two goals. One goal is to study the e ect of randomness in the evaluation of algorithms. The second goal is to see whether geographic information can help in improving the performance of news recommendation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In this report we describe the results of experiments we conducted during our
participation in the CLEF NEWSREEL News Recommendations Evaluation
Lab, the task of Benchmark News Recommendations in a Living Lab [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The
experiments have the following goals. The rst goal is to investigate the e ect of
system and/or user behavior randomness on recommender systems evaluations.
This randomness refers to the selection of algorithms for providing
recommendation request by the Plista framework and the user click behavior which can
be in uenced by the situation of the user. The second is to study whether
geographic information can play a role in news recommendation systems.
      </p>
      <p>
        The literature on recommender systems shows that o ine and online
recommender system evaluations do not concur with each other [
        <xref ref-type="bibr" rid="ref1 ref3 ref6">3, 1, 6</xref>
        ]. This is to say
that recommender systems behave di erently in o ine and online evaluations,
both in terms of absolute and relative performance. This has a serious implication
for recommender system research, because the whole point of o ine evaluation
is the assumption that at least the relative performance of recommender systems
is indicative of their relative online performance and thus an important step for
selecting algorithms that can be deployed in a live recommendation setting.
      </p>
      <p>
        Many reasons can be mentioned for the disparate performances of o ine and
online evaluations. The three most important papers [
        <xref ref-type="bibr" rid="ref1 ref3 ref6">3, 1, 6</xref>
        ] that compare online
and o ine evaluations use di erent datasets for o ine evaluation and online
evaluation. It is possible that the di erence in performance may be attributed
to this di erence in dataset. The literature lists some reasons for the disparate
performance of recommender systems in o ine and online evaluations [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. The
rst reason is that o ine evaluations can measure only accuracy; they do not take
the user behavior into account. The second reason is that o ine data-sets are
incomplete and imperfect, and thus recommender systems are being evaluated
based on incomplete dataset.
      </p>
      <p>The rst reason, that is, that o ine evaluations can measure only accuracy
and thus can not take user behavior into account can be understood in two
ways. One way is that online evaluations can be in uenced by adapted
implementations of the algorithms that were used in o ine evaluation, to take user
behavior into account, in which case it becomes a di erent implementation. The
other way to understand it is that online evaluations happen in an environment
that involves user behavior and thus their performances can be in uenced by
the system and/or user behavior \randomness". The rst interpretation implies
that o ine and online algorithms are di erently implemented. The second
interpretation implies that the di erence in o ine and online evaluations can be
due to the system and/or user behavior randomness that the online evaluations
are subjected to. We attempt to explore this randomness in our participation</p>
      <p>
        The motivation for the second goal comes from a descriptive study we
conducted on a Plista dataset we collected in our previous participation. In the
study, there were two ndings [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. One is that there is a substantial di erence
in the geographical distribution of the readerships of traditional news portals,
and the second is that within the same portal, the geographical distribution of
the readerships of the local news category and the rest of the categories shows
substantial di erence. In our experiments, we tried to exploit the second nding
to improve news recommendation.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <p>For the study of the e ect of system and/or user behavior randomness on
recommender system evaluations, we run two instances of the same news recommender
algorithm with the view to measuring the e ect of \randomness" involved in news
recommendation clicks. For exploiting geographical information for news
recommendation, we employ a recommendation diversi cation approach. Speci cally,
we include geographically relevant recommendation into the recommendation
list, thus diversifying the recommendations in favor of geographic relevance
2.1</p>
      <p>Algorithms
We experimented with ve algorithms, all of them modi cations of the recency
algorithm. The recency algorithm takes into account recency and popularity
of an item, and it has been shown to be a strong baseline in previous online
evaluations. The algorithmic variations that we experimented with are listed
below.</p>
      <p>
        Recency: This algorithm keeps the 100 most recently viewed items for each
publisher in consideration for being recommended to the user. For any
recommendation request, the items that are , at the time of the recommendation
request, the most recently read (clicked)) are the ones that are recommended.
We run two instances of this algorithm to get a sense of the randomness
involved in the selection of algorithms by the Plista framework [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and/or clicks
on recommendations by users.
      </p>
      <p>GeoRec: The geographical recommender takes the geographical region (states
to be speci c) of users and the local category of news items into account when
generating recommendations. We generate two sets of recommendations, one by
the recency recommender and one by a purely geographical recommender. For
the purely geographical recommender, we take the 100 most recently viewed
items and sort them according to their geographic conditional likelihood scores
generated by Equation 1.</p>
      <p>rua;ik = P (cik jgua )
(1)</p>
      <p>Where cik is the local category of item ik. An item is either in the local
category, or not. gua is the state-level geographical information of the user ua,
that is, the state the user belongs to. Then, we take top twice the number of
requested recommendations as recommendation of the geographic recommender.
For the recency, we take the requested number of recommendations from the
from the most recently viewed items. Then we intersect the two sets of
recommendations. If the number of elements in the intersection is not equal with the
requested number of recommendations, we get half 1 from purely geographic
recommender and half + 1 from recency recommender.</p>
      <p>GeoRecHistory: This is a modi cation of the GeoRec recommender; it
excludes from the recommendation list those items that the user has already
visited.</p>
      <p>RecencyRandom: This recommender recommends items randomly selected
from the circular-bu er. This was started a bit late, and we used it a baseline
algorithm.</p>
      <p>The two instances of Recency and the GeoREc algorithms were run for a
consecutive period of 53 days, from 2015-04-12 to 2015-06-03. The RecencyRandom
algorithm was started 12 days later on 2015-04-24.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussions</title>
      <p>We present two types of performance scores: cumulative and daily click-through
rates (CTR). The cumulative CTR is presented in Table 1. We see that the
performances are all close to each other. However, if we rank them, we see that
the GeoRec recommender followed by Recency followed by GeoRecHistory end
up as the rst three best performing algorithms. The daily performances of the
algorithms are shown in Figure 1, and the cumulative CTR as a function of the
number of days is shown in gure 2.</p>
      <p>From the daily ( gure 1) and cumulative ( gure 2) plots, we see that the
performance of any of the algorithms varies greatly. In the cumulative plot,
we see that Recency and Recency2 performed di erently for a long time, but
generally getting closer towards each other until some point after which they
seem to stabilize. If one was continuously monitoring the performances of this
●
●</p>
      <p>●
●
● ●
● ● ●</p>
      <p>●
●● ●</p>
      <p>● ● ● ● ●● ● ●● ●● ●● ●
●● ● ● ● ●● ● ●● ● ●
● ● ● ●● ● ● ●● ●
● ● ● ● ● ● ● ●
● ● ●</p>
      <p>●
● ● ● ●
● ● ● ●
●</p>
      <p>●
● ● ● ●
●
●●</p>
      <p>●●
● ●</p>
      <p>●
●</p>
      <p>● ●
R
T
C 1.0
2.0
1.5
0.5
0.0
1 5 9 13
18
23
28
33
38
43
48</p>
      <p>53</p>
      <p>Days
● ● ●● ● ●● ● ●●● ●●● ●● ● ●●●● ●● ● ● ● ●● ●●● ●●● ●● ●● ●● ●● ●●● ● ● ●● ● ●●● ●● ●●
● ● ● ●● ●● ●● ●● ●●● ● ● ● ● ●● ●● ●●● ● ● ● ● ● ● ● ●● ● ● ●● ● ●
●
●
●
R
T
C
0.8
0.6
0.4
0.2
recency
recency2
geoRec
geoRecHistory
recencyRandom
1
5
9
13
18
23
28
33
38
43
48</p>
      <p>53
Fig. 2. The cumulative CTR performances of the ve algorithms as they progress on
a daily basis
two algorithms, then one would say that the Recency algorithm is better than
Recency2 algorithm. This would happen, at least some times, even if one employs
statistical signi cance tests.</p>
      <p>Let's assume that the experimenter was peeking at the experiments everyday
to make a decision on which algorithm is better. How many times would the
experimenter declare statistically signi cant di erences between the di erent
algorithms? We examined this by using two baselines: the random recommender
(RecencyRandom) and Recency2. The results, when using the RecencyRandom
recommender as a baseline, are given in Table 2. Similarly, the results for the
baseline of Recency2 are given in Table 3.</p>
      <p>We see that, when RecencyRandom is used a a baseline, Recency, GeoRec
and GeoRecHistory achieve statistically signi cant performance almost more
than half of the time. With Recency2 as baseline, we see that only Recency
achieves a statistically signi cant result twice.</p>
      <p>What is truly interesting is that the same algorithm (Recency) can end up
achieving statistically signi cant performance over another instance of itself
(Recency2). The two instances of the same algorithm show so big di erence in
performance that there is a big chance of concluding one is better than itself. The
implication of this raises questions on improvement of performance of algorithms
that involve live users in general. Speci cally, to what extent is one algorithm's
improvement over another a real improvement? It seems to us that statistical
signi cance tests alone are not su ciently indicative of performance
improvements. Given the randomness involved in user's clicks on recommendations (and
the system), a certain range of performance di erence is not worth taking
seriously. Studies that involve users must account for some level of randomness, in
addition to statistical signi cance tests.
We set out to study two factors in news recommendation: geographical
information, and user and/or system randomness in news recommendation clicks.
Although the geographical recommendation did not show any strikingly useful
improvement over the recency algorithm, the user-system randomness seems to
indicate that care must be taken to take into account some degree of
randomness in recommender systems evaluation that involve users in a live setting, in
addition to statistical signi cance tests. In the future, we would like to extend
this work to include a better statistical testing system that can account for this
randomness. We also would like to compare the o ine and online evaluations,
Speci cally, we would like run one algorithm on the logs of another to see to
what extent they can replace each other.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Beel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Genzmehr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Langer</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Nurnberger, and</article-title>
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          .
          <article-title>A comparative analysis of o ine and online evaluations and discussion of research paper recommender system evaluation</article-title>
          .
          <source>In Proceedings of the International Workshop on Reproducibility and Replication in Recommender Systems Evaluation</source>
          , pages
          <volume>7</volume>
          {
          <fpage>14</fpage>
          . ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>T.</given-names>
            <surname>Brodt</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          .
          <article-title>Shedding light on a living lab: the clef newsreel open recommendation platform</article-title>
          .
          <source>In Proceedings of the 5th Information Interaction in Context Symposium</source>
          , pages
          <volume>223</volume>
          {
          <fpage>226</fpage>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>F.</given-names>
            <surname>Garcin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Faltings</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Donatsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alazzawi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bruttin</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Huber</surname>
          </string-name>
          .
          <article-title>O ine and online evaluation of news recommender systems at swissinfo. ch</article-title>
          .
          <source>In Proceedings of the 8th ACM Conference on Recommender systems</source>
          , pages
          <volume>169</volume>
          {
          <fpage>176</fpage>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Gebremeskel</surname>
          </string-name>
          and
          <string-name>
            <surname>A. P. de Vries</surname>
          </string-name>
          .
          <article-title>The role of geographic information in news consumption</article-title>
          .
          <source>In Proceedings of the 24th International Conference on World Wide Web Companion. International World Wide Web Conferences Steering Committee</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>F.</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kille</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lommatzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Plumbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Brodt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Heintz</surname>
          </string-name>
          .
          <article-title>Benchmarking news recommendations in a living lab</article-title>
          .
          <source>In Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Interaction, pages
          <volume>250</volume>
          {
          <fpage>267</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>E.</given-names>
            <surname>Kirshenbaum</surname>
          </string-name>
          , G. Forman, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Dugan</surname>
          </string-name>
          .
          <article-title>A live comparison of methods for personalized article recommendation at forbes</article-title>
          .
          <source>com. In Machine Learning and Knowledge Discovery in Databases</source>
          , pages
          <volume>51</volume>
          {
          <fpage>66</fpage>
          . Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S. M.</given-names>
            <surname>McNee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kapoor</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          .
          <article-title>Don't look stupid: avoiding pitfalls when recommending research papers</article-title>
          .
          <source>In Proceedings of the 2006 20th anniversary conference on Computer supported cooperative work</source>
          , pages
          <volume>171</volume>
          {
          <fpage>180</fpage>
          . ACM,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>