<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Recommender Systems Evaluations: O Online, Time and A/A Test ine,</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gebrekirstos G. Gebremeskel</string-name>
          <email>gebre@cwi.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arjen P. de Vries</string-name>
          <email>arjen@acm.org</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Access, CWI</institution>
          ,
          <addr-line>Amsterdam</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Radboud University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present a comparison of recommender systems algorithms along four dimensions. The rst dimension is o ine evaluation where we compare the performance of our algorithms in an o ine setting. The second dimension is online evaluation where we deploy recommender algorithms online with a view to comparing their performance patterns. The third dimension is time, where we compare our algorithms in two di erent years: 2015 and 2016. The fourth dimension is the quanti cation of the e ect of non-algorithmic factors on the performance of an online recommender system by using an A/A test. We then analyze the performance similarities and di erences along these dimensions in an attempt to draw meaningful patterns and conclusions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Recommender systems can be evaluated o ine or online. The purpose of
recommender system evaluation is to select algorithms for use in a production setting.
O ine evaluations test the e ectiveness of recommender system algorithms on
a certain dataset. Online evaluation attempts to evaluate recommender systems
by a method called A/B testing where a part of users are served by recommender
system A and the another part of users by recommender system B. The
recommender system that achieves a higher score according to a chosen metric ( for
example, Click-Through-Rate) is chosen as a better recommender system, given
other factors such as latency and complexity are comparable.</p>
      <p>The purpose of o ine evaluation is to select recommender systems for
deployment online. O ine evaluations are easier and reproducible. But do o ine
evaluations predict online performance behaviors and trends? Do the absolute
performances of algorithms o ine hold online too? Do the relative rankings of
algorithms according to o ine evaluation hold online too? How do o ine
evaluation compare and contrast with online evaluations?</p>
      <p>
        CLEF NewsREEL [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], a campaign-like news recommendation evaluation,
provides opportunities to investigate recommender system performance from several
angles. CLEF NewsREEL 2016 campaign, in particular, is focused on comparing
recommender system performance in online and o ine settings [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. CLEF
NewsREEL 2016 provides two tasks: Benchmark News Recommendations in a Living
Lab (Task 1) which enables evaluation of systems in a production setting [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
and Benchmarking News Recommendations in a Simulated Environment (Task
2) which enables the evaluation of systems in a simulated (o ine) setting using
dataset collected from the online interactions.
      </p>
      <p>In 2015, we participated in Task 1. In 2016, we participated in both CLEF
NewsREEL tasks. In this working notes, we report both o ine and online
evaluations and how they relate to each other. We also present the challenges of online
evaluation from the dimensions of time and non-algorithmic causes of
performance di erences. On the time dimension, we speci cally investigate online
performances behaviors in 2015 and 2016, and on the dimension of non-algorithmic
causes of performance di erences, we employ an A/A test where we run two
instances of the same algorithm to gauge the extent of performance di erence
resulting from non-algorithmic causes.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Tasks and Objectives</title>
      <p>
        The objective of our participation this year is to investigate performance
behaviors of recommender system along several dimensions. We are interested in
di erences and similarities in o ine and online performances, the variations in
performance over time, and the estimation of performance di erences caused
by non-algorithmic factors in the online recommender system evaluations. This
work can be see as an extension of studies that have previously investigated the
di erences between o ine and online recommender system evaluations [
        <xref ref-type="bibr" rid="ref1 ref11 ref2">2, 1, 11</xref>
        ].
      </p>
      <p>
        In 2015, we participated in CLEF NewsREEL News Recommendations
Evaluation, the task of Benchmark News Recommendations in a Living Lab [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
We reported the presence of substantial non-algorithmic factors that cause two
instances of the same algorithm to end up having statistically signi cant
performance di erences. The results are presented in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This year, we run four of
our 2015 recommender systems without change. This allows us to compare the
performance of the systems in 2015 and 2016. In 2016, we participated also in
Task 2, which allows us to evaluate the recommender systems in a simulated
environment and then compare the o ine performance measurements with the
corresponding online performance measurements. In this report, we present the
results of these evaluations along the four dimensions and highlight similarities,
di erences and patterns or the lack thereof.
      </p>
      <p>For the study of the e ect of non-algorithmic factors on online recommender
system performances, we run two instances of the same news recommender
algorithm with the view to quantifying the extent of performance di erences. To
compare the online and o ine performance behaviors, we conduct o ine
evaluations on a snapshot of a dataset collected from the same system. To investigate
performance in the dimension of time, we rerun last year's recommender
systems. This means that we can compare the performance of the recommender
systems in 2016 with their corresponding performance in 2015.</p>
      <p>
        The four recommender systems are two instances of Recency, one instance
of GeoRec and one instance of RecencyRandom. Recency keeps the 100
most recently viewed items for each publisher, and upon recommendation
request, the most recently read (clicked) are recommended. GeoRec is a modi
cation of the Recency recommender to diversify the recommendations by taking
into account the users' geographic context and estimated interest in local news.
RecencyRandom recommends items randomly selected from the 100 most
recently viewed items. For a detailed description of the algorithms, refer to [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussions</title>
      <p>We present the results and analysis from the di erent dimensions here. In 2015,
the recommender systems ran from 2015-04-12 to 2015-07-06, a total of 86 days.
RecencyRandom started 12 days later in 2015. In 2016, the systems ran from
2016-02-22 to 2016-05-21, a total of 70 days. We present three types of results for
2016: the daily performances, the incremental performances, and the cumulative
performances. The plot for the daily performance is presented in Figure 1. From
the plot, we observe large variations between the maximum and minimum
performance measurements of the tested recommender systems; the minimum value
equals 0 in every test, while the maximum varies between 12:5% for Recency2,
5:6% for GeoRec, 4:3% for RecencyRandom, and 4:2% for Recency. The highest
performance measurements all occurred between the 18th day from the start of
our participation (2016-03-10) and the 31st day. The highest scores of Recency2,
and GeoRec occurred on March 21nd, for RecencyRandom on March 20th and
GeoRec on March 10th. We do not have a plausible explanation why the
evaluation resulted in increased performance during that period, nor why the highest
scores for two systems occurred on those two days in March. We did however
observe quite a reduction in the number of recommendation requests issued in
the period when the systems showed increased performance scores (such that
minor variations would lead to larger normalized performance di erences than in
the rest of the evaluation period). Some of the systems have no reported results
between 2016-03-24 and 2016-04-05; if this reduction in the number of
recommendation requests is the same for all teams and systems who participated, the
lower number of recommendations paired by an increase in CTR could indicate
that users are more likely to click on recommendations when recommendations
are o ered sparsely. If this is the case, it might suggest further investigation into
the relationship of the number of recommendations and user responses.</p>
      <p>Figure 2 for 2015 and in Figure 3 for 2016 plot the performance measurements
as the systems progress on a daily basis, which we call incremental performance.
The cumulative number of requests, clicks and CTR scores of the systems in both
years are presented in Table 1. The cumulative performance measurements
remain below 1% for all systems. The maximum performance di erences observed
between the systems equal 0:16% in 2015 and 0:07% in 2016.</p>
      <p>From the plots in Figure 2 and Figure 3 and the cumulative performance
measurements in Table 1, we observe that the performance measurements of the
di erent systems vary. Are these performance variations between the di erent
systems also statistically signi cant? We look at statistical signi cance on a</p>
      <sec id="sec-3-1">
        <title>Days</title>
        <p>6
.
0
8
.
0
6
.
0</p>
      </sec>
      <sec id="sec-3-2">
        <title>Legend</title>
      </sec>
      <sec id="sec-3-3">
        <title>Recency</title>
      </sec>
      <sec id="sec-3-4">
        <title>Recency2</title>
      </sec>
      <sec id="sec-3-5">
        <title>RecencyRandom</title>
        <p>GeoRec
0
10
20
30
40
50
60
70</p>
      </sec>
      <sec id="sec-3-6">
        <title>Days</title>
        <p>
          daily basis after the 14th day, which is considered the average time within which
industry A/B tests are conducted. To compute statistical signi cance, we used
Python module of Thumbtack's Abba, a test for binomial experiments [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] (for a
description of the implementation, please refer to [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ])
        </p>
        <p>We perform statistical signi cance tests on a daily basis to simulate the
notion of an experimenter checking whether one system is better than the other
at the end of every day. In testing for statistical signi cance on a daily basis,
we seek an answer to the question: `On how many days would an experimenter
seeking to select the better system nd out that one system is signi cantly
di erent from the chosen baseline?' We investigate this under two baselines:
Recency2, and RecencyRandom. Tables 2 and 3 present the actual number of
days and the percentage of days on which signi cant performance di erences
were observed.</p>
        <p>Next, we looked into the error noti cations received by our recommender
systems in the 2016 period. The error types and counts for each system are
presented in Table 4. Three types of errors occurred, the highest number for the
RecencyRandom recommender. According to the ORP documentation3, error
code 408 corresponds to connection timeouts, error code 442 to invalid format
of recommendation responses, while error 455 is not described.</p>
        <p>We aggregated error messages by day. Out of the 70 days, Recency received
error noti cations on 16 days, Recency2 on 19 days, GeoRec on 24 days, and
RecencyRandom on 51 days. All systems received high number of error messages
on speci c days, especially on 2016-04-07 and 2016-04-08. While we do not know
the explanation for errors on especially those days, we did observe that most of
the high-error days seem to be those that correspond to the beginning of the
start of the systems, or at the beginning of a change of load (from low to high).
3 http://orp.plista.com/documentation/download
Why the RecencyRandom recommender received a high number of `invalidly
formatted' responses is not clear, because the format is the same as for other
systems. The main di erence between RecencyRandom and the other systems
we deployed is that it has a lower response time, and we suspect the high number
of errors to be related to its lower response rate.
In the online evaluation, or Benchmark News Recommendations in a Living Lab
(Task 1) as it is called in CLEF NewsREEL, we investigate recommender systems
in two dimensions. One dimension is time where we compare and contrast the
performances of our systems in 2015 and 2016. The second dimension is an A/A
test where we attempt to study non-algorithmic e ects on the performance of
systems. Each of the dimensions are discussed in the following subsections.
Time Dimension: Performances in 2015 and 2016 Participation in the
Lab in 2015 and in 2016 gives us the opportunity to study the evaluation results
from a time dimension. We compare the systems both in terms of their absolute
and relative performance (in terms of their rankings). To compare the absolute
performances, we used the 2015 instances of the recommender systems as
baselines, and the corresponding 2016 instances as alternatives. The performance
measurements of the Recency and GeoRec instances of 2016 were signi cantly
di erent from the performance measurements of the Recency and GeoRec
instances of 2015 with a P-values of 0 :0001 and 0 :0009 respectively. The 2015
instances of Recency2 and RecencyRandom were not signi cantly di erent from
their corresponding instances in 2016.</p>
        <p>In 2015, Recency2 ranked third, but in 2016, it ranked rst. In 2015, almost
all systems started from a lower CTR performance, and slowly increased towards
the end where the performance measurements stabilized (see Figure 3). In 2016,
however, the evaluation results of the systems reached its high at the beginning,
and then decreased steadily towards the end, except for recommender Recency2,
which showed an increase after the rst half of its deployment and then decreased
(see Figure 3). In 2016, the performance measurements seemed to continue to
decrease, and not converge to a stable result like in 2015.</p>
        <p>When we compare the number of days for which the results are signi cantly
di erent according to the statistical test (see Table 2 and Table 3), we observe
that there is no consistency. In 2015, there were two days (2.7%) on which
signi cant performance di erences were observed between Recency and Recency2
while there are 25 days (34.3%) on which signi cant performance di erence
between GeoRec and Recency2. In 2016, Recency has shown 47.4% of the time
signi cant performance, and GeoRec only 14%. When using RecencyRandom as
a baseline, Recency has registered signi cant performance di erences 27.4% of
the time in 2015, and 0% in 2016. GeoRec has 56.2% in 2015 and 8.8% in 2016.4</p>
        <p>We conclude that it is di erent to generalize the performance measurements
over time. The patterns observed in 2015 and 2016 vary widely, both in terms of
absolute and relative performance, irrespective of the baseline considered. The
implication is that one can not rely on the absolute and relative rankings of
recommender systems at one time for a similar job in another time. The systems
have not changed between the two evaluations. The di erences in evaluation
results, therefore, can only be attributed to the setting in which the systems are
deployed. It is possible that the presentation of recommendation items by the
publishers, the users and content of the news publishers might have undergone
changes which can then a ect the performances in the two years, but we cannot
be certain without more in-depth analysis.</p>
        <p>A/A Testing In both 2015 and 2016, two of our systems were instances of the
same algorithm. The two instances were run from the same computer; the only
di erences between them were the port numbers by which they communicated
with CLEF NewsREEL's ORP5. The purpose of running two instances of the
same algorithm is to quantify the level of performance di erences due to
nonalgorithmic causes. From the participant's perspective, performance variation
between two instances of the same algorithm can be seen pure luck. The extent
of performance di erences between the instances can be seen as also happening
between the performances of the other systems. We can consider that he
performance di erence due to the e ectiveness of the algorithms is therefore the
overall performance minus the maximum performance di erence between the
performances of the two instances.</p>
        <p>
          The results of the two instances (Recency and recency2) can be seen in 1,
the incremental plots (Figure 2 and Figure 3. The cumulative performances on
the 86th day of the deployment in 2015 showed no signi cant di erence. In 2016,
however, Recency2 showed a signi cant performance over Recency with a
Pvalue of 0 :0005 . Checking for statistical signi cance on a daily basis after the
14th day (see 2), in 2015, there were 2 days (2.7%) on which the two instances
4 We would like to mention a correction here over the reported statistical signi cance
score of GeoRec in 2015. It was reported that Georec did not achieve any signi cant
performance over Recency2 [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], which was an error in calculation.
5 http://orp.plista.com/
di ered signi cantly. In 2016, however, the number of days was extremely higher,
a total of 27 days (47.4%). This is interesting for two reasons: 1) the fact that
two instances can end up having statistically signi cant performance di erences
and 2) that the signi cant di erence occurred. In 2016, one instance achieved
signi cant performance di erences over the other instance for almost half of the
time.
3.2
        </p>
        <p>O ine Evaluation
We present evaluations conducted o ine, or in Benchmarking News
Recommendations in a Simulated Environment (Task 2), as it is called in CLEF NewsREEL.
Evaluation in Task 2 di ers from other o ine evaluation setups in that Task 2
actually simulates the online evaluation setting for each of the systems.
Usually, systems are selected on the basis of o ine evaluation and deployed online.
Other things such as complexity and latency being equal, there is this implicit
assumption that the relative o ine performances of systems holds online too.
That is that if System one has performed better than system two in an o ine
evaluation, it is assumed that the same rank holds when the two algorithms are
deployed online. In this section, we investigate whether this assumption holds by
comparing the o ine evaluation results of the algorithms in Task 1 with their
online results.</p>
        <p>
          Task 2 of CLEF NewsREEL provides a reproducible environment for
participants to evaluate their algorithms in a simulated environment that uses user-item
interaction dataset recorded from the online interactions [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. In the simulated
environment, a recommendation is successful if the user has viewed or clicked
on the recommendations. This is di erent from Task 1 (online evaluation) where
a recommendation is a success only if the recommendation is clicked. The
performances of our algorithms in the simulated evaluation are presented in Table
5. The plots as they progress on a daily basis are presented in Figure 4. In this
evaluation, Recency leads followed by Georec and then RecencyRandom. Using
RecencyRandom as a baseline, there was no signi cant performance di erence
in both Recency and GeoRec. Comparing the ranking with those of the systems
in Task 2, there is no consistency. We conclude that the relative o ine
performance measurements do not generalize to those online, much less the absolute
performance.
        </p>
        <p>From Table 5, we observe that only RecencyRandom has invalid responses.
We also observed that RecencyRandom has higher error messages and lower
performance in Task 1. To understand why, we looked at the response times of the
systems under extreme load. The mean, min, max and standard deviations of the
response times of the three systems are presented in Table 6. We observe that
RecencyRandom has the slowest response time followed by GeoRec. We have
also plotted the number of recommendations within 250 milliseconds in Figure
5. Here too, we observe the lowest reponse times for RecencyRandom (attributed
to the randomization before selecting recommendation items). When we look at
the publisher-level breakdown of the recommendation response in Table 5, we see
that RecencyRandom has invalid responses for two publishers, but for publisher
Tagesspiegel (1677), all its recommendations are invalid. In the o ine
evaluation, invalid response means that the response generates an exception during
parsing. We looked into the recommendation responses of RecencyRandom, and
compared the response for publisher 694 and 1677. Almost all item responses
for publisher 1677 were empty, which we assume to be related with the extreme
load.
Our systems are very similar to each other, in that they are slight modi cations
of each other. This means that it is expected that their performances do not vary
much. We have analyzed the performance of our systems from the dimensions of
online, o ine, and time. We have also investigated the the extent of performance
di erence due to non-algorithmic causes in online evaluation by running two
instances of the same algorithms.</p>
        <p>We have observed substantial variation along the four dimensions. The
performance measurements in both absolute and relative sense varied signi cantly
in 2015 and in 2016. More surprisingly, the two instances of the same algorithm
varied signi cantly both in the two years and within the same year. This is
surprising and indicates how challenging it is to evaluate algorithms online. In
the online evaluation, non-algorithmic and non-functional factors impact
performance measurements. Non-algorithmic factors include variations in users and
items that systems deal with, and the variations in recommendation requests.
Non-functional factors include response times and network problems. The
performance di erence between the two instances of the same algorithms can be
considered to re ect the impact of non-algorithmic and non-functional factors
on performance. It can then be subtracted from the performances of online
algorithms before they are compared with baselines and each other. This can be
seen as a way of discounting the randomness in online system evaluation from
a ecting comparisons.</p>
        <p>The implication of the lack of pattern in the performance of the systems
across time and baselines, and more specially the performance di erences
between the two instances of the same algorithm highlights the challenge of
comparing systems online on the basis of statistical signi cance tests alone. The
results call for caution in the comparison of systems online where user-item
dynamism, operational decision choices and non-functional factors all play roles
s
e
s
n
o
p
s
e
R
.
o
N
0
0
0
0
6
0
0
0
0
4
0
0
0
0
2
0
0
50
100
150</p>
        <p>200</p>
      </sec>
      <sec id="sec-3-7">
        <title>Bins</title>
        <p>in causing performance di erences that are not due to the e ectiveness of the
algorithms.
Let us also compare evaluation results for our systems to those of other teams
that participated in 2016 CLEF NewsREEL's Task 1, based on the results over
the period between 28 April and 20 May (provided by CLEF NewsREEL). The
plot of the team ranking as provided by CLEF NewsREEL is provided in Figure
6. We examined whether the performance of the best performing systems from
the teams that are ranked above us were signi cantly di erent from ours. Only
the ABC's and Arti cial Intelligence's systems were signi cantly di erent from
Recency2 (our best performing system for 2016).
2500
2000
1500
1000
500
0 0</p>
        <p>NewsREEL 2016 | Results (28 April to 20 May)
We set out to investigate the performance of recommender system algorithms
online, o ine, and in two separate periods. The recommender systems'
performances in di erent dimensions indicate that there is no consistency. The o ine
performances were not predictive of the online performances in both absolute
and relative sense. Also the performance measurements of the systems in 2015
were not predictive of those in 2016, both in relative and absolute sense. Our
systems are slight variations of the same algorithm, and yet the performances
varied in all dimensions. We conclude that we should be cautious in
interpreting the results of performance di erences, especially considering the di erences
observed between the two instances of the same algorithm.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>This research was partially supported by COMMIT project In niti.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Beel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Genzmehr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Langer</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Nurnberger, and</article-title>
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          .
          <article-title>A comparative analysis of o ine and online evaluations and discussion of research paper recommender system evaluation</article-title>
          .
          <source>In Proceedings of the International Workshop on Reproducibility and Replication in Recommender Systems Evaluation</source>
          , pages
          <volume>7</volume>
          {
          <fpage>14</fpage>
          . ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>F.</given-names>
            <surname>Garcin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Faltings</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Donatsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alazzawi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bruttin</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Huber</surname>
          </string-name>
          .
          <article-title>O ine and online evaluation of news recommender systems at swissinfo. ch</article-title>
          .
          <source>In Proceedings of the 8th ACM Conference on Recommender systems</source>
          , pages
          <volume>169</volume>
          {
          <fpage>176</fpage>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>G.</given-names>
            <surname>Gebremeskel</surname>
          </string-name>
          and
          <string-name>
            <surname>A. P. de Vries</surname>
          </string-name>
          .
          <article-title>The degree of randomness in a live recommender systems evaluation</article-title>
          .
          <source>In Working Notes for CLEF 2015 Conference</source>
          , Toulouse, France. CEUR,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Gebremeskel</surname>
          </string-name>
          and
          <string-name>
            <surname>A. P. de Vries</surname>
          </string-name>
          .
          <article-title>Random performance di erences between online recommender system algorithms</article-title>
          . In N. Fuhr,
          <string-name>
            <given-names>P.</given-names>
            <surname>Quaresma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Larsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Goncalves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Balog</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cappellato</surname>
          </string-name>
          , and N. Ferro, editors,
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction 7th International Conference of the CLEF Association</source>
          ,
          <article-title>CLEF 2016, vora</article-title>
          ,
          <source>Portugal, September 5-8</source>
          ,
          <year>2016</year>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>F.</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Brodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Seiler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kille</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lommatzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Turrin</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Sereny</surname>
          </string-name>
          .
          <article-title>Benchmarking news recommendations: the clef newsreel use case</article-title>
          .
          <source>In SIGIR Forum</source>
          , volume
          <volume>49</volume>
          , pages
          <fpage>129</fpage>
          {
          <fpage>136</fpage>
          . ACM Special Interest Group,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>F.</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kille</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lommatzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Plumbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Brodt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Heintz</surname>
          </string-name>
          .
          <article-title>Benchmarking news recommendations in a living lab</article-title>
          .
          <source>In Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Interaction, pages
          <volume>250</volume>
          {
          <fpage>267</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S.</given-names>
            <surname>Howard</surname>
          </string-name>
          .
          <source>Abba 0.1</source>
          .0. https://pypi.python.org/pypi/ABBA/0.1.0. Accessed:
          <fpage>2016</fpage>
          -06-18.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>S.</given-names>
            <surname>Howard</surname>
          </string-name>
          . Abba:
          <article-title>Frequently asked questions</article-title>
          . https://www.thumbtack.com/labs/abba/. Accessed:
          <fpage>2016</fpage>
          -06-18.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>B.</given-names>
            <surname>Kille</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lommatzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Gebremeskel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Seiler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Malagoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sereny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Brodt</surname>
          </string-name>
          , and A. de Vries.
          <article-title>Overview of NewsREEL'16: Multi-dimensional Evaluation of Real-Time Stream-Recommendation Algorithms</article-title>
          . In N. Fuhr,
          <string-name>
            <given-names>P.</given-names>
            <surname>Quaresma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Larsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Goncalves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Balog</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macdonald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cappellato</surname>
          </string-name>
          , and N. Ferro, editors,
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction 7th International Conference of the CLEF Association</source>
          ,
          <article-title>CLEF 2016, vora</article-title>
          ,
          <source>Portugal, September 5-8</source>
          ,
          <year>2016</year>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>B.</given-names>
            <surname>Kille</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lommatzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Turrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sereny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Brodt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Seiler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          .
          <article-title>Stream-based recommendations: Online and o ine evaluation as a service</article-title>
          .
          <source>In International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , pages
          <volume>497</volume>
          {
          <fpage>517</fpage>
          . Springer,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. E. Kirshenbaum, G. Forman, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Dugan</surname>
          </string-name>
          .
          <article-title>A live comparison of methods for personalized article recommendation at forbes</article-title>
          .
          <source>com. In Machine Learning and Knowledge Discovery in Databases</source>
          , pages
          <volume>51</volume>
          {
          <fpage>66</fpage>
          . Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>