<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Recommender Systems Fairness Evaluation via Generalized Cross Entropy∗</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yashar Deldjoo</string-name>
          <email>yashar.deldjoo@poliba.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vito Walter Anelli</string-name>
          <email>vitowalter.anelli@poliba.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hamed Zamani</string-name>
          <email>zamani@umass.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alejandro Bellogín</string-name>
          <email>alejandro.bellogin@uam.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tommaso Di Noia</string-name>
          <email>tommaso.dinoia@poliba.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Autonomous, University of Madrid</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Polytechnic University of Bari</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Massachusetts</institution>
          ,
          <addr-line>Amherst</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Fairness in recommender systems has been considered with respect to sensitive attributes of users (e.g., gender, race) or items (e.g., revenue in a multistakeholder setting). Regardless, the concept has been commonly interpreted as some form of equality - i.e., the degree to which the system is meeting the information needs of all its users in an equal sense. In this paper, we argue that fairness in recommender systems does not necessarily imply equality, but instead it should consider a distribution of resources based on merits and needs. We present a probabilistic framework based on generalized cross entropy to evaluate fairness of recommender systems under this perspective, where we show that the proposed framework is flexible and explanatory by allowing to incorporate domain knowledge (through an ideal fair distribution) that can help to understand which item or user aspects a recommendation algorithm is over- or under-representing. Results on two real-world datasets show the merits of the proposed evaluation framework both in terms of user and item fairness.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION AND CONTEXT</title>
      <p>
        Recommender systems (RS) are widely applied across the modern
Internet, in e-commerce websites, movies and music streaming
platforms, or on social media to point users to items (products or
services) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. For evaluation of RS, accuracy metrics are typically
employed, which measure how much the presented items will be of
interest to the target user. One commonly raised concern is how much
the recommendations produced by RS are fair. For example, do users
of certain gender or race receive fair utility (i.e., benefit) from the
recommendation service? To answer this question, one has to recognize
the multiple stakeholders involved in such systems and that fairness
issues can be studied for more than one group of participants [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In
a job recommendation scenario, for instance, these multiple groups
can be the job seekers and prospective employers where fairness
toward both parties has to be recognized. Moreover, fairness in RS
can be measured towards items or users; in this context, user and
item fairness are commonly associated with an equal chance for
appearing in the recommendation results (items) or receiving results
of the same quality (users). As an example for the latter, an unfair
system may discriminate against users of a particular race or gender.
      </p>
      <p>One common characteristic of the previous literature focusing
on RS fairness evaluation is that fairness has been commonly
interpreted as some form of equality across multiple groups (e.g., gender,
∗Copyright 2019 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).</p>
      <p>
        Presented at the RMSE workshop held in conjunction with the 13th ACM Conference
on Recommender Systems (RecSys), 2019, in Copenhagen, Denmark.
race). For example, Ekstrand et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] studied whether RS produce
equal utility for users of diferent demographic groups. The authors
ifnd demographic diferences in measured efectiveness across two
datasets from diferent domains. Yao and Huang [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] studied various
types of unfairness that can occur in collaborative filtering models
where, to produce fair recommendations, the authors proposed to
penalize algorithms producing disparate distributions of prediction
error. For additional resources see [
        <xref ref-type="bibr" rid="ref17 ref30 ref31 ref8 ref9">8, 9, 17, 30, 31</xref>
        ].
      </p>
      <p>
        Nonetheless, although less common, there are a few works where
fairness has been defined beyond uniformity. For instance, in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the
authors proposed an approach focused on mining the relation
between relevance and attention in Information Retrieval by exploiting
the positional bias of search results. That work promotes the notion
that ranked subjects should receive attention that is proportional to
their worthiness in a given search scenario and achieve fairness of
attention by making exposure proportional to relevance. Similarly, a
framework formulation of fairness constraints is presented in [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] on
rankings in terms of exposure allocation, both with respect to group
fairness constraints and individuals. Another approach where
nonuniform fairness has been used is the work proposed in [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ], where
the authors aim to solve the top-k ranking problem by optimizing
a fair utility function under two conditions: in-group monotonicity
(i.e., rank more relevant items above less relevant within the group)
and group fairness (proportion of protected group items in the top-k
ranking should be above a minimum threshold). In summary, even
though these approaches use some notion of non-uniformity, they
are applied under diferent perspectives and purposes.
      </p>
      <p>In the present work, we argue that fairness does not necessarily
imply equality between groups, but instead proper distribution of utility
(benefits) based on merits and needs. To this end, we present a
probabilistic framework for evaluating RS fairness based on attributes of
any nature (e.g., sensitive or insensitive) for both items or users and
show that the proposed framework is flexible enough to measure
fairness in RS by considering fairness as equality or non-equality among
groups, as specified by the system designer or any other parties
involved in multistakeholder setting. As we shall see later, the discussed
approaches are diferent from our proposal in that we are able to
accommodate diferent notions of fairness, not only ranking, e.g., rating,
ranking and even-beyond accuracy metrics. In fact, the main
advantage of our framework is to provide the system designer with a high
degree of flexibility on defining fairness from multiple viewpoints.
Results on two real-world datasets show the merits of the proposed
evaluation framework, both in terms of user and item fairness.
2</p>
    </sec>
    <sec id="sec-2">
      <title>EVALUATING FAIRNESS IN RS</title>
      <p>
        In this section, we propose a framework based on generalized cross
entropy for evaluating fairness in recommender systems. Let U and
I denote a set of users and items, respectively. Suppose A be a set
of sensitive attributes in which fairness is desired. Each attribute can
and pf1 = [
        <xref ref-type="bibr" rid="ref13 ref23">23 , 13</xref>
        ],pf2 = [
        <xref ref-type="bibr" rid="ref13">13 , 32</xref>
        ] characterize the fair distribution as
uniform or non-uniform distributions between two groups.
      </p>
      <p>Rec 0
Rec 1
Rec 2
be defined for either users, e.g., gender and race, or items, e.g., item
provider (or stakeholder).</p>
      <p>
        The goal is to find an unfairness measure I that produces a
nonnegative real number for a recommender system. A recommender
system M is considered less unfair (i.e., more fair) than M ′ with
respect to the attribute a ∈ A if and only if |I (M,a)| &lt; |I (M ′,a)|.
Previous works have used inequality measures to evaluate algorithmic
unfairness, however, we argue that fairness does not always imply
equality. For instance, let us assume that there are two types of users
in the system – regular (free registration) and premium (paid) – and
the goal is to compute fairness with respect to the users’
subscription type. In this example, it might be more fair to produce better
recommendations for paid users, therefore, equality is not always
equivalent to fairness. We define fairness of a recommender system
as the generalized cross entropy (GCE) for some parameter α , 0,1:
1 ∫
α (1−α )
where p and pf respectively denote the probability distribution of
the system performance and the fair probability distribution, both
with respect to the attribute x = a [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The unfairness measure I is
minimized with respect to attribute x =a when p =pf , meaning that
the performance of the system is equal to the performance of a fair
system. In the next sections, we discuss how to obtain or estimate
these two probability distributions. If the attribute a is discrete or
categorical (as typical attributes, such as gender or race), then the
unfairness measure is defined as:
      </p>
      <p>pfα (x )p(1−α )(x ) dx −1
I (M,a) =
1 Õ</p>
      <p>
α (1−α )  aj



pfα (aj )p(1−α )(aj )−1


2.1</p>
      <p>
        Fair Distribution pf
The definition of a fair distribution pf is problem-specific and should
be determined based on the problem or target scenario in hand. For
example, a job or music recommendation website may want to
ensure that its premium users, who pay for their subscription, would
receive more relevant recommendations. In this case, pf should be
non-uniform across the user classes (premium versus free users). In
(1)
(2)
other scenarios, a uniform definition of pf might be desired.
Generally, when fairness is equivalent to equality, then pf should be
uniform and in that case, the generalized cross entropy would be the
same as generalized entropy (see [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] for more information).
2.2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Estimating Performance Distribution p</title>
      <p>
        The performance distribution p should be estimated based on the
output of the recommender system on a test set. In the following,
we explain how we can compute this distribution for item attributes.
We define the recommendation gain ( rдi ) for each item i as follows:
Õ
rдi =
ϕ(i, RecuK ) д(u,i,r )
(3)
u ∈U
where RecuK is the set of top-K items recommended by the system to
the user u ∈U . ϕ(i, RecuK ) = 1 if item i is present in RecuK ; otherwise
ϕ(i, Recuk ) = 0. The function д(u,i,r ) is the gain of recommending
item i to user u with the rank r . Such gain function can be defined
in diferent ways. In its simplest form, when д(u,i,r ) = 1, the
recommendation gain in Eq. 3 would boil down to recommendation
count (i.e., rдi =rci ). A binary gain in which д(u,i,r ) = 1 when item
i recommended to user u is relevant and д(u,i,r ) = 0 otherwise, is
another simple form of the gain function based on relevance. The gain
functionд can be also defined based on ranking information, i.e.,
recommending relevant items to users in higher ranks is given a higher
gain. In such case, we recommend the discounted cumulative gain
(DCG) function that is widely used in the definition of NDCG [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ],
given by l2orgel2(u(r,i+)−11) where rel(u,i) denotes the relevance label for the
user-item pair u and i. We can further normalize the above formula
based on the ideal DCG for user u to compute the gain function д.
      </p>
      <p>The performance probability distribution p is then proportional
to the recommendation gain for the items associated to an attribute
value aj . Formally, the performance probability p(aj ) used in Eq. (2)
is computed as: p(aj ) = Íi ∈aj rдi /Z where Z is a normalization
factor set equal to Z = Íi rдi to make sure that Íp(aj ) = 1. Under an
analogous formulation, we could define a variation of fairness for
users based on Eq. (3):
rдu =
Õ
ϕ(i, RecuK ) д(u,i,r )
(4)
i ∈I
where in this case, the gain function cannot be reduced to 1,
otherwise, all users would receive the same recommendation gain rдu .
3</p>
    </sec>
    <sec id="sec-4">
      <title>TOY EXAMPLE</title>
      <p>For the illustration of the proposed concept, in Table 1 we provide a
toy example on how our approach for fairness evaluation framework
could be applied in a real recommendation setting. A set of six users
belonging to two groups (each group is associated with an attribute
value a1 (red) or a2 (green)) who are interacting with a set of items
are shown in Table 1. Let us assume the red group represents users
with a regular (free registration) subscription type on an e-commerce
website while the green group represents users with a premium (paid)
subscription type. A set of recommendations produced by diferent
systems (Rec0, Rec1, and Rec2) are shown in the last columns. The
goal is to compute fairness using the proposed fairness evaluation
metric based on GCE given by Eq. (2). The results of evaluation using
three diferent evaluation metrics are shown in Table 2. The metrics
used for the evaluation of fairness and accuracy of the system
include: (i) GCE (absolute value), (ii) Precision@3 and (iii) Recall@3.
Note that GCE = 0 means the system is completely fair, and the closer
the value is to zero, the more fair the respective system is.</p>
      <p>
        By looking at the recommendation results of Rec0, one can note
that if fairness is defined in a uniform way between two groups , defined
through fair distribution pf = [
        <xref ref-type="bibr" rid="ref12 ref12">12 , 12</xref>
        ], then Rec0 is not a completely
fair system, since GCE = 0.08 , 0. In contrast, if fairness is defined as
providing recommendation of higher utility (usefulness) to green users
who are users with paid premium membership type, (e.g., by setting pf
= [
        <xref ref-type="bibr" rid="ref13 ref23">13 , 23</xref>
        ]) then since GCE ≈ 0, we can say that recommendations
produced by Rec0 are fair. Both of the above conclusions are drawn with
respect to attribute “subscription type” (with categories free/paid
premium membership). This is an interesting insight which shows
the evaluation framework is flexible enough to capture fairness based
on the interest of system designer by defining what she considers
as fair recommendation through the definition of pf . While in many
application scenarios we may define fairness as equality among
different classes (e.g., gender, race), in some scenarios (such as those
where the target attribute is not sensitive, e.g., regular v.s. premium
users) fairness may not be equivalent to equality.
      </p>
      <p>Furthermore, by comparing the performance results of Rec1 and
Rec2, we observe that, even though precision and recall improve for
Rec2 and becomes the most accurate recommendation list, it fails
to keep a decent amount of fairness with respect to any parameter
settings of GCE, as in both cases it is outperformed by the other
methods. Moreover, GCE never reaches the optimal value, which
in this case is attributed to the unequal distribution of resources
among classes, since there are more relevant items on green than
red users. This evidences that optimizing an algorithm to produce
relevant recommendations does not necessarily result in more fair
recommendation rather, conversely, a trade-of between the two
evaluation properties can be noticed.
4</p>
    </sec>
    <sec id="sec-5">
      <title>EXPERIMENTS AND RESULTS</title>
      <p>In the section, we discuss our experimental setup and the results.
4.1</p>
    </sec>
    <sec id="sec-6">
      <title>Data Descriptions</title>
      <p>
        We conduct experiments on two real-world datasets, Xing job
recommendation dataset [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Amazon Review dataset [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The datasets
represent diferent item recommendation scenarios for job and
ecommerce domains. We used Xing dataset to study the item-related
notion of fairness, while Amazon is used to study the user-related
notion of fairness.
      </p>
      <p>
        Xing Job Recommendation Dataset (Xing-REC 17): The dataset
was first released by XING as part of the ACM RecSys Challenge
2017 for a job recommendation task [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The dataset contains 320M
of interactions happened in over 3 months. The reason for choosing
this dataset is that it provides several user-related attributes, such as
membership types (regular vs. premium), education degree, and
working country, that can be useful for the study of fairness. For example,
membership type allows us to study the non-equal (non-uniform)
notion of fairness, as a recruiter may want to ensure premium users
obtain better quality in their recommendations.
      </p>
      <p>
        Amazon: We used the toy and games subset which contains 53K
preference scores by 1K users for 24K items, with a sparsity of 99.8%. We
wanted the training set to be as close as possible to an on-line real
scenario in which the recommender system is deployed, with this goal in
mind we used a time-aware splitting. The most rigorous one would be
the fixed-timestamp splitting method [
        <xref ref-type="bibr" rid="ref10 ref18">10, 18</xref>
        ]. In these experiments,
however, we adopted the methodology proposed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] where a single
timestamp is chosen, which represents the moment when test users
are on the platform waiting for recommendations. The training set
corresponds to the past interactions, and the performance is
evaluated with data which correspond to future interactions. The splitting
timestamp is selected to maximize the number of users involved in
the evaluation according to two constraints: the training should
retain at least 15 ratings, and the test set should contain at least 5 ratings.
4.2
      </p>
    </sec>
    <sec id="sec-7">
      <title>Experimental Setup</title>
      <p>
        Two recommendation scenarios are considered to evaluate the
efectiveness of the proposed fairness evaluation framework with respect
to item-centric or user-centric notion of fairness [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Item fairness evaluation: It applies the proposed fairness
evaluation metrics based on GCE on the winner of the ACM RecSys
Challenge 2017. The challenge was formulated as “given a job posting,
recommend a list of candidates that are suitable candidates for the job”.
As such, the user candidates are considered as target items for
recommendation. In order to compute GCE, we used Eq. (3) by considering
a simplified case д(u,i,r ) = 1, in which the recommendation gain rдi
boils down to recommendation count rci for item i, i.e., the number
of times each user appears in the recommendation lists of all jobs.
      </p>
      <p>We compare two recommendation approaches: the winner
submission and a random submission, and evaluate the systems’ fairness
from the perspective of users membership types, education, and
location. As for membership type, premium users (or paid members)
are expected to receive better quality of recommendation.
User fairness evaluation: Here, we experiment with the more
traditional item recommendation task where we study the user fairness
dimension. We consider a scenario where a business owner may
want to ensure superior recommendation quality for its more
engaged users over less engaged (or new) users (or vice versa). In order
to have a more intuitive sense about how fair diferent
recommendation models are recommending to users of diferent classes, we study
the fairness of diferent CF recommendation models with respect to
users’ interactions, defined in 4 categories: (i) very inactive (VIA), (ii)
slightly inactive (SIA), (iii) slightly active (SA), and (iv) very active
(VA). For each user, we compute the score nR (u) that corresponds to
the total number of ratings provided by user u. We group the users
in four groups according to the quartile that this score belongs to.</p>
      <p>
        We have experimented with several recommendation models such
as UserKNN [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], ItemKNN [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] (considering binarized and cosine
similarity metric, Jaccard coeficient [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], and Pearson correlation
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]), SVD++ [
        <xref ref-type="bibr" rid="ref21 ref22">21, 22</xref>
        ], BPRMF [
        <xref ref-type="bibr" rid="ref22 ref24">22, 24</xref>
        ], BPRSlim [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], and two
nonpersonalized models, most-popular and random recommender.
      </p>
      <p>
        For comparison with the proposed GCE metric, we include two
complementary baseline metrics based on the absolute deviation
between the mean ratings of diferent groups as defined in [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]
MAD(R(i), R(j)) = ÍR(i) ÍR(j) where R(i) denotes the predicted
|R(i) | − |R(j) |
ratings for all user-item combinations in group i and R(i) is its
size. Larger values for MAD mean larger diferences between the
groups, interpreted as unfairness. Given that our proposed GCE in
user-fairness evaluation is based on NDCG, we adapt this
definition to also compare between average NDCG for each group. We
refer to these two baselines as MAD-rating and MAD-ranking.
Finally, the reported MAD corresponds to the average MAD
between all the pairwise combinations within the groups involved, i.e.,
MAD = avgi, j (MAD(R(i),R(j))).
4.3
      </p>
    </sec>
    <sec id="sec-8">
      <title>Results and Discussion</title>
      <p>We start our analysis with the results for the item fairness
evaluation as described in Section 4.2, presented in Tables 3 and 4. The
counts in these tables represent the total number of users with a
given category that each submission recommends. We observe in
Table 4 that recommendations produced by the RecSys Challenge
winner performs better with pf0 than with pf2 , since the GCE value is
closer to 0. This evidences that the proposed winner system produces
balanced recommendations across the two membership classes. This
is in contrast to our expectation that premium users should be
provided better recommendations. Therefore, even though the winning
submission could produce higher recommendation quality from a
global perspective, it does not comply with our expectation of a fair
recommendation for this attribute, which is to recommend better
recommendations to premium users.</p>
      <p>Furthermore, in Table 4 we present the recommendation fairness
evaluation results using GCE across two other attributes: Country
and Education; each of these attributes takes 4 categories. We define
ifve variations of the fair distribution pf : while pf0 considers all
attribute categories equally important, the others give one attribute
category a higher importance compared to the rest. After applying
the GCE on the winner submission, we observe that with respect to
the Country attribute, the lowest value of GCE (best case) is produced
for the German companies (GCE = 0.061) while for the Education
attribute the category Unknown (GCE = 0.97) produces the best
outcome, in both cases, these categories are the most frequently
recommended by the analyzed submission. These results show that for a
given target application, if the system is looking for candidates with
certain nationality (in this case, German) or education-level (here
any), the system recommendations coming from the winner
submission are closer to a fair system. In fact, due to the inherent biases in
the dataset, the random submission is obtaining better results
according to our definition of fairness for several of the fair distributions
analyzed. However, it is worth mentioning that if the system designer
wants to promote those users with BSc or PhD, the GCE would show
that the winner submission provides better recommendations to
those users than the random submission. To the best of our
knowledge, this is the first fairness-oriented evaluation metric that allows to
capture these nuances, which as a consequence, helps on
understanding how the recommendation algorithms work on each user group.</p>
      <p>Now Table 5 shows the results for the proposed user fairness
evaluation as described in Section 4.2. We observe in this table that each
recommender obtains a GCE value on a diferent range, an obvious
consequence of the diferent performance obtained in each case for
the diferent groups (as we observe in the NDCG@10 columns for
each user type). For instance, BPRMF is the one found by GCE to
perform in a fair way when assuming uniformity with respect to the
user groups (pf0 ), however, if the system designer aims to promote
those recommenders that provide better suggestions to the most
active users then Random, followed by ItemKNN and SVD++ are the
most fair algorithms.</p>
      <p>Comparing MAD against GCE, we observe that MAD-ranking
produces lower results when NDCGs in each class are close to each
other (e.g., in the case of Random recommender), which corresponds
to the already discussed notion of fairness as equality/uniformity;
similarly, MAD-rating obtains better results for the random
algorithm because, as expected, such method has no inherent bias with
respect to the defined user groups, but also for SVD++, probably
because this recommender tends to predict ratings in a small range. In
both cases, it becomes evident that MAD, in contrast to our proposed
GCE metric, cannot incorporate other definitions of fairness in its
computation, hence, its flexibility is very limited.</p>
      <p>In summary, we have shown that our proposed fairness
evaluation metric is able to unveil whether a recommendation algorithm
satisfies our definition of fairness, where we argue that it should
emphasize a proper distribution of utility based on merits and needs.
We demonstrate this in both notions of fairness: based on users
and based on items. Therefore, we conclude that this metric could
help better explaining the results of the algorithms towards specific
groups of users and items, and as a consequence, it could increase
the transparency of the recommender systems evaluation.
5</p>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSION</title>
      <p>Fairness-aware recommendation research requires appropriate
evaluation metrics to quantify fairness. Furthermore, fairness in RS can
be associated with either items or users, even though this
complementary view has been underrepresented in the literature. In this
work, we have presented a probabilistic framework to measure
fairness of RS under the perspective of users and items. Experimental
results on two real-world datasets show the merits of the proposed
evaluation framework. In particular, one of the key aspects of our
proposed evaluation metric is its transparency and flexibility, since
it allows to incorporate domain knowledge (by means of an ideal fair
distribution) that helps on understanding which item or user aspects
the recommendation algorithms are over- or under-representing.</p>
      <p>
        In the future, we plan to exploit the proposed fairness and
relevance aware evaluation system to build recommender systems that
directly optimize for this objective criterion. Also, it is of our
interest to consider studying various fairness of recommendation under
various content-based filtering or CF models using item content as
side information [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ] on diferent domains (e.g., tourism [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
entertainment [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], social recommendation among others). Finally, we
are considering to investigate the robustness of CF models against
shilling attacks [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] crafted to undermine not only the accuracy of
recommendations but also fairness of these models.
      </p>
    </sec>
    <sec id="sec-10">
      <title>ACKNOWLEDGEMENTS</title>
      <p>This work was supported in part by the Center for Intelligent
Information Retrieval and in part by project TIN2016-80630-P (MINECO).
Any opinions, findings and conclusions or recommendations
expressed in this material are those of the authors and do not
necessarily reflect those of the sponsors.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[1] 2014 (accessed June 5</source>
          ,
          <year>2019</year>
          ).
          <article-title>Amazon product data</article-title>
          . http://jmcauley.ucsd.edu/ data/amazon/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Fabian</given-names>
            <surname>Abel</surname>
          </string-name>
          , Yashar Deldjoo, Mehdi Elahi, and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Kohlsdorf</surname>
          </string-name>
          .
          <year>2017</year>
          . RecSys Challenge 2017:
          <article-title>Ofline and Online Evaluation</article-title>
          .
          <source>In Proc. RecSys (RecSys '17)</source>
          . ACM, New York, NY, USA,
          <fpage>372</fpage>
          -
          <lpage>373</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Jens</given-names>
            <surname>Adamczak</surname>
          </string-name>
          ,
          <string-name>
            <surname>Gerard-Paul Leyson</surname>
            , Peter Knees, Yashar Deldjoo, Farshad Bakhshandegan Moghaddam, Julia Neidhardt, Wolfgang Wörndl, and
            <given-names>Philipp</given-names>
          </string-name>
          <string-name>
            <surname>Monreal</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Session-Based Hotel Recommendations: Challenges and Future Directions</article-title>
          . arXiv preprint arXiv:
          <year>1908</year>
          .
          <volume>00071</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Vito</given-names>
            <surname>Walter</surname>
          </string-name>
          <string-name>
            <surname>Anelli</surname>
          </string-name>
          , Tommaso Di Noia, Eugenio Di Sciascio, Azzurra Ragone, and
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Trotta</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Local Popularity and Time in top-N Recommendation</article-title>
          .
          <source>In Advances in Information Retrieval - 41st European Conference on IR Research</source>
          , ECIR
          <year>2019</year>
          , Cologne, Germany, April 14-
          <issue>18</issue>
          ,
          <year>2019</year>
          , Proceedings, Part I.
          <fpage>861</fpage>
          -
          <lpage>868</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Asia</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Biega</surname>
            ,
            <given-names>Krishna P.</given-names>
          </string-name>
          <string-name>
            <surname>Gummadi</surname>
            , and
            <given-names>Gerhard</given-names>
          </string-name>
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Equity of Attention: Amortizing Individual Fairness in Rankings</article-title>
          .
          <source>In Proc. SIGIR</source>
          , Ann Arbor, MI, USA, July
          <volume>08</volume>
          -
          <issue>12</issue>
          ,
          <year>2018</year>
          . ACM,
          <volume>405</volume>
          -
          <fpage>414</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Zdravko</surname>
            <given-names>I</given-names>
          </string-name>
          <string-name>
            <surname>Botev and Dirk P Kroese</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>The generalized cross entropy method, with applications to probability density estimation</article-title>
          .
          <source>Methodology and Computing in Applied Probability</source>
          <volume>13</volume>
          ,
          <issue>1</issue>
          (
          <year>2011</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>John</surname>
            <given-names>S. Breese</given-names>
          </string-name>
          , David Heckerman, and Carl Myers Kadie.
          <year>1998</year>
          .
          <article-title>Empirical Analysis of Predictive Algorithms for Collaborative Filtering</article-title>
          .
          <source>In UAI '98: Proceedings of the Fourteenth Conference on Uncertainty in Artificial Intelligence</source>
          , University of Wisconsin Business School, Madison, Wisconsin, USA, July
          <volume>24</volume>
          -
          <issue>26</issue>
          ,
          <year>1998</year>
          . Morgan Kaufmann,
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Robin</given-names>
            <surname>Burke</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Multisided fairness for recommendation</article-title>
          .
          <source>arXiv preprint arXiv:1707.00093</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Robin</given-names>
            <surname>Burke</surname>
          </string-name>
          , Nasim Sonboli, and
          <string-name>
            <surname>Aldo</surname>
          </string-name>
          Ordonez-Gauger.
          <year>2018</year>
          .
          <article-title>Balanced Neighborhoods for Multi-sided Fairness in Recommendation</article-title>
          . In Conference on Fairness, Accountability and Transparency,
          <string-name>
            <surname>FAT</surname>
          </string-name>
          <year>2018</year>
          ,
          <volume>23</volume>
          -24
          <source>February</source>
          <year>2018</year>
          , New York, NY,
          <source>USA (Proceedings of Machine Learning Research)</source>
          , Vol.
          <volume>81</volume>
          . PMLR,
          <fpage>202</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Pedro</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Campos</surname>
            , Fernando Díez, and
            <given-names>Iván</given-names>
          </string-name>
          <string-name>
            <surname>Cantador</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Time-aware recommender systems: a comprehensive survey and analysis of existing evaluation protocols</article-title>
          .
          <source>User Model. User-Adapt. Interact</source>
          .
          <volume>24</volume>
          ,
          <issue>1</issue>
          -
          <fpage>2</fpage>
          (
          <year>2014</year>
          ),
          <fpage>67</fpage>
          -
          <lpage>119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Yashar</surname>
            <given-names>Deldjoo</given-names>
          </string-name>
          , Mihai Gabriel Constantin, Hamid Eghbal-Zadeh, Bogdan Ionescu, Markus Schedl, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Audio-visual encoding of multimedia content for enhancing movie recommendations</article-title>
          .
          <source>In Proc. RecSys</source>
          , RecSys
          <year>2018</year>
          , Vancouver, BC, Canada, October 2-
          <issue>7</issue>
          ,
          <year>2018</year>
          . ACM,
          <volume>455</volume>
          -
          <fpage>459</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Yashar</surname>
            <given-names>Deldjoo</given-names>
          </string-name>
          , Maurizio Ferrari Dacrema, Mihai Gabriel Constantin, Hamid Eghbal-zadeh, Stefano Cereda, Markus Schedl, Bogdan Ionescu, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Movie genome: alleviating new item cold start in movie recommendation</article-title>
          .
          <source>User Model. User-Adapt. Interact</source>
          .
          <volume>29</volume>
          ,
          <issue>2</issue>
          (
          <year>2019</year>
          ),
          <fpage>291</fpage>
          -
          <lpage>343</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Yashar</surname>
            <given-names>Deldjoo</given-names>
          </string-name>
          , Tommaso Di Noia, and Felice Antonio Merra.
          <year>2019</year>
          .
          <article-title>Assessing the Impact of a User-Item Collaborative Attack on Class of Users</article-title>
          .
          <source>In Workshop on the Impact of Recommender Systems (ImpactRecSys'19) - 13th ACM Conference of Recommender Systems.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Yashar</surname>
            <given-names>Deldjoo</given-names>
          </string-name>
          , Markus Schedl, Paolo Cremonesi, and
          <string-name>
            <given-names>Gabriella</given-names>
            <surname>Pasi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Content-Based Multimedia Recommendation Systems: Definition and Application Domains</article-title>
          .
          <source>In Proceedings of the 9th Italian Information Retrieval Workshop</source>
          , Rome, Italy, May,
          <fpage>28</fpage>
          -
          <lpage>30</lpage>
          ,
          <year>2018</year>
          .
          <source>(CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2140</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Wei</surname>
            <given-names>Dong</given-names>
          </string-name>
          , Charikar Moses, and
          <string-name>
            <given-names>Kai</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Eficient k-nearest neighbor graph construction for generic similarity measures</article-title>
          .
          <source>In Proceedings of the 20th international conference on World wide web. ACM</source>
          ,
          <volume>577</volume>
          -
          <fpage>586</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Michael</surname>
            <given-names>D</given-names>
          </string-name>
          <string-name>
            <surname>Ekstrand</surname>
          </string-name>
          , Mucun Tian, Ion Madrazo Azpiazu,
          <string-name>
            <surname>Jennifer D Ekstrand</surname>
          </string-name>
          , Oghenemaro Anuyah,
          <string-name>
            <surname>David McNeill</surname>
            ,
            <given-names>and Maria Soledad</given-names>
          </string-name>
          <string-name>
            <surname>Pera</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>All The Cool Kids, How Do They Fit In?: Popularity and Demographic Biases in Recommender Evaluation and Efectiveness</article-title>
          . In Conference on Fairness,
          <source>Accountability and Transparency</source>
          .
          <volume>172</volume>
          -
          <fpage>186</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Golnoosh</surname>
            <given-names>Farnadi</given-names>
          </string-name>
          , Pigi Kouki,
          <string-name>
            <surname>Spencer K. Thompson</surname>
            , Sriram Srinivasan, and
            <given-names>Lise</given-names>
          </string-name>
          <string-name>
            <surname>Getoor</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A Fairness-aware Hybrid Recommender System</article-title>
          . CoRR abs/
          <year>1809</year>
          .09030 (
          <year>2018</year>
          ). arXiv:
          <year>1809</year>
          .09030
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Asela</given-names>
            <surname>Gunawardana</surname>
          </string-name>
          and
          <string-name>
            <given-names>Guy</given-names>
            <surname>Shani</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Evaluating Recommender Systems</article-title>
          .
          <source>In Recommender Systems Handbook</source>
          . Springer,
          <fpage>265</fpage>
          -
          <lpage>308</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Jonathan</surname>
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Herlocker</surname>
          </string-name>
          , Joseph A.
          <string-name>
            <surname>Konstan</surname>
            ,
            <given-names>and John</given-names>
          </string-name>
          <string-name>
            <surname>Riedl</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>An Empirical Analysis of Design Choices in Neighborhood-Based Collaborative Filtering Algorithms</article-title>
          .
          <source>Inf. Retr. 5</source>
          ,
          <issue>4</issue>
          (
          <year>2002</year>
          ),
          <fpage>287</fpage>
          -
          <lpage>310</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Kalervo</given-names>
            <surname>Järvelin</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jaana</given-names>
            <surname>Kekäläinen</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Cumulated gain-based evaluation of IR techniques</article-title>
          .
          <source>ACM Transactions on Information Systems (TOIS) 20</source>
          ,
          <issue>4</issue>
          (
          <year>2002</year>
          ),
          <fpage>422</fpage>
          -
          <lpage>446</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Yehuda</given-names>
            <surname>Koren</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Factorization meets the neighborhood: a multifaceted collaborative filtering model</article-title>
          .
          <source>In Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , Las Vegas, Nevada, USA,
          <year>August</year>
          24-
          <issue>27</issue>
          ,
          <year>2008</year>
          . ACM,
          <volume>426</volume>
          -
          <fpage>434</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Yehuda</surname>
            <given-names>Koren</given-names>
          </string-name>
          , Robert Bell, and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Volinsky</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Matrix Factorization Techniques for Recommender Systems</article-title>
          .
          <source>Computer 42</source>
          ,
          <issue>8</issue>
          (
          <year>2009</year>
          ),
          <fpage>30</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Xia</given-names>
            <surname>Ning</surname>
          </string-name>
          and
          <string-name>
            <given-names>George</given-names>
            <surname>Karypis</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>SLIM: Sparse Linear Methods for Top-N Recommender Systems</article-title>
          .
          <source>In 11th IEEE International Conference on Data Mining, ICDM</source>
          <year>2011</year>
          , Vancouver, BC, Canada,
          <source>December 11-14</source>
          ,
          <year>2011</year>
          . IEEE Computer Society,
          <fpage>497</fpage>
          -
          <lpage>506</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Stefen</surname>
            <given-names>Rendle</given-names>
          </string-name>
          , Christoph Freudenthaler, Zeno Gantner, and
          <string-name>
            <surname>Lars</surname>
          </string-name>
          Schmidt-Thieme.
          <year>2009</year>
          .
          <article-title>BPR: Bayesian Personalized Ranking from Implicit Feedback</article-title>
          .
          <source>In UAI 2009, Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence</source>
          , Montreal, QC, Canada, June 18-21,
          <year>2009</year>
          . AUAI Press,
          <fpage>452</fpage>
          -
          <lpage>461</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Badrul</surname>
            <given-names>Sarwar</given-names>
          </string-name>
          , George Karypis, Joseph Konstan,
          <string-name>
            <given-names>and John</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Analysis of recommendation algorithms for e-commerce</article-title>
          .
          <source>In Proceedings of the 2nd ACM conference on Electronic commerce. ACM</source>
          ,
          <volume>158</volume>
          -
          <fpage>167</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Ashudeep</given-names>
            <surname>Singh</surname>
          </string-name>
          and
          <string-name>
            <given-names>Thorsten</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Fairness of Exposure in Rankings</article-title>
          .
          <source>In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining, KDD</source>
          <year>2018</year>
          , London, UK,
          <year>August</year>
          19-
          <issue>23</issue>
          ,
          <year>2018</year>
          . ACM,
          <volume>2219</volume>
          -
          <fpage>2228</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Till</surname>
            <given-names>Speicher</given-names>
          </string-name>
          , Hoda Heidari, Nina Grgic-Hlaca, Krishna P Gummadi, Adish Singla, Adrian Weller, and Muhammad Bilal Zafar.
          <year>2018</year>
          .
          <article-title>A Unified Approach to Quantifying Algorithmic Unfairness: Measuring Individual &amp;Group Unfairness via Inequality Indices</article-title>
          .
          <source>In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining. ACM</source>
          ,
          <volume>2239</volume>
          -
          <fpage>2248</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Sirui</given-names>
            <surname>Yao</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bert</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Beyond parity: Fairness objectives for collaborative ifltering</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          .
          <volume>2921</volume>
          -
          <fpage>2930</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Meike</surname>
            <given-names>Zehlike</given-names>
          </string-name>
          , Francesco Bonchi, Carlos Castillo, Sara Hajian, Mohamed Megahed, and
          <string-name>
            <surname>Ricardo</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Baeza-Yates</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>FA*IR: A Fair Top-k Ranking Algorithm</article-title>
          .
          <source>In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management</source>
          ,
          <string-name>
            <surname>CIKM</surname>
          </string-name>
          <year>2017</year>
          , Singapore,
          <source>November 06 - 10</source>
          ,
          <year>2017</year>
          . ACM,
          <volume>1569</volume>
          -
          <fpage>1578</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Yong</surname>
            <given-names>Zheng</given-names>
          </string-name>
          , Tanaya Dave, Neha Mishra, and
          <string-name>
            <given-names>Harshit</given-names>
            <surname>Kumar</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Fairness In Reciprocal Recommendations: A Speed-Dating Study</article-title>
          .
          <source>In Adjunct Publication of the 26th Conference on User Modeling, Adaptation and Personalization</source>
          ,
          <string-name>
            <surname>UMAP</surname>
          </string-name>
          <year>2018</year>
          , Singapore,
          <source>July 08-11</source>
          ,
          <year>2018</year>
          . ACM,
          <volume>29</volume>
          -
          <fpage>34</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Ziwei</surname>
            <given-names>Zhu</given-names>
          </string-name>
          , Xia Hu, and
          <string-name>
            <given-names>James</given-names>
            <surname>Caverlee</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Fairness-Aware Tensor-Based Recommendation</article-title>
          .
          <source>In Proc. CIKM</source>
          , CIKM
          <year>2018</year>
          , Torino, Italy,
          <source>October 22-26</source>
          ,
          <year>2018</year>
          . ACM,
          <volume>1153</volume>
          -
          <fpage>1162</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>