<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Empirical Analysis of Collaborative Recommender Systems Robustness to Shilling Attacks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anu Shrestha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesca Spezzano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Soledad Pera</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Boise State University</institution>
          ,
          <addr-line>Boise, ID</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <fpage>45</fpage>
      <lpage>57</lpage>
      <abstract>
        <p>Recommender systems play an essential role in our digital society as they suggest products to purchase, restaurants to visit, and even resources to support education. Recommender systems based on collaborative filtering are the most popular among the ones used in e-commerce platforms to improve user experience. Given the collaborative environment, these recommenders are more vulnerable to shilling attacks, i.e., malicious users creating fake profiles to provide fraudulent reviews, which are deliberately written to sound authentic and aim to manipulate the recommender system to promote or demote target products or simply to sabotage the system. Therefore, understanding the efects of shilling attacks and the robustness of recommender systems have gained massive attention. However, empirical analysis thus far has assessed the robustness of recommender systems via simulated attacks, and there is a lack of evidence on what is the impact of fraudulent reviews in a real-world setting. In this paper, we present the results of an extensive analysis conducted on multiple real-world datasets from diferent domains to quantify the efect of shilling attacks on recommender systems. We focus on the performance of various well-known collaborative filtering-based algorithms and their robustness to diferent types of users. Trends emerging from our analysis unveil that, in the presence of spammers, recommender systems are not uniformly robust for all types of benign users.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Recommender system</kwd>
        <kwd>shilling attack</kwd>
        <kwd>robustness</kwd>
        <kwd>fraudulent reviews</kwd>
        <kwd>Collaborative filtering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The efect of shilling attacks on recommender systems, where malicious users create fake
profiles so that they can then manipulate algorithms by providing fake reviews or ratings, have
been long studied [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. So far, recommender system researchers have: (1) Characterized
and modeled recommender system shilling attacks (where malicious users insert fake profiles
to manipulate recommendations), (2) Defined new metrics to quantify the impacts of these
attacks on commonly used recommender systems, and (3) Applied a detect + filtering approach
to mitigate the efects of spammers on recommendations. Nevertheless, we observe from the
literature that the analysis thus far has focused on assessing the robustness of recommender
systems via simulated attacks [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. Unfortunately, there is a lack of evidence on what is the
impact of fake reviews or fake ratings in a real-world setting.
      </p>
      <p>In this paper, we present an analysis conducted to understand the influence of fraudulent
reviews on the recommendation process in real-world scenarios. We do this through a study of
known datasets with gold standards in diferent domains and several commonly-used
recommendation algorithms. Specifically, we utilize data from two widely-used e-commerce platforms,
Yelp! and Amazon. Among various recommendation algorithms, we consider collaborative
ifltering-based approaches as they are the most eficient and popular recommenders in such
platforms. Thus, we focused our exploration on the robustness of these algorithms to shilling
attacks.</p>
      <p>The main contribution of this paper is two-fold. First, we analyze the performance of widely
used five collaborative filtering-based algorithms in the presence of spammers and compared
them when spammers are removed. By doing so, we seek to answer whether shilling attacks
afect the robustness of the considered recommender algorithms. Second, we investigate if there
is a specific user group (non-mainstream users) that are afected more than others (mainstream
users) by spammers.</p>
      <p>Our results are validated by an empirical evaluation using classical measures for evaluating
predictive and top-N recommendation strategies. We show that RMSE scores decrease and
NDCG@5 scores increase when removing spammers in the majority of the considered algorithms
and datasets. This serves as an indication that the performances of considered
collaborativeifltering-based recommender algorithms are indeed afected by spam ratings/reviews. Further, a
deep investigation to quantify the efects of spammers on recommendations received by certain
groups of users lead us to conclude that, for the Yelp! datasets removing spammers improves
the predictive ability (RMSE) of all the considered recommender algorithms regardless of the
type of users, i.e., mainstream or non-mainstream. However, in the case of Amazon datasets,
we observed a trend where removing spammers lessen the predictive ability for mainstream
users based on RMSE whereas improves for non-mainstream users according to both RMSE
and NDCG@5. Therefore, non-mainstream users whose rating behavior does not align with
the majority of the users are the most afected ones by spam ratings for Amazon datasets.
Overall, we observed that 25%-29% of benign users in Amazon datasets are users who would
not be equally satisfied by recommenders afected by shilling attacks. Thus, recommender
algorithms are not uniformly robust for all types of benign users in the presence of spammers
ratings/reviews.</p>
      <p>The rest of this paper is organized as follows. In Section 2, we summarize related work; we
then outline the dataset, algorithms, and evaluation strategies used in our empirical analysis in
Section 3. In Section 4 we report on our results and, finally, conclusions are drawn in Section 5.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        Collaborative filtering-based recommender systems are widely used to provide recommendations
to users in opinion-based systems, yet they are vulnerable to shilling attacks [
        <xref ref-type="bibr" rid="ref2 ref6">6, 2</xref>
        ]. These attacks
consist of fake user profiles injected into the system with the goal of providing spam ratings
or reviews to promote or demote specific products. While some Shilling attacks promote the
recommendations of certain attacked items (referred to as the push attack), others might demote
the predictions that are made for attacked items (referred to as the nuke attack) [
        <xref ref-type="bibr" rid="ref4 ref7">4, 7</xref>
        ]. Previous
work has defined several attack strategies, including random, average, bandwagon, love/hate,
segmented, and probe attacks. These strategies difer in the way fake profiles choose filler items,
i.e., other rated items chosen beyond attacked items to camouflage the fraudulent behavior.
More sophisticated attacks have been recently proposed [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], for instance, the one by Fang et
al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] looks at how to choose filler items to recommend an attacker-chosen targeted item to as
many users as possible. In the field of machine learning, many eforts have been devoted over the
years to develop techniques for automatic detection of such fraudulent profiles; the techniques
presented in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] are among the most recent ones. In the field of recommender systems,
researchers have focused on studying the efects of such shilling attacks mainly on collaborative
ifltering-based recommenders since the early 2000s [
        <xref ref-type="bibr" rid="ref11 ref4">11, 4</xref>
        ] and developed strategies to make
such algorithms more robust to shilling attacks [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We highlight, for example, outcomes of
the research conducted by Seminario and Wilson [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ] who explicitly look at power user
and power items attacks, i.e., attacks targeting influential users and items, respectively, within
collaborative filtering-based recommender systems.
      </p>
      <p>
        Most recently, the concept of diferential privacy has been explored to make matrix
factorizationbased collaborative filtering recommender algorithms more robust [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The vulnerability of
deep-learning-based recommender systems to shilling attacks has been studied in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In
particular, Lin et al. introduce a framework that considers complex attacks aimed towards
specific user groups. On a diferent perspective, Deldjoo et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] explore dataset characteristics
to explain an observed change in the performance of recommendation under shilling attacks.
      </p>
      <p>Our work add to this body of knowledge by exploring the robustness of collaborative
recommender systems to shilling attacks by using real-world data with spam reviews ground truth, as
opposed to attack simulation and investigating if some users are more vulnerable than others.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Settings</title>
      <p>As previously stated, our goal is to analyze commonly-used memory-based and model-based
collaborative filtering-based recommendation algorithms’ robustness to shilling attacks using a
number of datasets with ground truth on spam reviews. In the rest of this section, we describe
the experimental protocol for our analysis.</p>
      <sec id="sec-3-1">
        <title>3.1. Datasets</title>
        <p>For analysis purposes, we rely on four datasets (described below) produced based upon data
from two well-known e-commerce platforms: Yelp! and Amazon.</p>
        <p>
          Yelp! We consider Yelp! reviews from two domains: hotels (YH) and restaurants (YR) [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
Yelp filters fake/suspicious reviews and puts them on a spam list. A study found the Yelp filter
to be highly accurate [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], and many researchers have used filtered spam reviews as ground
truth for spammer detection [
          <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
          ]. Spammers, in our case, are users who wrote at least one
ifltered review. We removed users who rated the same products multiple times and reviews
with a rating of zero.
        </p>
        <sec id="sec-3-1-1">
          <title>Dataset</title>
          <p>YH
YR
AB
AH</p>
          <p>
            Amazon We also consider Amazon reviews from two domains: beauty (AB) and health
(AH) [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ]. In this case, we define ground truth based on helpfulness votes following the
approach suggested by [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] and based on the findings provided by Fayazi et al. [
            <xref ref-type="bibr" rid="ref22">22</xref>
            ]. Thus, we
treat as a spammer every user who wrote at least one spam review. We define a review as spam
if the rating is 4 or 5 and the helpfulness ratio is ≤ 0.4.
          </p>
          <p>We provide descriptive statistics for the four datasets in Table 1. It is important to note that
rating distribution is not similar across the datasets. As illustrated in Figure 1, rating trends
from AH are dissimilar to the other counterparts, with a vast number of users rating only 1
item. Moreover, the rating distribution of benign users vs. spammers on attacked products (i.e.,
products receiving at least one spam review) is captured in Figure 2. From this figure, it emerges
that benign users and spammer counterparts exhibit similar rating behavior in YR, AB, and AH,
whereas in the case of YH, spammers noticeably assign ratings of ‘1’ more often than benign
counterparts; the opposite is true for ratings of ‘4’.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Algorithms</title>
        <p>
          We focus our analysis on well-known and widely-used collaborative filtering-based
recommendation algorithms implemented using Lenskit for python [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], except for Probabilistic Matrix
Factorization, for which we relied on the implementation provided by Mnih et al. [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].
        </p>
        <p>
          (b) Yelp!-Restaurant
(c) Amazon-Beauty
(d) Amazon-Health
Item-item [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] is the popular item-based collaborative filter algorithm. It utilizes an
itemitem matrix to determine the similarity between the target item and other items (neighbors).
We used this algorithm with 20 neighbors and cosine similarity as similarity measure.
Probabilistic Matrix Factorization (PMF) [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] is a commonly-used latent factor-based
recommendation algorithm. Specifically, probabilistic matrix factorization decomposes the
sparse user-item matrix into low-dimensional matrices with latent factors to generate
recommendations. We used this algorithm with 40 latent factors and 150 iterations. This algorithm is
known for its accuracy, scalability, and dealing with sparsity.
        </p>
        <p>
          Alternating Least Squares (ALS) [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] is a matrix factorization-based algorithm designed to
improve recommendation algorithms performance in large-scale collaborative filtering problems.
This algorithm gain recognition following its success on the Netflix Challenge [
          <xref ref-type="bibr" rid="ref27 ref28">27, 28</xref>
          ]. In our
case, we consider 40 latent factors, 5 damping factors, and 150 iterations for our experiment.
Bayesian Personalized Ranking (BPR) [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] is a rank-based matrix factorization algorithm,
with 40 latent factors, 5 damping factors, and 150 iterations for our experiment. Note that as
top-N recommendation algorithm, i.e., based on rating information is in the form of implicit
feedback [
          <xref ref-type="bibr" rid="ref29 ref30">29, 30</xref>
          ], BPR scores items, but does not produce rating predictions. Thus, we are
forced to exclude BPR from the RMSE-based analysis discussed in Section 4.
FunkSVD [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ] is the well-known gradient descent matrix factorization technique with 40
latent features and 150 training iterations per feature.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Evaluation Framework</title>
        <p>
          By following the classical evaluation framework for shilling attacks on recommender systems [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ],
we measured the performances on the original dataset (with spammers) and when we remove
all the spammers (shilling attack), using well-known performance metrics. In all cases, we
performed 5-fold cross-validation. We tested whether diferences in the metric values with and
without spammers were statistically significant using a paired t-test.
        </p>
        <p>
          Metrics. For assessment, we turn to Root Mean Square Error (RMSE) and Normalized
Discounted Cumulative Gain (NDCG), which are classical measures for evaluating predictive and
top-N recommendations. We also consider measures explicitly defined to quantify the impact of
spammer attacks on recommenders: Prediction Shift (PS), which captures the average absolute
changes in predicted ratings for attacked items and Hit Ratio (HR), which considers if attacked
items are promoted to user top-n recommendations (cf. Burke et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] for formal definitions of
these metrics).
        </p>
        <p>
          We first examined recommender performance by considering all users in the respective
datasets. We then segmented users into fairness categories as computed by the Fairness and
Goodness algorithm described in the next paragraph in order to allow for more in-depth
explorations. By segmenting users based on fairness scores, it is possible to identify mainstream
and non-mainstream users. The latter are those whose rating patterns do not align with the
majority, i.e., liking what most people dislike and vice versa [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ].
        </p>
        <p>
          Fairness and Goodness The Fairness and Goodness algorithm (F&amp;G) [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] provides a measure
for capturing user rating behavior. While many measures exist for this task [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], we chose to use
F&amp;G as Serra et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] recently show it to be the best measure to identify trustworthy users in
opinion-based systems. F&amp;G computes a fairness score for each user and a goodness score for
each item. Specifically, the fairness  () of a user  is a measure of how fair or trustworthy the
user is in rating items. Intuitively, a ‘fair’ or ‘trustworthy’ rater should give an item the rating
that it deserves, while an ‘unfair’ one would deviate from that value. In the case of benign users,
the latter could be the case of an uninformative or non-mainstream user. The goodness ()
of an item  specifies how much users in the system like the item and what its true quality is.
Fairness and goodness are mutually recursively computed as:
 () = 1 −
|()|  ()
1
∑︁ | (, ) − ()|
        </p>
        <p>() =
1</p>
        <p>
          ∑︁  () ×  (, )
|()|  ()
(1)
(2)
where  (, ) is the rating given by the user  to the item , () is the set of ratings given
by user , () is the set of ratings received by item , and  = 4 in this case which corresponds
to the maximum error in a five-star rating system. Thus, the goodness of an item is given by
the average of its rating, where each rating is weighted by the fairness of the rater, while the
fairness of a user considers how much the ratings a user gives are far from the goodness of the
items. The higher the fairness, the more trustworthy the user is. Fairness scores of the user lie
in the [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] interval, and goodness scores lie in the [
          <xref ref-type="bibr" rid="ref1 ref5">1, 5</xref>
          ] interval.
cant diferences are shaded in gray,  ≤ 0.001.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Discussion</title>
      <p>In this section, we present our experimental evaluation of five recommendation algorithms on
four datasets of diferent domains. We discuss the efect of shilling attacks on recommendations
ofered to users in real-world scenarios, as opposed to simulated attacks. Specifically, we aim to
answer the following research questions:
RQ1 Do spammer’s ratings impact recommendations?
RQ2 Who is really afected by spammers?
By investigating recommender algorithm performance in the presence of spammers as well as
when spammers are removed, the first question enables us to gauge the shilling attacks’ efect
on the robustness of the considered recommendation algorithms. For the second question, we
used the fairness metric to determine mainstream and non-mainstream users and quantify the
efect of spammers on recommendations received by non-mainstream users.</p>
      <sec id="sec-4-1">
        <title>4.1. Do spammers ratings impact recommendations?</title>
        <p>
          To answer RQ1, we consider the performance of the recommender algorithms yielded on four
diferent datasets, as reported in Table 2. It comes across from the reported scores that removing
spammers indeed leads to lower RMSE scores, i.e., better predictions. Previous works have
shown PS values ranging from 0.5 to 1.5 when shilling attacks are simulated [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ]. However,
we observe very low values in real-world scenarios: in our case, considered, PS ranges from
0.087 to 0.160, which we argue might not be enough to promote or demote products attacked
by the spammers. When looking at algorithm performance from a top-N recommendation
standpoint, from reported NDCG@5 we see that, often, NDCG@5 scores tend to increase when
removing spammers. This means that users’ preferred items are more likely to appear within
the top-5 recommendations when spammers are excluded. Unfortunately, improvement is not
always meaningful, i.e., from Table 2 we see that improvements are not always significant,
especially on YH. We anticipated lower HR@5 scores when excluding spammers–we assumed
fewer attacked items would be promoted among the top-5 recommendations. Instead, we see
similar trends among HR@5 results as those observed for NDCG@5. In other words, for YH
and YR, performance is comparable regardless of the presence of spammers (i.e., diferences in
performance are not significant); for AB and AH we see significant diferences in performance.
        </p>
        <p>Overall, we can say that, in theory, the performance of collaborative filtering-based
recommender algorithms is afected by spammers’ ratings/reviews. This is particularly noticeable for
predictive recommenders (i.e., all algorithms yielded significant diferences across the datasets).
In practice, however, performance improvements are in their majority barely perceptible. This
leads us to question whether algorithm robustness is reflected by average metrics like RMSE
or NDCG. In the end, looking at recommender performance as a whole may not clearly
quantify how much spammers are able to deceit recommenders and, more importantly, if there
are specific user groups that are afected more than others. With this in mind, we conduct a
more thorough analysis with the aim of understanding if the aforementioned diferences in
performance are more pronounced among certain types of users (i.e., non-mainstream ones).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Who is really afected by spammers?</title>
        <p>To better understand which users are really afected by spammers, we analyzed users based on
their fairness: the ability of a user to rate a product according to what it deserves. It is worth
noting, however, that the rating a product deserves often aligns with what the majority of
benign users (mainstream users) think about that product, as mainstream users often outnumber
non-mainstream and spam users. We investigate trends according to RMSE, NDCG@5, and hit
ratio. As noted in the prior subsection, prediction shift values were small, so we excluded this
metric from our analysis).</p>
        <p>Figure 3 illustrates how the RMSE varies according to the fairness of benign users; for ease of
readability, we highlight statistically significant diferences in performance when spammers
are excluded in Table 3. We start by observing that, regardless of the algorithm for both Yelp!
datasets and with just one exception (YR, ALS, and FunkSVD, (0.4 − 0.5]), removing spammers
reduces the RMSE for all users. For the Amazon datasets, instead, when the user fairness
is greater than 0.4, removing spammers increases the RMSE for all users for each algorithm.
We posit these results could be due to the rating distributions of spammers vs. benign users
across these two platforms. As previously shown in Figure 2, spammer and benign users are
more similar in Amazon than Yelp!, with the majority of ratings being 4 and 5. Therefore,
removing spam could cause the recommender to lose information from mainstream users. On
the other end, when fairness is less than or equal to 0.4 among Amazon users, in most cases
where the diference is statistically significant, i.e., 14 out of 19 cases, removing spammers
YH
YR
AB
AH
with spammers (W) across diferent user fairness ranges (* means  &lt; 0.03 and ** means
 ≤ 0.01). Cases where removing spammers reduces the RMSE are shaded.
enables algorithms to avoid noise signals and thus perform better for these users (lower RMSE).
Note that there are more cases in AH than AB (4 out of 11 vs. 1 out of 8) where removing the
spammers is not beneficial for non-mainstream users. This could be due to the fact that, as
shown in Figure 1, AH data is more sparse than other datasets, making the process of generating
recommendations more dificult for most of the users in such a setting, independently of the
presence of spam.</p>
        <p>When we look at trends for NDCG@5, Figure 4 and Table 4 show that the quality of the
generated recommendations seldom improves on the Yelp! datasets for users having fairness
YH
YR
AB</p>
        <sec id="sec-4-2-1">
          <title>Item-Item PMF BPR ALS</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>FunkSVD</title>
          <p>with spammers (W) across diferent user fairness ranges (* means  &lt; 0.03 and ** means
 ≤ 0.01). Cases where removing spammers increases the NDCG@5 are shaded.
greater than 0.5, whereas for the Amazon datasets, the value of NDCG@5 is higher when
spammers are removed in the majority of the cases (28 out of 47) and independently of the user
type.</p>
          <p>Overall, our analysis reveals that removing spammers helps in reducing the number of
attacked items that hit the top-5 recommendations for all the users in all the datasets (see HR@5
analysis in Table 5).1 Moreover, removing spammers in Yelp! is beneficial for all the users when
considering the predictive performance of algorithms (based on RMSE); for Amazon, top-N
algorithms are better (according to NDCG@5) among mainstream users, who see more tailored
1For brevity, we exclude a figure akin to those complementing Tables 3 and 4.
recommended items in their top-5 item list when spammers are removed from the system. Also,
we see performance improvement in terms of both RMSE and NDCG@5 scores for Amazon
non-mainstream users, i.e., the ones with low fairness, hence the ones whose ratings are very
far from the ones of the majority of the users. Non-mainstream users afected by spammers
represent 29% (resp. 25%) of benign users in AB (resp. AH). In a real-world scenario, these
percentages would translate into hundreds of thousands of users who would not be equally
satisfied by recommenders that are not robust to shilling attacks.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>In this paper, we have taken a deeper look into how shilling attacks afect recommender systems
in a real-world scenario. For this, we conducted an in-depth exploration of the performance of
ifve well-known collaborative filtering algorithms on four diferent datasets.</p>
      <p>We saw similar trends among the performance of predictive and top-N recommenders: users
are exposed to better recommendations when spammers are excluded (RQ1). This highlights the
importance of further research on spammer detection and robust recommender systems. At the
same time, we question if the small diferences in performance (albeit statistically significant)
would be evident to recommender systems’ users and whether metrics considered for assessment
which aggregate performance for all users could obfuscate users who are more deeply afected
by spammers. This leads us to explore diferences in performance between mainstream vs.
non-mainstream users (RQ2). We saw that Amazon non-mainstream users are the ones most
afected by spam ratings according to both RMSE and NDCG@5.</p>
      <p>
        Based on the findings emerging from the analysis presented in this paper, it follows that future
work will be devoted to looking at other types of recommender algorithms, beyond collaborative
ifltering-based, to see if the trends we have observed in our analysis remain. Moreover, we plan
also to test the efectiveness of adversarial training for recommender systems [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] under the
real-world attacks considered in this paper.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This work has been supported by the National Science Foundation under Awards no. 1943370
and 1820685.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.-A.</given-names>
            <surname>Chirita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Nejdl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zamfir</surname>
          </string-name>
          ,
          <article-title>Preventing shilling attacks in online recommender systems</article-title>
          ,
          <source>in: Proceedings of the 7th annual ACM international workshop on Web information and data management</source>
          ,
          <year>2005</year>
          , pp.
          <fpage>67</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Si</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Shilling attacks against collaborative recommender systems: a review</article-title>
          ,
          <source>Artificial Intelligence Review</source>
          <volume>53</volume>
          (
          <year>2020</year>
          )
          <fpage>291</fpage>
          -
          <lpage>319</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Deldjoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. D.</given-names>
            <surname>Noia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Merra</surname>
          </string-name>
          ,
          <article-title>A survey on adversarial recommender systems: from attack/defense strategies to generative adversarial networks</article-title>
          ,
          <source>ACM Computing Surveys (CSUR) 54</source>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Burke</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. O'Mahony</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Hurley</surname>
          </string-name>
          ,
          <article-title>Robust collaborative recommendation</article-title>
          ,
          <source>in: Recommender systems handbook</source>
          , Springer,
          <year>2015</year>
          , pp.
          <fpage>961</fpage>
          -
          <lpage>995</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Seminario</surname>
          </string-name>
          ,
          <article-title>Accuracy and robustness impacts of power user attacks on collaborative recommender systems</article-title>
          , in: RecSys, ACM,
          <year>2013</year>
          , pp.
          <fpage>447</fpage>
          -
          <lpage>450</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Sundar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. D.</given-names>
            <surname>Russomanno</surname>
          </string-name>
          ,
          <article-title>Understanding shilling attacks and their detection traits: a comprehensive survey</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>171703</fpage>
          -
          <lpage>171715</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>M. P. O'Mahony</surname>
            ,
            <given-names>N. J.</given-names>
          </string-name>
          <string-name>
            <surname>Hurley</surname>
            ,
            <given-names>G. C.</given-names>
          </string-name>
          <string-name>
            <surname>Silvestre</surname>
          </string-name>
          ,
          <article-title>Recommender systems: Attack types and strategies</article-title>
          , in: AAAI,
          <year>2005</year>
          , pp.
          <fpage>334</fpage>
          -
          <lpage>339</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. Z.</given-names>
            <surname>Gong</surname>
          </string-name>
          , J. Liu,
          <article-title>Influence function based data poisoning attacks to top-n recommender systems</article-title>
          ,
          <source>in: Proceedings of The Web Conference</source>
          <year>2020</year>
          ,
          <year>2020</year>
          , pp.
          <fpage>3019</fpage>
          -
          <lpage>3025</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hooi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Makhija</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Faloutsos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. S.</given-names>
            <surname>Subrahmanian</surname>
          </string-name>
          ,
          <article-title>REV2: fraudulent user prediction in rating platforms</article-title>
          , in: WSDM, ACM,
          <year>2018</year>
          , pp.
          <fpage>333</fpage>
          -
          <lpage>341</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Serra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shrestha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Spezzano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Squicciarini</surname>
          </string-name>
          ,
          <string-name>
            <surname>Deeptrust:</surname>
          </string-name>
          <article-title>An automatic framework to detect trustworthy users in opinion-based systems</article-title>
          , in: CODASPY, ACM,
          <year>2020</year>
          , pp.
          <fpage>29</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Lam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          ,
          <article-title>Shilling recommender systems for fun and profit</article-title>
          , in: WWW, ACM,
          <year>2004</year>
          , pp.
          <fpage>393</fpage>
          -
          <lpage>402</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Seminario</surname>
          </string-name>
          , D. C. Wilson,
          <article-title>Attacking item-based recommender systems with power items</article-title>
          ,
          <source>in: Proceedings of the 8th ACM Conference on Recommender systems</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Seminario</surname>
          </string-name>
          , D. C. Wilson,
          <article-title>Nuke'em till they go: Investigating power user attacks to disparage items in collaborative recommenders</article-title>
          ,
          <source>in: Proceedings of the 9th ACM Conference on Recommender Systems</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>293</fpage>
          -
          <lpage>296</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wadhwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chaudhari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Achan</surname>
          </string-name>
          ,
          <article-title>Data poisoning attacks against diferentially private recommender systems</article-title>
          ,
          <source>in: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1617</fpage>
          -
          <lpage>1620</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Attacking recommender systems with augmented user profiles</article-title>
          ,
          <source>in: Proceedings of the 29th ACM International Conference on Information &amp; Knowledge Management</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>855</fpage>
          -
          <lpage>864</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Deldjoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Sciascio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Merra</surname>
          </string-name>
          ,
          <article-title>How dataset characteristics afect the robustness of collaborative recommendation models</article-title>
          ,
          <source>in: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>951</fpage>
          -
          <lpage>960</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Venkataraman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Glance</surname>
          </string-name>
          ,
          <article-title>What yelp fake review filter might be doing?</article-title>
          ,
          <source>in: ICWSM</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>409</fpage>
          -
          <lpage>418</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>K.</given-names>
            <surname>Weise</surname>
          </string-name>
          ,
          <article-title>A lie detector test for online reviewers</article-title>
          ,
          <source>in: Bloomberg BusinessWeek</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rayana</surname>
          </string-name>
          , L. Akoglu,
          <article-title>Collective opinion spam detection: Bridging review networks and metadata</article-title>
          , in: SIGKDD,
          <year>2015</year>
          , pp.
          <fpage>985</fpage>
          -
          <lpage>994</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>S. K. C.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <article-title>On the temporal dynamics of opinion spamming: Case studies on yelp</article-title>
          ,
          <source>in: WWW</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>369</fpage>
          -
          <lpage>379</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>J. McAuley</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <article-title>From amateurs to connoisseurs: modeling the evolution of user expertise through online reviews</article-title>
          ,
          <source>in: WWW</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>897</fpage>
          -
          <lpage>908</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fayazi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Caverlee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Squicciarini</surname>
          </string-name>
          ,
          <article-title>Uncovering crowdsourced manipulation of online reviews</article-title>
          ,
          <source>in: SIGIR</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>233</fpage>
          -
          <lpage>242</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>M. D. Ekstrand</surname>
          </string-name>
          ,
          <article-title>Lenskit for python: Next-generation software for recommender systems experiments</article-title>
          , in: CIKM,
          <year>2020</year>
          , pp.
          <fpage>2999</fpage>
          -
          <lpage>3006</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mnih</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <article-title>Probabilistic matrix factorization</article-title>
          ,
          <source>in: NIPS</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>1257</fpage>
          -
          <lpage>1264</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sarwar</surname>
          </string-name>
          , G. Karypis,
          <string-name>
            <given-names>J.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          ,
          <article-title>Item-based collaborative filtering recommendation algorithms</article-title>
          , in: WWW,
          <year>2001</year>
          , pp.
          <fpage>285</fpage>
          -
          <lpage>295</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Koren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Volinsky</surname>
          </string-name>
          ,
          <article-title>Collaborative filtering for implicit feedback datasets</article-title>
          , in: IEEE ICDM,
          <year>2008</year>
          , pp.
          <fpage>263</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Koren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Volinsky</surname>
          </string-name>
          ,
          <article-title>Matrix factorization techniques for recommender systems</article-title>
          ,
          <source>Computer</source>
          <volume>42</volume>
          (
          <year>2009</year>
          )
          <fpage>30</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wilkinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schreiber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <article-title>Large-scale parallel collaborative filtering for the netflix prize</article-title>
          ,
          <source>in: AAIM</source>
          , Springer,
          <year>2008</year>
          , pp.
          <fpage>337</fpage>
          -
          <lpage>348</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rendle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Freudenthaler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gantner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt-Thieme</surname>
          </string-name>
          ,
          <article-title>Bpr: Bayesian personalized ranking from implicit feedback</article-title>
          ,
          <source>arXiv preprint arXiv:1205.2618</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Wang</surname>
          </string-name>
          , D.-Y. Yeung,
          <article-title>Collaborative deep learning for recommender systems</article-title>
          ,
          <source>in: Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>1235</fpage>
          -
          <lpage>1244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>A.</given-names>
            <surname>Paterek</surname>
          </string-name>
          ,
          <article-title>Improving regularized singular value decomposition for collaborative filtering</article-title>
          ,
          <source>in: Proceedings of KDD cup and workshop</source>
          , volume
          <year>2007</year>
          ,
          <year>2007</year>
          , pp.
          <fpage>5</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>R. Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Urbano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanjalic</surname>
          </string-name>
          ,
          <article-title>Leave no user behind: Towards improving the utility of recommender systems for non-mainstream users</article-title>
          , in: WSDM, ACM,
          <year>2021</year>
          , pp.
          <fpage>103</fpage>
          -
          <lpage>111</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Spezzano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. S.</given-names>
            <surname>Subrahmanian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Faloutsos</surname>
          </string-name>
          ,
          <article-title>Edge weight prediction in weighted signed networks</article-title>
          ,
          <source>in: IEEE ICDM</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>221</fpage>
          -
          <lpage>230</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Lam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          ,
          <article-title>Shilling recommender systems for fun and profit</article-title>
          , in: WWW,
          <year>2004</year>
          , pp.
          <fpage>393</fpage>
          -
          <lpage>402</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>