<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>IIR</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Replication of Collaborative Filtering Generative Adversarial Networks on Recommender Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Discussion Paper</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fernando B. Pérez Maurera</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maurizio Ferrari Dacrema</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Cremonesi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ContentWise</institution>
          ,
          <addr-line>Via Simone Schiafino 11, Milano, 20158, Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Politecnico di Milano</institution>
          ,
          <addr-line>Piazza Leonardo da Vinci 32, 20133 Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>12</volume>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>CFGAN and its family of models (TagRec, MTPR, and CRGAN) learn to generate personalized and fake-but-realistic preferences for top-N recommendations by solely using previous interactions. The work discusses the impact of certain diferences between the CFGAN framework and the model used in the original evaluation. The absence of random noise and the use of real user profiles as condition vectors leaves the generator prone to learn a degenerate solution in which the output vector is identical to the input vector, therefore, behaving essentially as a simple auto-encoder. This work further expands the experimental analysis comparing CFGAN against a selection of simple and well-known properly optimized baselines, observing that CFGAN is not consistently competitive against them despite its high computational cost. This work is an extended abstract of the paper presented in [1].</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Generative Adversarial Networks</kwd>
        <kwd>Recommender Systems</kwd>
        <kwd>Collaborative Filtering</kwd>
        <kwd>Replicability</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Start
z
c</p>
      <p>Update G</p>
      <p>G</p>
      <p>Generated
Profiles</p>
      <p>Mask</p>
      <p>Probability
Masked Profiles are Real
Masked
Profiles</p>
      <p>D</p>
      <p>c</p>
      <p>Probability</p>
      <p>Real Profiles are Real
given user, CFGAN constructs their recommendations by selecting the top-N items with the
highest generated preference score.</p>
      <p>
        This work presents and discusses the results of several experiments on CFGAN with the
goal of addressing two research objectives. First, to describe inconsistencies found between
the formulation of CFGAN and the implementation of it used in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Second, to replicate
the claimed progress made in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] by measuring the CFGAN quality under a traditional top-N
recommendation scenario against properly-tuned baselines. The discussions presented here are
aligned with the exhortation given in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]: research works should focus on understanding and
analyzing the proposed models.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Inconsistencies in CFGAN</title>
      <p>
        Figure 1 presents the architecture and training process of CFGAN as described in the reference
of CFGAN [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. From the figure, two vectors are part of the architecture of CFGAN: the random
and condition vectors, denoted as  and , respectively. This work highlights inconsistencies
between the implementation and the reference CFGAN presented in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]: the use of user profiles
as the condition vector and the absence of random noise for the empirical evaluation. These
inconsistencies raise concerns about the model’s ability to generalize and provide personalized
recommendations.
      </p>
      <p>
        First, the condition vector is used to provide personalized recommendations. To achieve
this, this vector is encoded with users features, e.g., location, social information, identifiers,
among others. Due to the collaborative nature of the datasets used in the experiments of the
reference CFGAN [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the user profiles are used as condition vectors, i.e., the data points that
the generator and discriminator learn from. Using the user profiles as the condition vector
makes both networks prone to learn trivial solutions. Essentially, the generator fundamentally
becomes an auto-encoder and the discriminator may degenerate into learning a function that
compares the condition with the real or generated profile.
      </p>
      <p>
        Second, from a theoretical standpoint, the random noise is required on traditional GANs to
explore several points and to create a mapping between the random to the data spaces. The
random noise vector is also part of the CFGAN reference and it serves the same purpose as for
traditional GANs. However, the implementation of CFGAN in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] removes the random noise
from the model. Due to the absence of random noise, CFGAN is trained on highly sparse user
profiles without the exploration of diferent input spaces. Furthermore, removing the random
noise implicitly makes the assumption user profiles are static over time. As a consequence,
CFGAN is less robust to evolving users preferences and dataset shifts [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Methodology</title>
      <p>
        This work presents an evaluation study comprised of several experiments on CFGAN. The goal
of this evaluation study is two-fold. First, to replicate the progress claims made in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], where
“replicability” is defined as in the ACM Artifact Review and Badging, version 1.1. 2 Second, to
measure the efects in recommendation quality caused by the inconsistencies between CFGAN
description and its implementation. The supplemental material provided in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] solely contain
the implementation of CFGAN and its data splitting, training, and evaluation. The details of the
experimental methodology of the evaluation study is as follows:
Datasets and Splits: The experiments used the same open-source datasets and random holdout
splits in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], i.e., a sampled version of Ciao [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and ML100K and ML1M versions of
Movielens [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. A validation split was created for hyper-parameter tuning purposes following the
same split-creation steps as in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Evaluation: All recommenders were evaluated on traditional accuracy and beyond-accuracy
metrics [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] in the standard top-N recommendation scenario. Hyper-parameters were searched
using bayesian search with 16 random cases, 50 total cases, and optimizing NDCG [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Baseline Recommenders: Neighborhood-based (Item KNN and User KNN) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], graph-based
( 3) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], auto-encoders (SLIM ElasticNet [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and EASE R [14]), and machine learning
recommenders (PureSVD [15] and MF BPR [16]). The description of these recommenders, their
hyper-parameters, and their ranges is found in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        CFGAN Recommenders: CFGAN as implemented in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] was optimized. Two diferent variants
were trained using the optimal hyper-parameters of the previous: CFGAN with random noise,
and CFGAN using user identifiers as condition vectors. 3
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Discussion</title>
      <p>2Available online at https://www.acm.org/publications/policies/artifact-review-and-badging-current.
3Due to space limitations, this work omits the list of hyper-parameters of CFGAN.
4Due to space limitations, only a subset of accuracy and beyond-accuracy metrics are shown.</p>
      <sec id="sec-4-1">
        <title>PRECISION</title>
      </sec>
      <sec id="sec-4-2">
        <title>RECALL</title>
      </sec>
      <sec id="sec-4-3">
        <title>UserKNN</title>
      </sec>
      <sec id="sec-4-4">
        <title>ItemKNN</title>
      </sec>
      <sec id="sec-4-5">
        <title>RP3beta</title>
      </sec>
      <sec id="sec-4-6">
        <title>PureSVD</title>
      </sec>
      <sec id="sec-4-7">
        <title>SLIM ElasticNet MF BPR</title>
      </sec>
      <sec id="sec-4-8">
        <title>EASE R</title>
      </sec>
      <sec id="sec-4-9">
        <title>CFGAN</title>
      </sec>
      <sec id="sec-4-10">
        <title>CFGAN UI</title>
        <p>
          CFGAN RN
than CFGAN. In particular, User KNN, SLIM ElasticNet, and EASE R have relative higher
NDCG than CFGAN by 2.34 %, 8.53 %, and 10.34 %, respectively. Furthermore, these more
accurate baselines also trained faster than CFGAN, with diferences in training time between
two or three orders of magnitude. The results indicate that the progress claims made in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]
could not be replicated in the experiments of this evaluation study.
        </p>
        <p>Regarding the absence of random noise, the results of the experiments are varied. Across
datasets and variants, including random noise to CFGAN (CFGAN RN in Table 1) led to both
relative increases or decreases in accuracy without a clear pattern.</p>
        <p>Clear patterns resulted by changing the condition vector from user profiles to user identifiers
(CFGAN UI in Table 1). Particularly, across datasets and variants, CFGAN UI consistently
obtained relative lower accuracy metrics with respect to the base CFGAN. These results impose
the following dichotomy. On one hand, using user profiles as condition vector may lead to
both networks learn a trivial solution, as discussed in Section 2. On the other hand, using
user identifiers as condition vectors when learning from pure collaborative data is possible on
CFGAN at the cost of providing accurate recommendations.</p>
        <p>Further studies are needed to address the recommendation quality of CFGAN and the
inconsistencies presented in this work. For instance, a revision of the architecture of CFGAN can
be addressed in future works. In this work, the results suggest that the current architecture
does not work when changing the condition vectors to be the user identifiers. All these aspects
are still open research questions and addressing will be beneficial for the maturity of this
recommendation model.
2011.134.
[14] H. Steck, Embarrassingly shallow autoencoders for sparse data, in: The World Wide
Web Conference, WWW 2019, San Francisco, CA, USA, May 13-17, 2019, ACM, 2019, pp.
3251–3257. doi:10.1145/3308558.3313710.
[15] P. Cremonesi, Y. Koren, R. Turrin, Performance of recommender algorithms on top-n
recommendation tasks, in: Proceedings of the 2010 ACM Conference on Recommender
Systems, RecSys 2010, Barcelona, Spain, September 26-30, 2010, ACM, 2010, pp. 39–46.
doi:10.1145/1864708.1864721.
[16] S. Rendle, C. Freudenthaler, Z. Gantner, L. Schmidt-Thieme, BPR: bayesian personalized
ranking from implicit feedback, in: UAI 2009, Proceedings of the Twenty-Fifth Conference
on Uncertainty in Artificial Intelligence, Montreal, QC, Canada, June 18-21, 2009, AUAI
Press, 2009, pp. 452–461.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F. B.</given-names>
            <surname>Pérez Maurera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Ferrari</given-names>
            <surname>Dacrema</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          ,
          <article-title>An evaluation study of generative adversarial networks for collaborative filtering</article-title>
          ,
          <source>in: Advances in Information Retrieval - 44th European Conference on IR Research</source>
          , ECIR
          <year>2022</year>
          , Stavanger, Norway,
          <source>April 10-14</source>
          ,
          <year>2022</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          , volume
          <volume>13185</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2022</year>
          , pp.
          <fpage>671</fpage>
          -
          <lpage>685</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -99736-6\_
          <fpage>45</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ferrari Dacrema</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Boglio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>A troubling analysis of reproducibility and progress in recommender systems research</article-title>
          ,
          <source>ACM Trans. Inf. Syst</source>
          .
          <volume>39</volume>
          (
          <year>2021</year>
          )
          <volume>20</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          :
          <fpage>49</fpage>
          . doi:
          <volume>10</volume>
          .1145/3434185.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ferrari Dacrema</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>Are we really making much progress? A worrying analysis of recent neural recommendation approaches</article-title>
          ,
          <source>in: Proceedings of the 13th ACM Conference on Recommender Systems, RecSys</source>
          <year>2019</year>
          , Copenhagen, Denmark,
          <source>September 16-20</source>
          ,
          <year>2019</year>
          , ACM,
          <year>2019</year>
          , pp.
          <fpage>101</fpage>
          -
          <lpage>109</lpage>
          . doi:
          <volume>10</volume>
          .1145/3298689.3347058.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>The neural hype and comparisons against weak baselines</article-title>
          ,
          <source>SIGIR Forum 52</source>
          (
          <year>2019</year>
          )
          <fpage>40</fpage>
          -
          <lpage>51</lpage>
          . doi:
          <volume>10</volume>
          .1145/3308774.3308781.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>The neural hype, justified! a recantation</article-title>
          ,
          <source>SIGIR Forum 53</source>
          (
          <year>2021</year>
          )
          <fpage>88</fpage>
          -
          <lpage>93</lpage>
          . doi:
          <volume>10</volume>
          . 1145/3458553.3458563.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Critically examining the "neural hype": Weak baselines and the additivity of efectiveness gains from neural ranking models</article-title>
          ,
          <source>in: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <string-name>
            <surname>SIGIR</surname>
          </string-name>
          <year>2019</year>
          , Paris, France,
          <source>July 21-25</source>
          ,
          <year>2019</year>
          , ACM,
          <year>2019</year>
          , pp.
          <fpage>1129</fpage>
          -
          <lpage>1132</lpage>
          . doi:
          <volume>10</volume>
          . 1145/3331184.3331340.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Chae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>CFGAN:</surname>
          </string-name>
          <article-title>A generic collaborative filtering framework based on generative adversarial networks</article-title>
          ,
          <source>in: Proceedings of the 27th ACM International Conference on Information and Knowledge Management</source>
          ,
          <string-name>
            <surname>CIKM</surname>
          </string-name>
          <year>2018</year>
          , Torino, Italy,
          <source>October 22-26</source>
          ,
          <year>2018</year>
          , ACM,
          <year>2018</year>
          , pp.
          <fpage>137</fpage>
          -
          <lpage>146</lpage>
          . doi:
          <volume>10</volume>
          .1145/3269206.3271743.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Z. C.</given-names>
            <surname>Lipton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Steinhardt</surname>
          </string-name>
          ,
          <article-title>Troubling trends in machine learning scholarship</article-title>
          ,
          <source>ACM Queue 17</source>
          (
          <year>2019</year>
          )
          <article-title>80</article-title>
          . doi:
          <volume>10</volume>
          .1145/3317287.3328534.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Quionero-Candela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sugiyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Schwaighofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. D.</given-names>
            <surname>Lawrence</surname>
          </string-name>
          , Dataset Shift in Machine Learning, The MIT Press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gao</surname>
          </string-name>
          , H. Liu,
          <article-title>mTrust: discerning multi-faceted trust in a connected world</article-title>
          ,
          <source>in: Proceedings of the Fifth International Conference on Web Search and Web Data Mining, WSDM</source>
          <year>2012</year>
          , Seattle, WA, USA, February 8-
          <issue>12</issue>
          ,
          <year>2012</year>
          , ACM,
          <year>2012</year>
          , pp.
          <fpage>93</fpage>
          -
          <lpage>102</lpage>
          . doi:
          <volume>10</volume>
          .1145/2124295.2124309.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Harper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <article-title>The MovieLens datasets: History and context</article-title>
          ,
          <source>ACM Trans. Interact. Intell. Syst</source>
          .
          <volume>5</volume>
          (
          <year>2016</year>
          )
          <volume>19</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          :
          <fpage>19</fpage>
          . doi:
          <volume>10</volume>
          .1145/2827872.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Christofel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Paudel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Newell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bernstein</surname>
          </string-name>
          ,
          <article-title>Blockbusters and wallflowers: Accurate, diverse, and scalable recommendations with random walks</article-title>
          ,
          <source>in: Proceedings of the 9th ACM Conference on Recommender Systems, RecSys</source>
          <year>2015</year>
          , Vienna, Austria,
          <source>September 16-20</source>
          ,
          <year>2015</year>
          , ACM,
          <year>2015</year>
          , pp.
          <fpage>163</fpage>
          -
          <lpage>170</lpage>
          . doi:
          <volume>10</volume>
          .1145/2792838.2800180.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>X.</given-names>
            <surname>Ning</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Karypis, SLIM: sparse linear methods for top-n recommender systems</article-title>
          ,
          <source>in: 11th IEEE International Conference on Data Mining, ICDM</source>
          <year>2011</year>
          , Vancouver, BC, Canada,
          <source>December 11-14</source>
          ,
          <year>2011</year>
          , IEEE Computer Society,
          <year>2011</year>
          , pp.
          <fpage>497</fpage>
          -
          <lpage>506</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICDM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>