<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>S. Stoikov)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Better than Bieber? Measuring Song Quality Using Human Feedback</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sasha Stoikov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cornell Financial Engineering Manhattan</institution>
          ,
          <addr-line>New York</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Music recommendation algorithms on streaming platforms tend to reinforce popularity biases over time. This leads to a perceived unfairness for a majority of artists. In this article we address the fairness problem through an assessment of actual user preferences collected on a music app designed to capture the intensity of user sentiment. Users of the app provide explicit feedback on songs, designating them as dislike, like, and superlike (a function which saves songs to a playlist). We find that while the like rate increases monotonically with artist popularity, this does not hold true for superlike rates - those are highest for artists with much lower monthly streams. These findings have implications for the growth potential in the music market and the legitimacy of current recommendation algorithms.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Algorithmic auditing</kwd>
        <kwd>Fairness</kwd>
        <kwd>Music Platforms</kwd>
        <kwd>Recommendation Systems</kwd>
        <kwd>Information Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Why are some songs played billions of times, while others remain unheard? Algorithms on
streaming platforms have the power to make or break songs, in ways that may seem unfair and
opaque, particularly to the artists whose songs have not gone viral. Recommendation systems
are prone to feedback loops where popular songs are recommended disproportionately more
often than undiscovered ones. Simulation studies have shown that systems trained on data
produced by users exposed to recommendations can lead to less diversity and lower utility from
the perspective of users of these platforms [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. From the perspective of artists, field studies
have found that popularity bias and lack of transparency in music recommendation engines
are perceived as a major source of unfairness [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Since most streaming platforms pay a
fraction of a penny per stream, exposure is at the heart of the cultural and economic capital
of an artist. For this reason, exposure fairness, ranking fairness and popularity bias have been
studied extensively in the literature [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Other notions of fairness
related to gender, diversity and genres have also received attention [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] but there
is no consensus among artists as to how they should be addressed [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        The concept of algorithmic fairness has been studied by scholars in other fields like
college admissions or mortgage applications. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] define algorithmic unfairness in terms
of instances where a candidate of higher quality is ranked lower than someone with lower
quality. However, in a domain as subjective as music, the notion of quality needs to be properly
defined. To determine if the popularity of a given artist is fair, exposure data obtained on
streaming platforms is not enough: explicit human feedback is required. What percentage of
people like the top hits? What percentage of people love them? If a lesser known artist is more
broadly liked and more deeply loved, it may be possible to establish that they have been unfairly
under-exposed. In this paper, we address these questions by proposing a measure of quality for
music, using an explicit music ratings dataset [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] collected by Piki, a music ratings app.
      </p>
      <p>
        User generated datasets fall into two broad categories: explicit datasets, where the users
express their opinions on the quality of an item and implicit datasets, where user behavior is
surveilled by a platform and the opinions are implied by the behavior of the user. Well-known
explicit datasets include the Yelp restaurant dataset [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and the Movielens dataset [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Since
feedback is voluntary and since users typically have access to a search bar, this kind of dataset
is often subject to self-selection bias: if there is no obligation to rate, users tend to rate only
when they feel very strongly about an item, and rate more items they like than items they
dislike. Likewise, implicit datasets commonly used to train modern recommendation systems on
streaming apps like Spotify [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] will tend to collect more positive signals, i.e. streams of a song,
than negative signals, say song skips, which are not necessarily expressing an opinion about a
song. The Piki music dataset studied in this paper has been shown to mitigate these positive
biases, which have been shown to increase the accuracy of recommendations algorithms [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>The paper is organized as follows. In section 2, we describe the Piki music interface and
dataset. In section 3, we aggregate this data into two quality metrics, validate their consistency
across populations and compare them for a range of artist popularity and song release dates.
We conclude in section 4.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data collection</title>
      <p>The circumstances motivating Piki users to give ratings are very diferent from those of users
of other ratings or streaming apps.</p>
      <p>1. Piki users are presented with sets of 30-second music video clips which they rate while
they are listening to the music. This is in contrast to restaurant and movie rating apps
where there may be a significant time lag between the experience and the rating. Note
that some of the songs have music videos while others only have audio. The start time is
in the middle of the songs and may be diferent from song to song.
2. Users must provide explicit ratings, disliking, liking, or superliking a song – until they do
so, the clip plays in a loop. Immediately after a rating is provided, the user is presented
the next song. This is in contrast with streaming apps where implicit actions such as
listens, skips, shares and saves to a playlist are a noisy reflection of a user’s tastes.
3. Users are paid small amounts of cash, to rate a large amount of songs. This incentivizes
them to rate, even if some of the songs are not to their liking. Since streaming apps aim to
maximize user retention, they are likely to recommend songs that are predictably likely
to please, not more surprising songs that are outside of their comfort zone.
4. Users don’t have access to a search bar, which could naturally lead to users self-selecting
artists that first come to mind, often celebrities, thus leading to a popularity bias. The
distribution of songs in Piki are approximately 16% random (from a list of songs with
high superlike rates), 76% personalized (a collaborative filtering algorithm) and 8%
hyperpersonalized (an algorithm that presents songs by artists the user has already superliked).
5. A system of timers incentivizes users to rate more uniformly, despite the very diverse
behaviors of people rating a very subjective media like music. Note that the dislike, like
and superlike buttons appear sequentially in that order, several seconds after the song
begins playing. Figure 1 illustrates the way the timers work. The timer on the dislike
button ensures that each song is given a fair listen by the rater. The timers on the like
and superlike buttons ensure that users are willing to invest time in their most preferred
songs. The timers may also be slowed down to throttle users who dislike indiscriminately
rate to get rewards.
6. Superliked songs are saved to a playlist that the user may export to other applications.</p>
      <p>This indicates that they are invested in listening to the song again soon.</p>
      <p>
        The Piki dataset [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] has a similar structure to datasets produced by explicit ratings apps like
Movielens and Yelp, except for each user id and song id, there are 3 possible ratings (dislike, like
and superlike), instead of 5. The dataset consists of 1.5 million ratings on 231,800 songs rated
by 7,519 users. The catalog was build by querying top songs from artists popular on streaming
and live music apps. The songs presented on Piki have a diversity in artist Spotify popularity
(Figure 2) and release decade (Figure 3).
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Song quality measures</title>
      <p>Since there are 3 possible ratings and they appear in sequential order, we define two natural
conditional probabilities, the like rate and the superlike rate. If ,  and  are the number
of dislikes, likes and superlikes for a song , we define the like rate for song  as:
 = + ++  =  (|)</p>
      <p>= +  =  (|)
represents the conditional probability that a song is superliked, given that it is liked. Since
superlikes require users to invest more time and the songs are saved to a playlist, we can assume
that they indicate a more significant valuation on the part of the user.</p>
      <p>In Figure 4, we compute the like rates of songs based on their first 100 ratings and plot
them against the like rates based on the next 100 ratings for that same song (note that we only
considered songs with more than 200 ratings of which there are 669). This shows that despite
the diversity of tastes of listeners, these two quality metrics are consistent: if a song scores
highly among 100 listeners, it will score highly among the next 100 listeners.</p>
      <p>Having shown consistency in the like and superlike rates for individual songs, we now
aggregate like and superlike rates by grouping by decade of release and Spotify popularity.
Both the like rates and superlike rates are increasing with age (see Figure 5), possibly due to a
survivorship bias. Notice in Figure 6, the like rate is monotonic in artist popularity: this seems to
justify the popularity of the most streamed artists with more than 70 million monthly listeners
(the likes of Justin Bieber, Taylor Swift, Drake, Doja Cat, Bad Bunny and Ed Sheeran), according
to this quality metric. More surprisingly, the superlike rates seem to peak at a popularity of 70,
which typically corresponds to artists with a few million monthly listeners (the likes of Krewella,
Louis Tomlinson, The Marias, Khruangbin, Sting, Raekwon and Kaytranada). These results
seem to indicate that like rates measure a type of familiarity which grows with popularity, while
superlike rates measure a type of excitement that comes from deeper discovery, which is likely
to happen for lesser known artists. If that is the case, many of these middle class artists may be
more deserving of exposure than many artists in the 90-100 popularity range.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>
        Many fairness metrics have been proposed in the information retrieval literature [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and,
from the point of view of algorithmic design, fairness to producers is often framed as a constraint
on the utility of the consumers. In this paper, we show empirically how fairness to artists may
not need to come at the expense of music listeners. The fact that some under-exposed artists
are superliked at higher rates than top celebrities, indicates that an algorithmic system that
gives these artists more exposure may increase the utility of its users.
      </p>
      <p>Establishing fairness of exposure on streaming apps requires transparency. Artists will only
feel the system is fair, if they understand why some songs go viral, while others don’t. We
propose achieving this transparency by defining metrics based on human feedback, with strict
principles on the interfaces collecting the feedback and consistency tests on the quality metrics.
Only then can we deploy them as auditing or regulatory tools to keep algorithms accountable.</p>
      <p>The notion of quality metrics for works of art is likely to be controversial. The old saying that
“in matters of taste, there can be no disputes” will always rear its head. Despite the challenge in
defining tools for assessing creative works, fairness for the creative class can only be achieved
if artists and platforms agree on transparent notions of music quality. We argue in this paper,
that these metrics need to focus on aggregating critical human opinions, rather than simply
measuring past exposure and replicating it into the future.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A. J. B.</given-names>
            <surname>Chaney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Stewart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. E.</given-names>
            <surname>Engelhardt</surname>
          </string-name>
          ,
          <article-title>How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility</article-title>
          ,
          <source>in: Proceedings of the 12th ACM Conference on Recommender Systems</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>224</fpage>
          -
          <lpage>232</lpage>
          . URL: http: //arxiv.org/abs/1710.11214. doi:
          <volume>10</volume>
          .1145/3240323.3240370, arXiv:
          <fpage>1710</fpage>
          .11214 [cs, stat].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferraro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Serra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <article-title>Artist and style exposure bias in collaborative ifltering based music recommendations</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1911</year>
          .04827.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>Dinnissen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bauer</surname>
          </string-name>
          , Fairness in Music Recommender Systems:
          <string-name>
            <given-names>A</given-names>
            <surname>Stakeholder-Centered Mini</surname>
          </string-name>
          <string-name>
            <surname>Review</surname>
          </string-name>
          ,
          <source>Frontiers in Big Data</source>
          <volume>5</volume>
          (
          <year>2022</year>
          )
          <article-title>913608</article-title>
          . URL: https://www.frontiersin.org/ articles/10.3389/fdata.
          <year>2022</year>
          .913608/full. doi:
          <volume>10</volume>
          .3389/fdata.
          <year>2022</year>
          .
          <volume>913608</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferraro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Serra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bauer</surname>
          </string-name>
          ,
          <article-title>What is fair? Exploring the artists' perspective on the fairness of music streaming platforms</article-title>
          ,
          <year>2021</year>
          . URL: http://arxiv.org/abs/2106.02415, arXiv:
          <fpage>2106</fpage>
          .02415 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Raj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Ekstrand</surname>
          </string-name>
          , Comparing fair ranking metrics,
          <year>2022</year>
          . arXiv:
          <year>2009</year>
          .01311.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Abdollahpouri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mansoury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Burke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mobasher</surname>
          </string-name>
          ,
          <article-title>The unfairness of popularity bias in recommendation</article-title>
          ,
          <year>2019</year>
          . arXiv:
          <year>1907</year>
          .13286.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Diaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mitra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Ekstrand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Biega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Carterette</surname>
          </string-name>
          ,
          <article-title>Evaluating Stochastic Rankings with Expected Exposure</article-title>
          ,
          <source>in: Proceedings of the 29th ACM International Conference on Information &amp; Knowledge Management</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>275</fpage>
          -
          <lpage>284</lpage>
          . URL: http://arxiv.org/abs/
          <year>2004</year>
          . 13157. doi:
          <volume>10</volume>
          .1145/3340531.3411962, arXiv:
          <year>2004</year>
          .13157 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Joachims</surname>
          </string-name>
          , Fairness of Exposure in Rankings,
          <source>in: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>2219</fpage>
          -
          <lpage>2228</lpage>
          . URL: http://arxiv.org/abs/
          <year>1802</year>
          .07281. doi:
          <volume>10</volume>
          .1145/3219819.3220088, arXiv:
          <year>1802</year>
          .07281 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Biega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. P.</given-names>
            <surname>Gummadi</surname>
          </string-name>
          , G. Weikum, Equity of Attention: Amortizing Individual Fairness in Rankings,
          <source>in: The 41st International ACM SIGIR Conference on Research &amp; Development in Information Retrieval</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>405</fpage>
          -
          <lpage>414</lpage>
          . URL: http://arxiv.org/abs/
          <year>1805</year>
          . 01788. doi:
          <volume>10</volume>
          .1145/3209978.3210063, arXiv:
          <year>1805</year>
          .01788 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Patro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Porcaro</surname>
          </string-name>
          , L. Mitchell,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zehlike</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Garg</surname>
          </string-name>
          ,
          <article-title>Fair ranking: a critical review, challenges</article-title>
          , and future directions,
          <year>2022</year>
          . URL: http://arxiv.org/abs/2201.12662, arXiv:
          <fpage>2201</fpage>
          .12662 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tan</surname>
          </string-name>
          , S. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Fairness in recommendation: A survey,
          <year>2022</year>
          . arXiv:
          <volume>2205</volume>
          .
          <fpage>13619</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fenton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Caverlee</surname>
          </string-name>
          ,
          <article-title>Fairness among New Items in Cold Start Recommender Systems</article-title>
          ,
          <source>in: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , Virtual Event Canada,
          <year>2021</year>
          , pp.
          <fpage>767</fpage>
          -
          <lpage>776</lpage>
          . URL: https://dl.acm.org/doi/10.1145/3404835.3462948. doi:
          <volume>10</volume>
          .1145/ 3404835.3462948.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>A. B. Melchiorre</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Rekabsaz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Parada-Cabaleiro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Brandl</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Lesota</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Schedl</surname>
          </string-name>
          ,
          <article-title>Investigating gender fairness of recommendation algorithms in the music domain</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>58</volume>
          (
          <year>2021</year>
          )
          <article-title>102666</article-title>
          . URL: https://linkinghub.elsevier.com/retrieve/ pii/S0306457321001540. doi:
          <volume>10</volume>
          .1016/j.ipm.
          <year>2021</year>
          .
          <volume>102666</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferraro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Serra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bauer</surname>
          </string-name>
          , Break the Loop:
          <article-title>Gender Imbalance in Music Recommenders</article-title>
          ,
          <source>in: Proceedings of the 2021 Conference on Human Information Interaction and Retrieval</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <source>Canberra ACT Australia</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>249</fpage>
          -
          <lpage>254</lpage>
          . URL: https://dl.acm.org/doi/10.1145/ 3406522.3446033. doi:
          <volume>10</volume>
          .1145/3406522.3446033.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Shakespeare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Porcaro</surname>
          </string-name>
          , E. GÃ³mez, C. Castillo,
          <source>Exploring Artist Gender Bias in Music Recommendation</source>
          ,
          <year>2020</year>
          . URL: http://arxiv.org/abs/
          <year>2009</year>
          .01715, arXiv:
          <year>2009</year>
          .01715 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Epps-Darling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. T.</given-names>
            <surname>Bouyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cramer</surname>
          </string-name>
          ,
          <article-title>Artist gender representation in music streaming</article-title>
          ,
          <source>In Proceedings of the 21st ISMIR Conference. International Society for Music Information Retrieval</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kearns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z. S.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <article-title>Meritocratic fairness for cross-population selection</article-title>
          , in: D.
          <string-name>
            <surname>Precup</surname>
            ,
            <given-names>Y. W.</given-names>
          </string-name>
          <string-name>
            <surname>Teh</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 34th International Conference on Machine Learning</source>
          , volume
          <volume>70</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1828</fpage>
          -
          <lpage>1836</lpage>
          . URL: https://proceedings.mlr.press/v70/kearns17a.html.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Joseph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kearns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Morgenstern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roth</surname>
          </string-name>
          , Fairness in Learning: Classic and
          <string-name>
            <given-names>Contextual</given-names>
            <surname>Bandits</surname>
          </string-name>
          ,
          <year>2016</year>
          . URL: http://arxiv.org/abs/1605.07139, arXiv:
          <fpage>1605</fpage>
          .07139 [cs, stat].
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <article-title>The piki dataset (</article-title>
          <year>2023</year>
          ). URL: https://github.com/sstoikov/piki-music-dataset.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <article-title>The yelp dataset (</article-title>
          <year>2023</year>
          ). URL: https://www.yelp.com/dataset.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Harper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <article-title>The movielens datasets: History and context</article-title>
          ,
          <source>ACM Trans. Interact. Intell. Syst</source>
          .
          <volume>5</volume>
          (
          <year>2015</year>
          ). URL: https://doi.org/10.1145/2827872. doi:
          <volume>10</volume>
          .1145/2827872.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Koren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Volinsky</surname>
          </string-name>
          ,
          <article-title>Collaborative filtering for implicit feedback datasets, ICDM '08</article-title>
          , IEEE Computer Society, USA,
          <year>2008</year>
          , p.
          <fpage>263</fpage>
          -
          <lpage>272</lpage>
          . URL: https://doi.org/10.1109/ICDM.
          <year>2008</year>
          .
          <volume>22</volume>
          . doi:
          <volume>10</volume>
          .1109/ICDM.
          <year>2008</year>
          .
          <volume>22</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Stoikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <article-title>Evaluating music recommendations with binary feedback for multiple stakeholders</article-title>
          , in: H.
          <string-name>
            <surname>Abdollahpouri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Elahi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Mansoury</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sahebi</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Nazari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Chaney</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          Loni (Eds.),
          <source>Proceedings of the 1st Workshop on Multi-Objective Recommender Systems (MORS</source>
          <year>2021</year>
          )
          <article-title>co-located with 15th ACM Conference on Recommender Systems (RecSys</article-title>
          <year>2021</year>
          ), Amsterdam, The Netherlands,
          <year>September 25</year>
          ,
          <year>2021</year>
          , volume
          <volume>2959</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2021</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2959</volume>
          /paper9.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>