<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Efects of Human-curated Content on Diversity in PSM: ARD-M Dataset</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marcel Hauck</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ahtsham Manzoor</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sven Pagel</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ARD Online</institution>
          ,
          <addr-line>Isaac-Fulda-Allee 1, 55124 Mainz</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Johannes Gutenberg University Mainz</institution>
          ,
          <addr-line>Jakob-Welder-Weg 9, 55128 Mainz</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Mainz University of Applied Sciences</institution>
          ,
          <addr-line>Lucy-Hillebrand-Str. 2, Mainz, Germany, 55128</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Klagenfurt</institution>
          ,
          <addr-line>P.O. Box 1212, Klagenfurt, Austria, 43017-6221</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Public service media (PSM) providers, like ARD, are continuously undergoing digital disruptions due to the ubiquitous availability of the internet and technological breakthroughs. To keep up with such advancements, PSM providers face new challenges, particularly in adopting artificial intelligence technologies such as recommender systems. However, due to the heterogeneous nature of the content, the unavailability of PSM datasets, and human involvement in content curation and decision-making, it remains unclear how internal domain knowledge experts (e.g., publishers and editors) can be bridged with external technologists. This poses further challenges in creating recommendation technology for PSM providers that automatically deliver relevant, transparent, and diverse content for their consumers. To this end, real-world datasets can be pivotal in investigating and closing this information gap. In this work, therefore, we release a real-world dataset for researchers containing features like items and their metadata, user interaction feedback, and complementary features like item position, timestamps, and item categories, in the movie domain. Furthermore, through a series of analyzes, we show how diferent data sources (publisher, editors, TMDB) and metadata features (genres, keywords, plot) impact the diversity of items. Finally, we present an outlook of future directions to be explored by the research community using the ARD-M dataset. This is an important step forward to strengthen relevance, transparency and diversity in the recommendation process, as part of the PSM core values.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Public service media providers</kwd>
        <kwd>recommender systems</kwd>
        <kwd>dataset</kwd>
        <kwd>user interaction features</kwd>
        <kwd>diversity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>in competition with Netflix, Amazon Prime, and other
online streaming providers. One reason could be that PSM
As of today, over 90%1 of the population in Western coun- platforms provide a variety of content, for example, daily
tries are internet users (e.g., 92% in the USA). Moreover, news, stories, movies, sports, or cultural shows. Thus,
around 90%2 of them are also smartphone users (e.g., 89% exercising recommendation technology for such mixed
in the USA). Such ubiquitous advancements have led to content delivery requires modern yet holistic solutions.
the digital transformation of internet services such as Therefore, PSM providers traditionally rely on editors as
e-commerce or streaming of video content. Additionally, domain experts for delivering the content [4]. On the one
advancements in technology, particularly in the media hand, editors play a key role in shaping content
delivindustry, have largely replaced editorial recommenda- ery procedures by providing feedback and implementing
tions with algorithmic content recommendations [1, 2, 3]. policies that align with legal and ethical mandates and
A prominent outcome of recent transformations is an guidelines [5]. Furthermore, algorithmic content
selecacross-the-board shift towards consuming media content tion and user personalization can introduce risks and
sopredominantly through online channels. cietal threats (e.g., misinformation, discrimination, bias,</p>
      <p>In this regard, Public Service Media (PSM) platforms, and privacy issues), see also [6].
like ARD Mediathek, are still striving to incorporate rec- PSM providers are generally committed to asserting
ommendation technology and maintaining market shares core values such as relevance, transparency, and diversity
while ensuring public trust. To this end, editors’ domain
Workshop on Learning and Evaluating Recommendations with Impres- knowledge and role in ensuring such values while
con*siConosrr(eLsEpRoIn)d@inRgeacuSythso2r0.23, September 18-22 2023, Singapore tent curation and selection ofer a great deal to retrospect
$ mhauck@uni-mainz.de (M. Hauck); ahtsham.manzoor@aau.at how recommender systems are designed. For instance,
(A. Manzoor); sven.pagel@hs-mainz.de (S. Pagel) digitizing media processes has created numerous
oppor0000-0003-1105-1519 (M. Hauck); 0000-0001-9418-7539 tunities to collect and analyze vast amounts of audience
(A. Manzoor)</p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License and consumption data. This data can be leveraged to
1 hCPWrEooUrctkReshtdoinpgpssIhStpN:/c1e:6u1r3-w/-0s.o7r3g/daACtttEarirbUuetRpionoW4r.0toaInrlt.ekcrnsoahmtioon/parl e(PCpCrooBrYcte4s.0e/)d.diignigtsal(-C2E02U3R--juWlyS-.ogrlogb)al-statshot
icnutsetroemstiszoefsienrdviivciedsuaalncdocnosnutmenertsbwasheidleoandatphteinpgerhcuemivaend2https://www.gsma.com/mobileeconomy/ in-the-loop design practices. For this reason, the need
for externally available real-world industry datasets is each titled with a category such as “Newly available films ”
crucial for the research community to make practical use or “Movies for the whole family”. So far, mainly human
of domain experts’ knowledge in recommender systems. editors are responsible for making decisions regarding</p>
      <p>In addition, relevance, transparency, and diversity of the specific position and listing of an item in a specific
item recommendations are highly vigilant topics in rec- list. Specifically, given a large pool of content provided
ommender systems research, see e.g., [7, 8, 9]. Relevance by publishers in regional broadcasters (so-called
“Lanrefers to the degree to which recommended items align desrundfunkanstalten”), editors select and integrate a
with the user’s preferences, needs, or interests, see also specific item in a particular list along with curated
mixed[10, 11]. Research has shown that by ofering relevant quality metadata features like keywords or genres. With
recommendations, these systems can enhance user satis- the goal of building a bridge between algorithmic and
hufaction and engagement and ultimately drive conversion man curation of content in PSM platforms, we collected
rates [12]. Moreover, diversity plays an equally impor- data for the first whole week of January 2023.
tant role in recommendation algorithms. It refers to the Apart from items and their metadata (genres, plots, and
variety and heterogeneity of items suggested to users keywords), we provide two user interaction features, i.e.,
preventing over-specialization and filter bubbles, where item views and item clicks. For simplicity, we represent
users are only exposed to a limited set of items [9]. Also, item impressions as item views, whereas it remains an
diversity helps to address the recommendation bias issue open question whether such an assumption is plausible
[13]. Therefore, achieving an optimal balance between in the context of an online streaming platform. In
addirelevance and diversity is crucial for delivering high- tion, we provide complementary features like timestamp,
quality recommendations in PSM platforms that cater item position in a particular list. The descriptive
statisto users’ diverse interests while providing exposure to tics of our dataset are shown in Table 1. The numbers
multifarious and potentially serendipitous content. of distinct values for each item metadata feature are in</p>
      <p>However, given the lack of real-world, human-curated Table 2. For both user interaction features, we show the
PSM datasets, it is unclear how internal editorial domain data characteristics like sparsity and density in Table 3.
knowledge and editors’ decisions can be bridged with rec- We believe a rich user feedback dataset will open new
ommender systems from external technologists serving directions to research recommendation technology for
users’ needs while achieving market share and re-growth. PSM providers by external technologists.</p>
      <p>In this work, therefore, we first collected a real-world To protect user privacy, we utilized random numbers
dataset in the movie domain containing items, their meta- as visitor and session identifiers to ensure decoupling
data, and users’ interaction feedback records, i.e., item from production data. To ensure data protection for rare
impressions and item clicks. Subsequently, using TMDB3 events (e.g., usage in the middle of the night), we rounded
database, we enriched the dataset with additional meta- timestamps to the closest minutes. Finally, to enable
data features like movie plots and genres. Second, with cross-source exploratory analysis, we enriched the data
the help of a series of analyzes, we show how item diver- with a summary of the plots, keywords, and genres from
sity and similarity vary in intra-list navigational inter- the community-built TMDB5 movies database. With a
faces over the period of one week. Finally, we present variety of metadata collected from multiple sources, i.e.,
an outlook of future directions for the recommender sys- publishers, editors, and TMDB, we believe new
algorithtems research community and release the ARD-M dataset mic dimensions of content or collaborative filtering-based
online under CC BY-NC license. approaches to recommendations can be explored.
Moreover, with a humans-in-the-loop approach, domain
experts’ knowledge can be leveraged in various ways, for
2. Dataset Collection example, by investigating content categorization, user
behavior, and item diversity and similarity in lists. We
ARD is a PSM provider in Germany, which ofers a vari- release our dataset publicly at https://www.kaggle.com
ety of online streaming content, including movies4. On /datasets/marcelhauck/ardmovies/ with kind permission
the movie page, items are ordered within horizontal lists, of ARD Online and its Data Protection Oficer.</p>
      <sec id="sec-1-1">
        <title>3https://www.themoviedb.org/</title>
        <p>4https://www.ardmediathek.de/filme</p>
      </sec>
      <sec id="sec-1-2">
        <title>5https://www.themoviedb.org/</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Comparison with Existing</title>
    </sec>
    <sec id="sec-3">
      <title>Datasets</title>
      <p>aware recommendations. However, a social network of
users sharing common preferences or interests can be
created with the ARD-M dataset as well, e.g., by
includIn this section, we briefly discuss how our ARD-M dataset ing consumed items, and their metadata features like
compares to existing predominant datasets available in keywords, or genres.
the literature on recommender systems. Overall, we ob- There are additional datasets that with impressions on
serve a number of datasets available in the literature to diferent domains. The “Microsoft News Dataset (MIND)”
carry out research on recommender systems. For ex- contains data from the Microsoft News website with
inample, the predominant tendency is to design recom- teractions (impressions, clicks) and metadata features
mender algorithms and evaluate them using one or more like title, body, and category [24]. The “ContentWise
datasets. Technically, a recommender system can serve Impressions” dataset includes interactions of movies and
multiple user objectives such as providing relevant item TV series that are collected from a recommender system
recommendations, endorsing transparency via explana- on an Over-The-Top media service [25]. The “FINN.no
tions, or ensuring diversity in the recommendation pro- slate dataset” contains data from the real estate
marketcess. Recommender system designers thereby leverage place FINN.no with sequential interaction features on
datasets like ARD-M to attain and assess such objectives recommendations and search results in lists [26].
by implementing specific computational tasks like Next In addition, our ARD-M dataset ofers a variety of
feaitem-recommendation, Next session prediction, Top-N tures, which can be further used to implement specific
recommendations, or Explanation generation. Overall, computational tasks to achieve certain recommender
sysas explained in Section 2, our ARD-M dataset includes tem objectives. For example, in the context of horizontal
various features like items, users, metadata (genres, key- list interfaces like in the case of commercial online
streamwords, plots), user interaction feedback (impressions and ing platforms such as Netflix, or Amazon Prime Video,
clicks), and complementary features (list IDs, item posi- we provide specific item positions in a specific list. Such
tions, timestamps, and item duration). ifne-grained details can be modeled in the recommender</p>
      <p>To this end, our ARD-M dataset resembles various algorithm for investigating aspects, for example, CTR
opdatasets in terms of available features and therefore timization with respect to item positions, user behavior
can serve multiple objectives in the recommendation analysis, and design optimization of diferent user
interprocess. For example, e-commerce datasets like YOO- faces. Furthermore, our dataset is exceptionally dense
CHOOSE, used in [14], DIGINETICA [15], and Retail- (42.9%) in terms of user impressions, presenting
varirocket [16] ofer similar user interaction features, and ous ways to aid personalization, user segmentation, or
are extensively used, specifically, for collaborative filter- comparing diferent recommendation strategies. Finally,
ing or session-based recommendations. In the movies unlike other datasets, we integrate additional metadata
domain, MovieLens 10M [17, 18], Douban [19], and Net- obtained through multiple sources, thereby providing
oplfix [ 20, 21] datasets share common features like cate- portunities to conduct metadata-based content modeling.
gories, keywords, and timestamps, yet include additional To the best of our knowledge, there is no dataset
cuuser interaction feedback, i.e., user ratings. Similarly, in rated from a PSM platform, where decisions regarding
the movies domain, unlike ARD-M, datasets like Yahoo! diverse content selection are made by domain experts.
Movies [22], EachMovie [19] and FilmTrust [23] include Thus, with a humans-in-the-loop approach, we believe
complementary user features like gender, age, and trust ARD-M extends substantial opportunities to investigate
network, which can be further used to design context- novel recommender system use cases.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Intra-list Similarity (ILS)</title>
    </sec>
    <sec id="sec-5">
      <title>Analysis</title>
      <p>Award winning movies 0.30 0.51 0.46 0.08 0.37
With a goal of understanding the diversity and
similarity of items curated by human editors, inspired by work Short films 0.29 0.80 0.13 0.10 0.52 0.19 0.15
in [27], we compute the intra-list similarity (ILS) score. CEoumroppeetaintioonnlfiinlme sfeosftitvhael 1.00 0.36 0.13 1.00 0.46 0.21 0.19
Diversity can be seen as the opposite of similarity. Specif- Dramas | Movies 0.54 0.62 0.32 0.07 0.49 0.21 0.22
ically, we compute the ILS for each pair of items in a
list with the available metadata features, e.g., plots and Enchanting fairy tales 0.89 0.91 0.73 0.87 0.55 0.22 0.30
genres. We separately use the Jaccard Coeficient [28] for Not online for much longer 0.42 0.37 0.27 0.13 0.50 0.22 0.27
similarity calculation of genres and keywords from all Arthouse Movies 0.25 0.47 0.31 0.12 0.51 0.22 0.21
data sources (i.e., publishers, editors, and TMDB). For the
Jaccard Coeficient, a value of 0 shows low similarity (= Newly available movies 0.36 0.39 0.21 0.07 0.48 0.22 0.23
high diversity), and 1 as high similarity (= low diversity). Movies for the whole family 0.33 1.00 0.32 0.05 0.57 0.23 0.19
For plots, we first produced embeddings using a Sentence Movies 0.31 0.36 0.29 0.12 0.50 0.23 0.22
Transformer [29] from Hugging Face6 and computed
cosine similarity score for plots between each pair of items Movies to relax 0.49 0.41 0.36 0.07 0.45 0.24 0.30
in a list. For the cosine similarity score, a value of − 1 Currently popular movies 0.50 0.36 0.15 0.16 0.49 0.25 0.22
shows low similarity (= high diversity), and 1 as high Thrillers and detective stories 0.40 0.68 0.12 0.10 0.52 0.25 0.17
similarity (= low diversity). More Drama Movies 0.49 0.43 0.40 0.12 0.52 0.25 0.25</p>
      <p>Figure 1 shows the ILS for every list in the dataset
based on Jaccard coeficient (left part) and cosine simi- Movies that tell history 0.42 0.48 0.38 0.19 0.50 0.31 0.25
larity (right part). Editors frequently removed or added Current TV movies 0.59 0.34 0.18 0.35 0.48 0.46 0.24
items to those lists, leading to a changed ILS. Therefore,
tthhiimgehevseaprlau(nees..gAi.n,sKtcheaeynwfigbuoerredsesaerfenr,othtmheeaEivdneidtroiavrgised) uionarlt hmloewecteaordmacptoalnehtseaiss- frrseeonGm lirsebuhP frrseeonGm itrsodE frrseeonGm TBDM frrsyeoodKwm lirsebuhP frrsyeoodKwm itrsodE CosinlftrooPmeslirsebuhPim.lftrooPmforTBDM...
tency (e.g., Keywords from Publishers). Editors select Jaccard coefficient for ...
content for lists so that they have a high semantic simi- diverse similar diverse similar
larity. Therefore, this is also reflected in the consistent 0.0 0.5 1.0 1 0 1
similarity based on their curated keywords. The situation
is diferent for the numerous regional broadcasters with Figure 1: ILS for all lists by their metadata features
human publishers, who generate and upload keywords
in diferent ways (e.g., use of default values or structured
keyword hierarchies). The average ILS scores for all
metadata features are shown in Table 4. The results suggest used for computing the correlation. Results indicate
that editorial features based on Jaccard coeficient (genres, that publishers’ metadata features (keywords and
genkeywords) with a value range from 0 (diverse) to 1 (sim- res) correlate strongly (0.91). A similar case is found
ilar) are reasonably better represented (genres = 0.52, in TMDB features (genres and plot = 0.53) and
editokeywords = 0.50) for the semantic connection between rial features (keywords and genres = 0.49), even when
items in a list than genres (0.30) from TMDB and key- their average ILS scores are diferent, see also Table 4.
words from publishers (0.21). As described before, the This implies that editors can rely on additional data
cuILS scores based on the movie plot are calculated using ration sources like publishers with substantial domain
the pairwise cosine similarities with a value range from knowledge and expertise while selecting the content for
-1 (diverse) to 1 (similar). Therefore, the results (avg. a particular list. Also, diferent metadata sources can
≈ 0.22) reveal a slightly similar content representation. be uniquely combined in a recommendation model to</p>
      <p>Next, we conduct a correlation analysis to investi- assert values like the diversity of items with a
humansgate the relationships of content diversity between all in-the-loop design approach. We present additional
anmetadata sources (publishers, editors, TMDB); see re- alyzes, e.g., the variation of the ILS scores for all
catesults in Figure 2. Specifically, first, we compute the gories in time, in our Kaggle data repository at https:
average ILS score for each item list and for all meta- //www.kaggle.com/code/marcelhauck/data-analyses.
data features. Subsequently, the average ILS scores are</p>
      <sec id="sec-5-1">
        <title>6https://huggingface.co/sentence-transformers/distiluse-base-mul</title>
        <p>tilingual-cased-v1</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusion and Future Research</title>
    </sec>
    <sec id="sec-7">
      <title>Directions</title>
      <p>Public Service Media (PSM) providers have been
continuously adapting to technological developments and mostly
provide digital services for their consumers. However,
the integration of external recommendation technology
for heterogeneous content presents further challenges
in incorporating internal domain experts’ knowledge to
assert core values such as content diversity while
projecting business and consumer value. To this end, real-world
PSM datasets can be pivotal to fostering research in this
domain. Therefore, in this work, we release a real-world
ARD-M dataset created from ARD Mediathek and analyze
the content regarding diversity and similarity aspects.</p>
      <p>Finally, we highlight a few directions for the research
community to build on the ideas. First, datasets like ours
can be used to tailor services and content through
automated recommendations to the perceived interests of
individual consumers. Also, understanding users’
behavior and their experiences with consumption data can
further help in meaningfully devising ways for publishers
to create content that interests their audience. Second,
PSM providers are looking for ways to create common
yet holistic evaluation metrics that illustrate
algorithmic performance, the relevance of content, and business
value for both providers and consumers. We believe our
released dataset can pave the way for researching and
devising such novel metrics in the future. Finally, in PSM
providers, editors possess domain expertise relating to
content curation and selection while adhering to PSM
responsibilities like educative and unbiased content
delivery. Therefore, it is largely unclear without popularity
bias content features how public-service values can be
protected and infused in recommendation algorithms
development. We believe this work opens new avenues of
innovations and collaborations between PSM actors like
editors, researchers, and recommender system designers.
[28] S. Niwattanakul, J. Singthongchai, E. Naenudorn,
S. Wanapu, Using of jaccard coeficient for
keywords similarity, in: Proceedings of the
international multiconference of engineers and computer
scientists, volume 1, 2013, pp. 380–384.
[29] N. Reimers, I. Gurevych, Sentence-bert: Sentence
embeddings using siamese bert-networks, in:
Proceedings of the 2019 Conference on Empirical
Methods in Natural Language Processing, Association
for Computational Linguistics, 2019.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>