<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>INRA</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>about Migration</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andreea Iana</string-name>
          <email>andreea.iana@uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mehwish Alam</string-name>
          <email>mehwish.alam@telecom-paris.fr</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Grote</string-name>
          <email>alexander.grote@kit.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nevena Nikolajevic</string-name>
          <email>nevena.nikolajevic@kit.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katharina Ludwig</string-name>
          <email>katharina.ludwig@uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Müller</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christof Weinhardt</string-name>
          <email>weinhardt@kit.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heiko Paulheim</string-name>
          <email>heiko.paulheim@uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data and Web Science, University of Mannheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Media and Communication Studies, University of Mannheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Karlsruhe Institute of Technology</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Télécom Paris, Institut Polytechique de Paris</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>11</volume>
      <fpage>18</fpage>
      <lpage>22</lpage>
      <abstract>
        <p>News recommendation plays a critical role in shaping the public's worldviews through the way in which it filters and disseminates information about diferent topics. Given the crucial impact that media plays in opinion formation, especially for sensitive topics, understanding the efects of personalized recommendation beyond accuracy has become essential in today's digital society. In this work, we present NeMig, a bilingual news collection on the topic of migration, and corresponding rich user data. In comparison to existing news recommendation datasets, which comprise a large variety of monolingual news, NeMig covers articles on a single controversial topic, published in both Germany and the US. We annotate the sentiment polarization of the articles and the political leanings of the media outlets, in addition to extracting subtopics and named entities disambiguated through Wikidata. These features can be used to analyze the efects of algorithmic news curation beyond accuracy-based performance, such as recommender biases and the creation of filter bubbles. We construct domain-specific knowledge graphs from the news text and metadata, thus encoding knowledge-level connections between articles. Importantly, while existing datasets include only click behavior, we collect user socio-demographic and political information in addition to explicit click feedback. We demonstrate the utility of NeMig through experiments on the tasks of news recommenders benchmarking, analysis of biases in recommenders, and news trends analysis. NeMig aims to provide a useful resource for the news recommendation community and to foster interdisciplinary research into the multidimensional efects of algorithmic news curation.</p>
      </abstract>
      <kwd-group>
        <kwd>news corpora</kwd>
        <kwd>user data</kwd>
        <kwd>recommender system</kwd>
        <kwd>bias analysis</kwd>
        <kwd>aspect diversity</kwd>
        <kwd>knowledge graph</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The large volume of news published online each day creates an information overload that
exceeds the consumptive capacities of readers. At the same time, the digitalization of news
consumption has sparked an unprecedented development of personalized recommendation
(H. Paulheim)
algorithms, used by online news platforms to process the continuously growing quantities of
news and ofer readers personalized suggestions. However, models which are optimized to
maximize congruity to users’ preferences and past behavior, tend to produce recommendations
which are highly similar in content to previously read/clicked ones [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. By filtering out
news deemed irrelevant to the user, news recommenders control how and which information
is disseminated to the public [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Consequently, algorithmic news curation has a large impact
on opinion formation and voting behavior [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], as well as the potential to fuel conflicts and
polarization [
        <xref ref-type="bibr" rid="ref2 ref6 ref7 ref8 ref9">6, 7, 8, 2, 9</xref>
        ], especially when dealing with controversial subjects such as war,
climate change, or migration. In addition to the power of recommender systems to shape
people’s perception of the world, the fundamental role that media plays in today’s society as
a public forum [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ], have deemed analyzing and understanding the implications of news
curation algorithms necessary for modern democracies.
      </p>
      <p>
        While plethora of recommendation algorithms have been proposed in recent years, resources
available to analyze their efects beyond accuracy performance and in multilingual scenarios
are scarce. The existing body of work exhibits two main shortcomings: (1) news datasets for
recommendation (i) mostly focus on general or less sensitive topics (e.g., entertainment, fashion,
sports) [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ] and (ii) leave largely unexplored features which are critical for analyzing the
algorithmic creation of ”filter bubbles” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or underlying recommender biases (e.g., sentiment or
political orientation); (2) news collections for media analysis (e.g., fake news detection [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ],
news narratives analysis [
        <xref ref-type="bibr" rid="ref16">16, 17</xref>
        ], news bias detection [18, 19]) lack user information, and thus,
cannot easily be used in a recommendation scenario.
      </p>
      <p>In this paper, we introduce NeMig a bilingual news collection and knowledge graphs (KGs)
about refugees and migration, in German and English, as well as corresponding real-world user
data. We collect over 7K, and respectively, 10K articles from German and US news outlets. The
articles span a large political spectrum in order to cover various viewpoints on the topic. 1)
Compared to existing resources, NeMig (i) targets the same polarizing topic in two languages
and (ii) is annotated with the sentiment polarization of articles and political leanings of the
media outlets, in addition to subtopics and disambiguated named entities. 2) We construct
corresponding knowledge graphs (KGs) from the news’ text, metadata, and extracted named
entities, which we further expand with up to two-hop neighbors from Wikidata of the named
entities. We provide the KGs in diferent variants. 3) We collect and publish in an anonymized
fashion both explicit feedback and socio-demographic data for 3K users for each of the two
datasets. Although our user dataset is small compared to large benchmark news datasets, it
contains information about the users’ media consumption, political attitudes and interests,
personality traits and demographics, which makes the dataset valuable for studying the efects
of recommender systems beyond accuracy-based performance, e.g., on political polarization.
Furthermore, the user data constitutes a starting point for the generation of synthetic user
datasets, which comprise not only click behavior information, but also explicit data about the
users’ background and preferences.</p>
      <p>We demonstrate the utility of the resource through experiments on diferent downstream
tasks. Firstly, we benchmark several news recommenders on NeMig. Secondly, we investigate
biases and polarization efects of the benchmarked recommendation models, and evaluate the
contribution of various features to the quality of the KG. Lastly, we analyze news trends by
examining the evolution over time and political orientation of the most frequent entities in our
corpora, identifying correlations between the patterns of evolution and worldwide events.</p>
      <p>NeMig is available in two versions, for research purposes: i) we release the full news datasets
in N-Triples format, under restricted access on Zenodo [20], as the news bodies can only be
provided upon request due to copyright policies, ii) NeMig, with all features but without the
news bodies, is freely available, under a CC BY-NC-SA 4.0 license, in tabular format on GitHub.1
The user data is available in anonymized tabular format on both Zenodo [20] and GitHub.2</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        News Recommendation Benchmarks. Various datasets have been constructed for developing
and benchmarking news recommender systems (NRS). Plista [21] comprises a collection of over
70K German articles gathered from 13 news portals, as well as click data for over 14 million users.
Globo [22, 23] is a Portuguese dataset for news recommendations collected from Globo.com,
and contains data for over 300K users and 46K news distributed in 1.2 million sessions. The
Norwegian dataset Adressa [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] contains news collected from the Adresseavisen’s news portal
and is provided in two variants: a light version with over 11K articles and click data for 561K
users, and a larger one with more than 48K articles and clicks for 3 million users. In addition to
metadata (e.g., categories, authors), the articles are annotated with named entities. MIND [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
the most recent large-scale dataset, is constructed from user click logs of Microsoft News. It
covers 1 million users and more than 161K news articles in English. Moreover, the dataset is
enriched with named entities linked to Wikidata.
      </p>
      <p>News Knowledge Graphs. Knowledge graphs have been shown to model knowledge-level
connections between news that cannot be captured by purely text-based models [24, 25]. In the
context of the Newsreader3 project, focusing on the multilingual processing of news articles,
the authors generate an event-centric KG from a multilingual news dataset that describes which
events took place, where, when, and who was involved in them [26, 27]. News Graph [28] is a
graph constructed specifically for news recommendation. It is generated from a news corpus
collected from MSN News containing over 621k articles, 594k news entities, as well as user-item
interaction logs. The graph contains news content, user behaviors, and news topic entities
enriched with neighboring triples from Microsoft Satori, along with three types of collaborative
relations for entities, i.e., co-occurring in the same news, clicked by the same user, and clicked
by the same user in the same browsing session. The use of semantic KGs for the production,
distribution, and consumption of news is summarized in [29].</p>
      <p>
        News Datasets for Media Discourse Analysis. The advent of news platforms as an ubiquitous
means of information for Internet users has established online news as a fundamental source
for various media discourse analysis tasks. For instance, Horne et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] has created a large
dataset of political news collected from mainstream and alternative sources, enriched with
content-based and social media engagement features, that can be used for news and engagement
characterization, news attribution, and content copying, or exploration of news narratives.
      </p>
      <sec id="sec-2-1">
        <title>1https://github.com/andreeaiana/nemig_nrs/tree/main/data 2https://github.com/andreeaiana/nemig/tree/main/data/user_data 3http://www.newsreader-project.eu</title>
        <p>
          Similarly, Lim et al. [19] and Färber et al. [18] collected and annotated articles with diferent
types of bias at the sentence level for analyzing news biases. In a related line of work, researchers
have designed datasets for monolingual [
          <xref ref-type="bibr" rid="ref14 ref16">14, 16</xref>
          ] and multilingual cross-domain [30] fake news
detection, or for distinguishing between fake and satire news published online [31, 32].
Limitations of Current News Resources. Existing news datasets have several shortcomings.
On the one hand, the large monolingual recommendation benchmarks benefit from sizable
user interaction data, but fail to explore news (e.g., sentiment orientation, political leaning)
or user features (e.g., demographics, political interests) which can be exploited in the analysis
of recommender system efects. On the other hand, the news collections for media discourse
analysis contain rich annotations (e.g., news biases, social media features, social context,
spatiotemporal information), but are ill-equipped to be seamlessly utilized in training recommendation
algorithms due to the lack of user feedback data. In contrast to these resources, NeMig comprises
a (i) bilingual news collection, annotated not only with text and metadata information, but also
with subtopic, sentiment and political orientation information, and (ii) implicit user feedback,
demographic, and political data. Moreover, we use Wikidata as an open source for entity linking,
and not a commercial knowledge base, as done in [28].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The NeMig News Corpora</title>
      <sec id="sec-3-1">
        <title>3.1. Data Collection</title>
        <p>We construct the news corpora by crawling news articles from 40 German4, and respectively 45
US5, legacy and alternative media outlets, selected by a team of researchers from the media and
communication domain. The news are sampled from the entire political spectrum, such as to
cover diferent viewpoints on the topic. We include German articles published between Jan.
1, 2019 and Dec. 31, 20216, and English ones between Jan. 1, 2021 and July 1, 20227. For each
language, we identify relevant articles using 13 keyword stems representative of the migration
topic, such as flüchtl* , asyl*, or migration* for the German outlets, and refugee*, asylum seeker*,
or migrant* for the US sources. Moreover, only articles written in either German or English
(for the German and English corpora, respectively), with a minimum length of 150 words, and
containing at least two keywords stems have been included in the two corpora in order to avoid
foreign language articles, disclaimers, advertisements, or reader comments. From the resulting
raw datasets, we further excluded duplicates, videos, live tickers (i.e., articles continuously
updated with news headlines), and outliers8 (i.e., articles which are abnormally long or short).
The datasets contain 7,346 (German), and respectively 10,814 (English) filtered articles. Figs. 1a
and 1b show the distribution of the articles over time and over media outlets. We observe that
the news is unevenly distributed over various publishers, with smaller or niche media outlets
being underrepresented. In comparison, with the exception of a few outlets, the articles are
relatively evenly distributed over the diferent time periods.
4The implementation of the German news crawlers is available at https://github.com/andreeaiana/german-news
5The implementation of the English news crawlers is available at https://github.com/andreeaiana/us-news
6The data was collected in Dec. 2020 and Jan. 2021, and updated in Jan. 2022.
7The data was collected in Aug. 2022.
8Outliers were identified by using two standard deviations from the mean.</p>
        <p>(a) German news corpus.</p>
        <p>Each news article contains a title, a body, and if available, an abstract. Additionally, its
metadata specifies its provenance (news outlet, URL), publishing and modification dates, authors,
and keywords provided by the media outlet. Furthermore, we classify the political leaning of the
news outlets. Concretely, we categorize the US sources into three political orientation classes,
namely left , center, and right, based on the AllSides Media Bias Chart9. Since no corresponding
classification exists for the German media, the same researcher team classifies the outlets into
four political orientation groups: left , center, right, conspiracy.10 We find (see Fig. 2) that the
distribution of news outlets over the political orientation classes difers between the German and
the English datasets. While the majority of outlets for German are center media, the majority of
outlets for English are left. On the article level, however, the distribution for the English dataset
is drastically diferent and similar between the two corpora.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Data Annotation</title>
        <p>
          Sentiment Analysis. Sentiment polarization information is leveraged by fairness-aware news
recommenders to produce sentiment-diverse or agnostic recommendations [
          <xref ref-type="bibr" rid="ref1">1, 33</xref>
          ]. Furthermore,
such information can be used to investigate whether algorithmic news curation propagates or
encourages the sentiment polarization of readers [
          <xref ref-type="bibr" rid="ref3">34, 3, 35</xref>
          ] or, more generally, the sentiment
bias of diferent news outlets [ 36, 37]. We use a multilingual XLM-R [38] language model,
9https://www.allsides.com/media-bias/media-bias-chart
10Disclaimer: We excluded the outlet Man Tau from further analysis regarding the political orientation as it cannot
be categorized in any of the four groups, but we include the corresponding news in the dataset.
11The implementation of the annotation and KG construction pipeline, and the intermediary files produced by the
annotation process are available at https://github.com/andreeaiana/nemig
12Sub-topic extraction constitutes the only exception, as it requires setting parameters (e.g., threshold for pruning
topics) based on intermediary results. We run this process separately and integrate its output in the pipeline.
trained on 198M tweets and fine-tuned for sentiment analysis [ 39], to classify news into one of
three sentiment classes: positive, neutral, or negative. We perform sentiment classification on
the concatenation of the news’ title and abstract.13 Table 1 shows the distribution of news over
sentiment classes in the two datasets. For both corpora, only a small number of articles have
a positive sentiment. Moreover, while the distribution of articles with neutral and negative
sentiments is relatively balanced in the German corpus, the English corpus has significantly
more negative articles.
        </p>
        <p>Sub-topic Modeling. We observe that news cover discourses about refugee migration from
diferent geographical areas or related to various political events. In order to extract
subtopics from each dataset, we use the neural topic modeling approach proposed in [40], with
a pre-trained English Sentence Transformer for the English news, and multilingual Sentence
13Note that we use the concatenation of title and the article’s first sentences as input when abstracts are not provided.</p>
        <p>(b) English news corpus.</p>
        <p>Transformer [41] for the German ones. We assign labels to the resulting topics based on the
topmost representative terms of the respective cluster of documents. We group articles that
cannot be assigned to any topic appearing in at least 15 news into a separate cluster. We extract
25 sub-topics from the German corpora, and 40 from the English one, respectively. We find, as
shown in Fig. 4, that the sub-topics indicate migration from (e.g., Libya, Ukraine) or to (e.g.,
Greece, US) diferent areas, migration-related political events (e.g., Russian invasion of Ukraine,
US army retreating from Afghanistan) or policies (e.g., US immigration and border control laws).
Named Entity Recognition. Events are generally described in news by means of named
entities (NEs) that indicate what, when, and where it happened, or who was involved in it
[42]. We extract NEs from both the textual content, and the metadata of articles, and classify
them into four classes: persons (PER), organizations (ORG), locations (LOC), and miscellaneous
(MISC). We use two XLM-R [38] language models to perform named entity recognition, one
ifne-tuned on the German 14 and the other on the English subset15 of the CoNLL03 dataset
[43]. Concretely, we input the title, abstract, or sentence-segmented news body [44] for text
components, whereas for metadata, we use author names, lists of keywords from the news
outlets, and those used to describe sub-topics. In the former case, we output only the NEs
recognized in these original input. In the latter scenario, the output consists not only of NEs
extracted by the model but also of the original inputs, where no entity was extracted. In the KG
construction step, we include NEs extracted from metadata, as well as other information (e.g.,
frequent keywords, unrecognized authors) to model connections between articles.
Entity Linking. Next, we disambiguate the NEs through named entity linking to Wikidata
[45] using a multilingual entity linking model [46], which is based on a sequence-to-sequence
14https://huggingface.co/xlm-roberta-large-finetuned-conll03-german
15https://huggingface.co/xlm-roberta-large-finetuned-conll03-english
architecture [47] that generates entity names in over 100 languages. As input, we use a text
sequence containing one named entity annotated with special start and end tokens. The input
sequence is either the title, abstract, or sentence-segmented news body, or for metadata, either
the extracted entity from the previous step or the original keyword or author name. For both
input types, the model outputs the linked entity, the corresponding Wikidata QID, the language
of the Wikidata entity label, as well as a score. We select Wikidata as the external knowledge
base for linkage and graph expansion due to (i) being open source, (ii) its wide coverage, and
(iii) its up-to-dateness [48], which is particularly important in the context of news.
Entity Filtering. The previous step can yield both entities not identified in Wikidata although
they exist (e.g., keyword academia corresponds to Wikidata QID Q1211427 ), as well as incorrect
links. Therefore, we filter incorrectly extracted or linked entities. For NEs originating from
textual components, we (1) threshold on the confidence score of the named entity recognition
model; (2) remove entities without a Wikidata page or linked to a Wikimedia disambiguation
page; (3) remove entities whose type does not correspond to the entity type in Wikidata based
on type-specific properties 16; (4) threshold on the score of the named entity linking model; (5)
remove entities whose Wikidata label is in a foreign language and are only observed once in
the dataset. For authors and keywords, we use similar filtering pipelines. Table 2 shows the
statistics for the named entity recognition and linking steps.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Knowledge Graph Construction</title>
      <p>Wikidata [45] is a large knowledge graph comprising a vast amount of information that is often
used to model and integrate common-sense knowledge into recommender systems [49, 25].
Nevertheless, it also contains data which can be irrelevant for domain-specific recommendation
16We used the following properties to detect the type of an entity in Wikidata: PER is instance of human; ORG has
headquarters location, inception, or is founded by; LOC has coordinate location.</p>
      <p>PER_2</p>
      <p>PER_3
keywords keywords
news_topic_4</p>
      <p>Title text Abstract text
headline abstract
news_title_213 news_abstract_213
author about</p>
      <p>hasPart
www.news_outlet.com</p>
      <p>DD.MM.YYYY</p>
      <p>DD.MM.YYYY
neutral
description news_sentiment_positive</p>
      <p>MISC_1</p>
      <p>url
dateModified
datePublished
sentiment
keywords
right
description
news_political_orientation_right</p>
      <p>hasPart
news_213 hasPart
publisher type
ORG_2</p>
      <p>NewsArticle
news_body_213
articleBody
Body text</p>
      <p>isReferencedBy
isReferencedBy</p>
      <p>PER_1
hasActor
news_event_213 hasActor ORG_1
type hasPlace
Event mentions
isReferencedBy</p>
      <p>LOC_1
isReferencedBy</p>
      <p>MISC_1</p>
      <p>Legend
blank
node
resource
literal
relation
class
tasks. Thus, we build a domain-specific KG from the annotated corpora, combining textual
content and metadata from news with Wikidata triples of the NEs and their k-hop neighbors.</p>
      <sec id="sec-4-1">
        <title>4.1. Base graph construction</title>
        <p>NeMigKG is a KG in which the nodes denote news content and real-world entities (e.g., authors,
locations), while the edges indicate diferent relation types between them. Fig. 5 illustrates an
example of an article in NeMigKG.</p>
        <p>Relations. We model relations using several schemas and vocabularies. Specifically, we
represent edges between articles and their textual content or metadata with schema.org17
relationships, and those involving extracted NEs with the Simple Event Model [50]. Furthermore,
we encode information about miscellaneous NEs with the schema:mentions relation, and map
the provenance of NEs to the news text with the dcterms:isReferencedBy.18 Additionally, we
create the relations nemig:sentiment and nemig:political_orientation to represent edges
between news and their sentiment labels, and between media outlets and their political leaning,
respectively. Lastly, we identify pairs of related news19, i.e., articles with identical titles,
overlapping bodies, but diferent provenance, which we map with the schema:isBasedOn relation.
We find that such related news are based on one another, with the most recent published one
representing an update, extension, or longer discussion of the original article.
Nodes. NeMigKG contains two kinds of nodes: (i) literals (e.g., title, dates), and (ii) resources
which encode NEs extracted from diferent parts of an article’s content (e.g. persons, locations) or
metadata (i.e., publisher, author). We further distinguish the latter into (i) disambiguated entities
from Wikidata (denoted as Wikidata resources) and (ii) custom resources created from entities that
17https://schema.org/
18https://www.dublincore.org/specifications/dublin-core/dcmi-terms/terms/isReferencedBy/
19Related news are identified using the overlap ratio between pairs of news’ bodies.
could not be identified or linked to Wikidata, but which still encode meaningful information
and provide knowledge-level connections between news (e.g., authors not in Wikidata, or
frequent keywords). We represent each article in NeMigKG with a unique identifier. For each
title, abstract, body, topic, sentiment, political leaning, or event mapped in the KG, we create
additional nodes, as shown in Fig. 5.</p>
        <p>Enrichment with External Information. We extend NeMigKG with up to 2-hop neighbors
from Wikidata of all Wikidata resources already included in our graph to inject additional
information from external knowledge bases. Firstly, we extract, for all Wikidata resources, all
triples in Wikidata containing another entity as their tail. We ignore literal neighbors. Secondly,
we incorporate the matching triples into the graph and retrieve a new neighbors set for the
added tail entities. We repeat this process iteratively until we enrich NeMigKG with all relevant
triples two hops away from the original Wikidata resources. We post-process the resulting graph
by removing all sink entities, i.e., nodes with a degree smaller than two, which cannot be used
to derive further connections between the news as they are leaf nodes or not linked to Wikidata.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Graph types</title>
        <p>The majority of KG embedding models used in downstream tasks encode only entities and
relations, while ignoring literals. Therefore, we create four variants of NeMigKG for each of
the two datasets, as follows: (1) base - contains literals and entities from the corresponding
news corpus; (2) entities - contains only resource nodes from the base graph, and no literals; (3)
enriched entities - contains only resource nodes and k-hop triples from Wikidata; (4) complete
represents the enriched entities graph with the literals re-added. Table 3 shows the statistics of
the diferent KGs for both languages.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. User Data</title>
      <p>We collect user data in terms of (i) explicit click feedback, (ii) demographics, and (iii) political
information. For each corpus, we conduct one online user study aimed at measuring the political
polarization efects of NRS. The participants are recruited through online-access panels and
selected using a quote procedure to create a representative sample of German, and respectively,
US Internet users aged 18 to 74. Among the German users, 50.2% are male and 49.6% are female,
while 44.6% have an Abitur (passed the secondary school final examinations). 46.2% of the US
users are male and 52.5% are female, and 33.2% have at least a high school degree.</p>
      <sec id="sec-5-1">
        <title>In the study, participants are firstly asked to rate 10 articles as</title>
        <p>
          worthy or not to be read further
(i.e., explicit feedback). We use the IDs of the news deemed worthy of reading to build the user
Click History, which is later used to construct each user’s profile. Afterwards, each participant
is shown a list of six news (impressions), for which a binary rating is recorded. This process
was repeated five times. We use the information gathered in the second step to construct the
Impressions log. Similar to [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], we structure the anonymized user behaviors in the format
[impression ID, user ID, Click History, Impression Log], where the Impression Log contains the IDs
of the news shown to the user and the label indicating whether the user clicked on them.
        </p>
        <p>
          Before the study’s stimulus phase, we collect information regarding the users’ media usage
(MED[
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">1-9</xref>
          ]), political attitudes (POL[
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">1-9</xref>
          ]), and empathy levels (EMP[
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8">1-8</xref>
          ]). After the stimulus, we
ask participants questions regarding their socio-demographic status (gender, age, qualification,
nationality, born in Germany or the US, parents born in Germany or the US, income), emotions
(EMO[
          <xref ref-type="bibr" rid="ref1 ref10 ref11 ref12 ref13 ref14 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">1-14</xref>
          ]), levels of ideological (IPO[
          <xref ref-type="bibr" rid="ref1 ref10 ref11 ref12 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">1-12</xref>
          ]), afective (i.e, identification with the main political
parties), and perceived (PPO[
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6">1-6</xref>
          ]) polarization, of political participation (PPA[
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8">1-8</xref>
          ]), and of
prosocial behavior (PRO[
          <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5">1-5</xref>
          ]). The resulting user datasets contain information about 3,432
German and 3,000 US users, respectively. The German dataset has a sparsity of only 0.02%,
20
meaning that out of the news included in the online user study, less than 1% were never seen
by the participants, in either of the two study phases. In contrast, the English dataset is much
sparser, with 9.27% of the news never interacted with by any of the users. Table 4 shows an
example of questions used during the user data collection process in the US. Please refer to
the project page
        </p>
        <p>or Zenodo [20] for a full set of questions used in both studies, as well as an
extensive explanation of each question, and the corresponding answer scale.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Experiments</title>
      <p>We analyze the performance of diferent news recommenders on NeMig on a range of
recommendation tasks. Afterwards, we perform news trends analysis using the named entities
identified in the corpora.</p>
      <sec id="sec-6-1">
        <title>6.1. Recommendation Models</title>
        <p>We benchmark various recommendation models, which difer in their news and user encoders. 21
The former converts input features (e.g., title, topics, entities) into a news embedding by
contextualizing pretrained word embeddings [51]. The later aggregates the embeddings of the
clicked news into a user-level representation.</p>
        <p>• NRMS [52] embeds news titles with a two-layer encoder consisting of multi-head
selfattention [53], followed by additive attention [54]; a similar architecture is used to learn
user representations.
• NAML [55] uses a sequence of convolutional neural network (CNN) [56] and additive
attention to learn representations from news titles, and abstracts, and additionally
leverages categories through a linear category encoder; its user encoder consists of an additive
attention layer.
• MINS [57] embeds titles and abstracts of news as NRMS [52], and categories through a
linear embedding layer; it learns user representations with a combination of multi-head
self-attention, multi-channel GRU-based recurrent network [58], and additive attention.
• CAUM [59] combines the text encoder used by NRMS [52] with an entity embedder
composed of attention layers in order to learn news representations; it produces
candidateaware user representations with a candidate-aware self-attention network which models
long-range dependencies between clicked news, conditioned on the candidate, combined
with a candidate-aware CNN that captures short-term user interests from adjacent clicks,
also conditioned on the candidate’s content.
• DKN [24] is a knowledge-aware recommender which learns representations from news
titles with a knowledge-aware CNN [56] over aligned embeddings of words and entities,
and of users with a candidate-aware attention network.
• TANR [60] has the same architecture as NAML [55], but does not use a category embedding
layer; additionally, it injects information on topical categories, by jointly optimizing the
recommender for content personalization and topic classification.
• SentiDebias [33], built on the architecture of NRMS [52], addresses the problem of
sentiment debiasing using adversarial learning to reduce the model’s sentiment bias
(originating from the user data) and generate sentiment-diverse recommendations.</p>
        <p>We evaluate the aforementioned NRS not only in terms of the standard content personalization
performance, but also w.r.t. aspect-based diversity and personalization of results, meaning that
we investigate how diverse or faithful the recommendations are to the user’s consumed news in
terms of aspect   . Concretely, following [61], we define aspect-based diversification as the level
of uniformity of an aspect’s distribution among the recommended news. In contrast, we define
aspect-based personalization as the level of homogeneity between a user’s recommendations
and clicked news w.r.t. the distribution of an aspect (e.g., sentiment). We experiment with three
aspects: political leaning, topical categories, and the sentiment of news. Note that we use the
political leaning of a news outlet as proxy for determining the political leaning of an article
originating from that outlet.
21The implementation of the news recommenders is available at https://github.com/andreeaiana/nemig_nrs</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Evaluation Metrics</title>
        <p>We use the common metrics AUC, MRR, nDCG@3, and nDCG@6 to report content
personalization performance. Following [61], we measure aspect-based diversity of recommendations at
position  using the normalized entropy of aspects   ’s distribution in the recommendation list:
where |  | denotes the number of classes of aspect   . We evaluate aspect-based
personalization22 with the generalized Jaccard similarity [62]:
∑
∈ 
() log ()
log(|  |)</p>
        <p>,</p>
        <p>|= 1 | min(ℛ , ℋ )
∑</p>
        <p>
          |= 1 | max(ℛ , ℋ )
∑
,
where   and   represent the probability of a news with class  of aspect   to be contained in
the recommendations list ℛ, and, respectively, in the user history ℋ. All metrics are bounded
to the [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] range.
        </p>
      </sec>
      <sec id="sec-6-3">
        <title>6.3. Experimental Setup</title>
        <p>We use pre-trained 300-dimensional German23 and English24 GloVe embeddings [51] and
100dimensional TransD embeddings [63] pretrained on NeMigKG to initialize the word and entity
embeddings of the NRS. Note that we use the entity embeddings trained on the variant of
NeMigKG enriched with 1-hop Wikidata neighbors. Moreover, we use the extracted sub-topics
as topical categories in NAML, MINS, and TANR, and embed them with 100-dimensional vectors.
Following [64], we sample four negatives per positive example. We train all models for 10
epochs, using a batch size of 4.25 We optimize with the Adam algorithm [65], with the learning
rate set to 1e-5. We set all other model-specific hyperparameters to optimal values reported in
the respective papers. We use 70% randomly sampled user interactions for training the NRS,
10% as validation data, and the remaining 20% as a test set. We repeat each experiment five
times with diferent random seeds and report averages and standard deviations. We normalize
the scores to the [0, 100] range.</p>
      </sec>
      <sec id="sec-6-4">
        <title>6.4. Results and Discussion</title>
        <p>We first discuss the recommendation performance of the aforementioned NRS on NeMig. We
then analyze whether they are prone to implicit biases in terms of three aspects: political leaning,
topical categories, and sentiment orientation. Lastly, we investigate the influence of diferent
input features on the quality of NeMigKG.
22Note that perfect aspect-based personalization would imply identical distributions of aspect   ’s in the
recommendations list and user history.
23https://www.deepset.ai/german-word-embeddings
24https://nlp.stanford.edu/projects/glove/
25We trained each model on a single NVIDIA RTX 2080Ti GPU.
(1)
(2)
Content Personalization. Table 5 summarizes the results on content personalization for
both datasets. We find that, despite the small dataset sizes, the state-of-the-art neural models
achieve high predictive performance. We notice a high similarity w.r.t. content
personalization performance of NAML and TANR, as well as of NRMS and SentiDebias. Both pairs of
models share nearly the same news and user encoder architectures, and difer only in their
optimization objectives. These results indicate that secondary optimization goals (i.e., topic
classification in the case of TANR, or sentiment debiasing in SentiDebias) have little efect on
pure recommendation performance.</p>
        <p>
          Generally, we find that the two knowledge-aware models, DKN and CAUM, outperform all
the other recommenders. This suggests that injecting additional external knowledge into the
model can help detect additional, knowledge-level connections between news, thus resulting
in better news representations. The results are partly at odds with the findings of [ 66], which
observed that while CAUM indeed outperforms other models with candidate-agnostic user
encoders, DKN underperforms. However, [66] conducted experiments on the MIND dataset [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ],
which is magnitudes larger than NeMig. This points to the fact that both knowledge-awareness
and candidate-awareness user modeling might be particularly beneficial in scenarios where
little training data is available. Lastly, we observe that the quality of content personalization
is higher for the German than for the English dataset. We hypothesize that this is due to the
higher sparsity of the English dataset, in which significantly more news have never been seen
by the users.
        </p>
        <p>Aspect-based Diversity and Personalization. Next, we analyze the diversity and
personalization of results, using topical categories (ctg), sentiment (snt), and political leaning (pol)
as aspects. Concretely, we report the performance for content personalization (nDCG@ ),
aspect-based diversity (Dctg for topical categories, Dsnt for sentiment polarization, and Dpol
for political leaning) and aspect-based personalization (PSctg for topical categories, PSsnt for
sentiment polarization, and PSpol for political leaning). We report results for  = 3 .</p>
        <p>Table 6 summarizes the results on content personalization and aspect diversity. As expected,
SentiDebias [33] achieves the highest sentiment diversity on both datasets. NAML [55]
outperforms, by a relatively large margin, the other models in terms of category diversity of results.
These results are somewhat surprising and at odds with those obtained on other datasets
[61], as NAML leverages information on topical categories to customize results to the users’
preferences, and not to address diversity. Furthermore, compared to the first two aspects, we
Content personalization and aspect diversity (in terms of topical categories, sentiments, and political
leaning) performance of diferent NRS. The best results per column are highlighted in bold, the second
best are underlined.</p>
        <p>Content and aspect personalization (in terms of topical categories, sentiments, and political leaning)
performance of diferent NRS. The best results per column are highlighted in bold, the second best are
ifnd that all recommenders perform nearly identical w.r.t. degree of political diversification
of recommendations. This could be due to the fact that none of the models explicitly target
political diversification.</p>
        <p>Aspect-based diversity is tightly correlated with aspect-based personalization [61]. More
specifically, higher levels of diversity come at the cost of personalization, as can be seen in Table
7, which illustrates the results on content and aspect-based personalization. We observe that
categorical and political personalization are much more aligned with content personalization
performance, thqn with the sentiment of news, as discussed also in [61]. We find this to be
intuitive, as users tend to choose which news to read based on their topics/categories of interest,
as well as political preferences, and not on their sentiment polarization.</p>
        <p>An earlier version of the German dataset [67] has been used in [68] and [69] to examine
sentiment and stance recommender bias, and to investigate polarization and filter bubble
creation through online studies, respectively. While in this study we have analyzed political
biases of recommenders only from the perspective of users’ click histories, in the future their
political information can be further explored in conjunction with the models’ predictions to
understand whether any existing biases are reinforced or amplified by the algorithms. Similarly,
NeMig can be used in future works to analyze these trends also in multilingual or cross-lingual
recommendation scenarios, as it comprises information on the same topic in both German and
English.
Knowledge Graph Ablations. The type of data contained in the knowledge graph is
paramount for training high quality knowledge graph embeddings, which are further exploited
by knowledge-aware recommenders to improve the accuracy of predictions. Thus, we analyze
the impact of input features on NeMigKG. Specifically, we pre-train entity embeddings on
diferent versions of NeMigKG using TransD [63], and study the recommendation and aspect-based
diversity performance of DKN using these embeddings. We summarize the results in Table 8 in
terms of content personalization (nDCG@ ), and aspect-based diversity for topical categories
(Dctg), sentiment (Dsnt), and political leaning (Dpol). Adding topical categories, sentiment and
political information in NeMigKG increases the diversity of recommendations w.r.t. these
aspects, while having minor efects on accuracy-based performance. We find that extending
NeMigKG with  -hop neighbors from Wikidata is beneficial only up to  = 1 hops. This indicates
that a certain degree of contextualization of named entities with general knowledge is helpful
when the data size is small. However, injecting too much external knowledge (i.e.,  = 2 ) can
dilute the original information and have detrimental efects on the recommender’s downstream
performance.</p>
      </sec>
      <sec id="sec-6-5">
        <title>6.5. News Trends Analysis</title>
        <p>Lastly, we demonstrate how NeMig can be used to analyze trends and correlations between
entities used in the news discourse and events. Concretely, we analyze the evolution of named
entities overtime and political groups. Figs. 6-7 show the evolution over time and political
orientation of the source outlets of some of the top-20 most frequent entities in our corpora. We
ifnd that generally the most frequently sampled entities are covered predominantly in center
media, which is also the source of the majority of news in both corpora. Notable exceptions in
the German corpus (Fig. 6) are Angela Merkel, appearing equally or more often in right-wing
media, and Die Linke, with a more equal distribution among all political classes, although
leftwing coverage dominates the discourse. The evolution of entities over time reveals correlations
with major events. In terms of organizations, German news report most frequently about the
main political parties, whose evolution coincides with elections on the federal and European
level, or about central political figures in international politics (e.g., the coverage of Recep Tayyip
Erdoğan appears to be potentially correlated with the Turkish ofensive in north-eastern Syria).
In contrast to German media, the representation of the top entities from the English corpus in
the left US media is generally lower.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>We introduced NeMig, a bilingual news collection and knowledge graphs on the topic of
migration, along with user data, including demographic and political information. The news is
collected from mainstream and alternative German and US media outlets spanning a wide
political scale. We extract sub-topics and named entities linked to Wikidata from the news text and
metadata. In contrast to existing datasets, we provide sentiment and political annotations, and
use Wikidata for incorporating background knowledge. We hope that this resource will inspire
future research on analyzing the multidimensional implications of news curation algorithms, in
both monolingual and cross-lingual settings.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This work has been conducted in the ReNewRS project, which is funded by the
BadenWürttemberg Stiftung in the Responsible Artificial Intelligence program. The authors would
also like to thank Julia Wildgans at the University of Mannheim for her legal advice on data
handling.
news content, social context, and spatiotemporal information for studying fake news on
social media, Big data 8 (2020) 171–188.
[17] E. Levi, G. Mor, S. Shenhav, T. Sheafer, Compres: A dataset for narrative structure in news,
arXiv preprint arXiv:2007.04874 (2020).
[18] M. Färber, V. Burkard, A. Jatowt, S. Lim, A multidimensional dataset based on
crowdsourcing for analyzing and detecting news bias, in: Proceedings of the 29th ACM International
Conference on Information &amp; Knowledge Management, 2020, pp. 3007–3014.
[19] S. Lim, A. Jatowt, M. Färber, M. Yoshikawa, Annotating and analyzing biased sentences in
news articles using crowdsourcing, in: Proceedings of the Twelfth Language Resources
and Evaluation Conference, 2020, pp. 1478–1484.
[20] A. Iana, M. Alam, A. Grote, N. Nikolajevic, K. Ludwig, P. Müller, C. Weinhardt, H. Paulheim,
NeMig - A Bilingual News Collection and Knowledge Graph about Migration, 2022. doi:10.
5281/zenodo.7442424.
[21] B. Kille, F. Hopfgartner, T. Brodt, T. Heintz, The plista dataset, in: Proceedings of the 2013
International News Recommender Systems Workshop and Challenge, NRS ’13, Association
for Computing Machinery, New York, NY, USA, 2013, p. 16–23.
[22] G. de Souza Pereira Moreira, F. Ferreira, A. M. da Cunha, News session-based
recommendations using deep neural networks, in: B. Hidasi, A. Karatzoglou, O. S. Shalom, B. Shapira,
D. Tikk, F. Vasile, S. Dieleman (Eds.), Proceedings of the 3rd Workshop on Deep Learning
for Recommender Systems, DLRS@RecSys 2018, Vancouver, BC, Canada, October 6, 2018,
ACM, 2018, pp. 15–23.
[23] G. de Souza Pereira Moreira, D. Jannach, A. M. da Cunha, Contextual hybrid
sessionbased news recommendation with recurrent neural networks, IEEE Access 7 (2019)
169185–169203.
[24] H. Wang, F. Zhang, X. Xie, M. Guo, Dkn: Deep knowledge-aware network for news
recommendation, in: Proceedings of the 2018 world wide web conference, 2018, pp.
1835–1844.
[25] A. Iana, M. Alam, H. Paulheim, A survey on knowledge-aware news recommender systems,</p>
      <p>Semantic Web (2022) 1–62.
[26] P. Vossen, R. Agerri, I. Aldabe, A. Cybulska, M. van Erp, A. Fokkens, E. Laparra, A.-L.</p>
      <p>Minard, A. P. Aprosio, G. Rigau, M. Rospocher, R. Segers, Newsreader: Using knowledge
resources in a cross-lingual reading machine to generate more knowledge from massive
streams of news, Special Issue Knowledge-Based Systems, Elsevier (2016).
[27] M. Rospocher, M. van Erp, P. Vossen, A. Fokkens, I. Aldabe, G. Rigau, A. Soroa, T. Ploeger,
T. Bogaard, Building event-centric knowledge graphs from news, Journal of Web Semantics
(2016).
[28] D. Liu, T. Bai, J. Lian, X. Zhao, G. Sun, J.-R. Wen, X. Xie, News graph: An enhanced
knowledge graph for news recommendation., in: KaRS@ CIKM, 2019, pp. 1–7.
[29] A. L. Opdahl, T. Al-Moslmi, D.-T. Dang-Nguyen, M. Gallofré Ocaña, B. Tessem, C. Veres,
Semantic knowledge graphs for the news: A review, ACM Comput. Surv. (2022). doi:10.
1145/3543508.
[30] G. K. Shahi, D. Nandini, Fakecovid–a multilingual cross-domain fact check news dataset
for covid-19, arXiv preprint arXiv:2006.11343 (2020).
[31] J. Golbeck, M. Mauriello, B. Auxier, K. H. Bhanushali, C. Bonk, M. A. Bouzaghrane, C.
Buntain, R. Chanduka, P. Cheakalos, J. B. Everett, et al., Fake news vs satire: A dataset and
analysis, in: Proceedings of the 10th ACM Conference on Web Science, 2018, pp. 17–21.
[32] Z. Liu, S. Shabani, N. G. Balet, M. Sokhn, Detection of satiric news on social media:
Analysis of the phenomenon with a french dataset, in: 2019 28th International Conference
on Computer Communication and Networks (ICCCN), IEEE, 2019, pp. 1–6.
[33] C. Wu, F. Wu, T. Qi, W.-Q. Zhang, X. Xie, Y. Huang, Removing ai’s sentiment manipulation
of personalized news delivery, Humanities and Social Sciences Communications 9 (2022)
1–9.
[34] J. McCoy, T. Rahman, M. Somer, Polarization and the global crisis of democracy: Common
patterns, dynamics, and pernicious consequences for democratic polities, American
Behavioral Scientist 62 (2018) 16–42.
[35] J. Cho, S. Ahmed, M. Hilbert, B. Liu, J. Luu, Do search algorithms endanger democracy?
an experimental investigation of algorithm efects on political polarization, Journal of
Broadcasting &amp; Electronic Media 64 (2020) 150–172.
[36] R. R. R. Gangula, S. R. Duggenpudi, R. Mamidi, Detecting political bias in news articles
using headline attention, in: Proceedings of the 2019 ACL Workshop BlackboxNLP:
Analyzing and Interpreting Neural Networks for NLP, 2019, pp. 77–84.
[37] K. Lazaridou, R. Krestel, Identifying political bias in news articles, Bulletin of the IEEE</p>
      <p>TCDL 12 (2016).
[38] A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, É. Grave,
M. Ott, L. Zettlemoyer, V. Stoyanov, Unsupervised cross-lingual representation learning at
scale, in: Proceedings of the 58th Annual Meeting of the Association for Computational
Linguistics, 2020, pp. 8440–8451.
[39] F. Barbieri, L. E. Anke, J. Camacho-Collados, Xlm-t: Multilingual language models in
twitter for sentiment analysis and beyond, in: Proceedings of the Thirteenth Language
Resources and Evaluation Conference, 2022, pp. 258–266.
[40] M. Grootendorst, Bertopic: Neural topic modeling with a class-based tf-idf procedure,
arXiv preprint arXiv:2203.05794 (2022).
[41] N. Reimers, I. Gurevych, Making monolingual sentence embeddings multilingual using
knowledge distillation, in: Proceedings of the 2020 Conference on Empirical Methods in
Natural Language Processing, Association for Computational Linguistics, 2020.
[42] L. Li, D.-D. Wang, S.-Z. Zhu, T. Li, Personalized news recommendation: a review and an
experimental investigation, Journal of computer science and technology 26 (2011) 754–766.
[43] E. F. Sang, F. De Meulder, Introduction to the conll-2003 shared task: Language-independent
named entity recognition, arXiv preprint cs/0306050 (2003).
[44] P. Qi, Y. Zhang, Y. Zhang, J. Bolton, C. D. Manning, Stanza: A python natural language
processing toolkit for many human languages, arXiv preprint arXiv:2003.07082 (2020).
[45] D. Vrandečić, M. Krötzsch, Wikidata: a free collaborative knowledgebase, Communications
of the ACM 57 (2014) 78–85.
[46] N. De Cao, L. Wu, K. Popat, M. Artetxe, N. Goyal, M. Plekhanov, L. Zettlemoyer, N.
Cancedda, S. Riedel, F. Petroni, Multilingual autoregressive entity linking, Transactions of the
Association for Computational Linguistics 10 (2022) 274–290.
[47] Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis, L. Zettlemoyer,
Multilingual denoising pre-training for neural machine translation, Transactions of the
Association for Computational Linguistics 8 (2020) 726–742.
[48] N. Heist, S. Hertling, D. Ringler, H. Paulheim, Knowledge graphs on the web-an overview.,
2020.
[49] Q. Guo, F. Zhuang, C. Qin, H. Zhu, X. Xie, H. Xiong, Q. He, A survey on knowledge
graphbased recommender systems, IEEE Transactions on Knowledge and Data Engineering 34
(2020) 3549–3568.
[50] W. R. Van Hage, V. Malaisé, R. Segers, L. Hollink, G. Schreiber, Design and use of the
simple event model (sem), Journal of Web Semantics 9 (2011) 128–136.
[51] J. Pennington, R. Socher, C. D. Manning, Glove: Global vectors for word representation, in:
Proceedings of the 2014 conference on empirical methods in natural language processing
(EMNLP), 2014, pp. 1532–1543.
[52] C. Wu, F. Wu, S. Ge, T. Qi, Y. Huang, X. Xie, Neural news recommendation with multi-head
self-attention, in: Proceedings of the 2019 Conference on Empirical Methods in Natural
Language Processing and the 9th International Joint Conference on Natural Language
Processing (EMNLP-IJCNLP), 2019, pp. 6389–6394.
[53] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I.
Polosukhin, Attention is all you need, Advances in neural information processing systems 30
(2017).
[54] D. Bahdanau, K. H. Cho, Y. Bengio, Neural machine translation by jointly learning to align
and translate, in: 3rd International Conference on Learning Representations, ICLR 2015,
2015.
[55] C. Wu, F. Wu, M. An, J. Huang, Y. Huang, X. Xie, Neural news recommendation with
attentive multi-view learning, in: Proceedings of the 28th International Joint Conference
on Artificial Intelligence, 2019, pp. 3863–3869.
[56] Y. Kim, Convolutional neural networks for sentence classification, in: Proceedings of
the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP),
Association for Computational Linguistics, Doha, Qatar, 2014, pp. 1746–1751.
[57] R. Wang, S. Wang, W. Lu, X. Peng, News recommendation via multi-interest news sequence
modelling, in: ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and
Signal Processing (ICASSP), IEEE, 2022, pp. 7942–7946.
[58] K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, Y. Bengio,
Learning phrase representations using rnn encoder–decoder for statistical machine
translation, in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language
Processing (EMNLP), Association for Computational Linguistics, 2014, p. 1724.
[59] T. Qi, F. Wu, C. Wu, Y. Huang, News recommendation with candidate-aware user
modeling, in: Proceedings of the 45th International ACM SIGIR Conference on Research and
Development in Information Retrieval, 2022, pp. 1917–1921.
[60] C. Wu, F. Wu, M. An, Y. Huang, X. Xie, Neural news recommendation with topic-aware
news representation, in: Proceedings of the 57th Annual meeting of the association for
computational linguistics, 2019, pp. 1154–1159.
[61] A. Iana, G. Glavaš, H. Paulheim, Train once, use flexibly: A modular framework for
multi-aspect neural news recommendation, 2023. arXiv:2307.16089.
[62] V. Bonnici, Kullback-leibler divergence between quantum distributions, and its
upperbound, arXiv preprint arXiv:2008.05932 (2020).
[63] G. Ji, S. He, L. Xu, K. Liu, J. Zhao, Knowledge graph embedding via dynamic mapping
matrix, in: Proceedings of the 53rd annual meeting of the association for computational
linguistics and the 7th international joint conference on natural language processing
(volume 1: Long papers), 2015, pp. 687–696.
[64] C. Wu, F. Wu, Y. Huang, Rethinking infonce: How many negative samples do you need?,
in: L. D. Raedt (Ed.), Proceedings of the Thirty-First International Joint Conference on
Artificial Intelligence, IJCAI-22, International Joint Conferences on Artificial Intelligence
Organization, 2022, pp. 2509–2515.
[65] D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, ICLR (2014).
[66] A. Iana, G. Glavas, H. Paulheim, Simplifying content-based neural news recommendation:
On user modeling and training objectives, in: Proceedings of the 46th International
ACM SIGIR Conference on Research and Development in Information Retrieval, 2023, pp.
2384–2388.
[67] A. Iana, A. Grote, K. Ludwig, M. Alam, P. Müller, C. Weinhardt, H. Sack, H. Paulheim,</p>
      <p>Geneg: German news knowledge graph, 2022. doi:10.5281/zenodo.5913171.
[68] M. Alam, A. Iana, A. Grote, K. Ludwig, P. Müller, H. Paulheim, Towards analyzing the
bias of news recommender systems using sentiment and stance detection, in: Companion
Proceedings of the Web Conference 2022, 2022, pp. 448–457.
[69] K. Ludwig, A. Grote, A. Iana, M. Alam, H. Paulheim, H. Sack, C. Weinhardt, P. Müller,
Divided by the algorithm? the (limited) efects of content-and sentiment-based news
recommendation on afective, ideological, and perceived polarization, Social Science
Computer Review (2023) 08944393221149290.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Huang,</surname>
          </string-name>
          <article-title>SentiRec: Sentiment diversity-aware neural news recommendation, in: Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th</article-title>
          <source>International Joint Conference on Natural Language Processing</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>44</fpage>
          -
          <lpage>53</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .aacl-main.
          <volume>6</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shivaram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Culotta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Shapiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bilgic</surname>
          </string-name>
          ,
          <article-title>The interaction between political typology and filter bubbles in news recommendation algorithms</article-title>
          ,
          <source>in: Proceedings of the Web Conference</source>
          <year>2021</year>
          ,
          <year>2021</year>
          , pp.
          <fpage>3791</fpage>
          -
          <lpage>3801</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>E. Pariser,</surname>
          </string-name>
          <article-title>The filter bubble: What the Internet is hiding from you</article-title>
          ,
          <string-name>
            <surname>Penguin</surname>
            <given-names>UK</given-names>
          </string-name>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Bartels</surname>
          </string-name>
          ,
          <article-title>Messages received: The political impact of media exposure</article-title>
          ,
          <source>American political science review 87</source>
          (
          <year>1993</year>
          )
          <fpage>267</fpage>
          -
          <lpage>285</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Dewenter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Linder</surname>
          </string-name>
          , T. Thomas,
          <article-title>Can media drive the electorate? the impact of media coverage on voting intentions</article-title>
          ,
          <source>European Journal of Political Economy</source>
          <volume>58</volume>
          (
          <year>2019</year>
          )
          <fpage>245</fpage>
          -
          <lpage>261</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Jensen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Naidu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wilse-Samson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gergen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zuckerman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Spirling</surname>
          </string-name>
          ,
          <article-title>Political polarization and the dynamics of political language: Evidence from 130 years of partisan speech [with comments</article-title>
          and discussion],
          <source>Brookings Papers on Economic Activity</source>
          (
          <year>2012</year>
          )
          <fpage>1</fpage>
          -
          <lpage>81</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bakshy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Messing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Adamic</surname>
          </string-name>
          ,
          <article-title>Exposure to ideologically diverse news and opinion on facebook</article-title>
          ,
          <source>Science</source>
          <volume>348</volume>
          (
          <year>2015</year>
          )
          <fpage>1130</fpage>
          -
          <lpage>1132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Boutyline</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Willer</surname>
          </string-name>
          ,
          <article-title>The social structure of political echo chambers: Variation in ideological homophily in online networks</article-title>
          ,
          <source>Political psychology 38</source>
          (
          <year>2017</year>
          )
          <fpage>551</fpage>
          -
          <lpage>569</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ludwig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Does social media use promote political mass polarization? a structured literature review, Questions of Communicative Change and Continuity</article-title>
          .
          <source>In Memory of Wolfram Peiser</source>
          (
          <year>2022</year>
          )
          <fpage>118</fpage>
          -
          <lpage>166</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>J. M. Balkin</surname>
          </string-name>
          ,
          <article-title>Free speech in the algorithmic society: Big data, private governance, and new school speech regulation</article-title>
          ,
          <source>UCDL rev. 51</source>
          (
          <year>2017</year>
          )
          <fpage>1149</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Helberger</surname>
          </string-name>
          ,
          <article-title>On the democratic role of news recommenders</article-title>
          ,
          <source>Digital Journalism</source>
          <volume>7</volume>
          (
          <year>2019</year>
          )
          <fpage>993</fpage>
          -
          <lpage>1012</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Gulla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , P. Liu,
          <string-name>
            <given-names>O.</given-names>
            <surname>Özgöbek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <article-title>The adressa dataset for news recommendation</article-title>
          ,
          <source>WI '17</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2017</year>
          , p.
          <fpage>1042</fpage>
          -
          <lpage>1048</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Zhou, MIND: A large-scale dataset for news recommendation</article-title>
          , in: D.
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chai</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Schluter</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          <string-name>
            <surname>Tetreault</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July</source>
          <volume>5</volume>
          -
          <issue>10</issue>
          ,
          <year>2020</year>
          , Association for Computational Linguistics,
          <year>2020</year>
          , pp.
          <fpage>3597</fpage>
          -
          <lpage>3606</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>W. Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          , “
          <article-title>liar, liar pants on fire”: A new benchmark dataset for fake news detection, in: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics</article-title>
          (Volume
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <year>2017</year>
          , pp.
          <fpage>422</fpage>
          -
          <lpage>426</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>B.</given-names>
            <surname>Horne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khedr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Adali</surname>
          </string-name>
          ,
          <article-title>Sampling the news producers: A large news and feature data set for the study of the complex media landscape</article-title>
          ,
          <source>in: Proceedings of the International AAAI Conference on Web and Social Media</source>
          , volume
          <volume>12</volume>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K.</given-names>
            <surname>Shu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mahudeswaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          , H. Liu,
          <article-title>Fakenewsnet: A data repository with</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>