<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benjamin Kille</string-name>
          <email>benjamin.u.kille@ntnu.no</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Lommatzsch</string-name>
          <email>andreas.lommatzsch@tu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Özlem Özgöbek</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mehdi Elahi</string-name>
          <email>mehdi.elahi@uib.no</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Duc-Tien Dang-Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Berlin Institute of Technology</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Kristiania University College</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Norwegian University of Science and Technology</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Bergen</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Most news outlets ofer a multi-modal user experience. Besides texts, readers encounter images, audio, video, and interactive elements. News Images strives to understand better how images afect news consumption. Participants gain access to a large scale data set of news articles and images. The task consists of two subtasks. Participants can engage in both or one of them. In the first subtask, participants must predict which images publishers paired given news articles with. In the second subtask, participants must estimate the chance that users will pay attention to pairs of articles and images. This paper describes the settings in detail and draws connections to existing research.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The news landscape features a multimodal mix of content.
Frequently, images accompany the text to draw attention. Research
concerning multimedia and recommender systems usually assumes
a simple relationship between images and text. For instance,
research on image captioning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] assumes the caption to quite literally
describe the image’s scenery. However, research on the connection
between text and images of news articles indicates a more
complicated relationship [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This is where the News Images task in
MediaEval 2021 comes into play. The task investigates this
relationship to understand its implications for journalism and news
personalisation.
      </p>
      <p>The task comprises two subtasks, both of which participants
can tackle using text-based or image-based features. For the first
task, the link between a set of articles and images has been
removed. Participants must re-establish which images the publisher
had assigned to articles. For the second task, participants must
estimate the chance that users paid attention to articles. Thereby,
we are trying to understand if images increase the users’ attention
to news articles. Ultimately, we seek to gain further insight about
the relationship of text and images and the reactions of users to
recommendations consisting of headlines and images. Particularly,
the task aims to surpass conventional work in the area of image
concept detection. In other words, we hope to capture aspects of
images that exceed the literally depicted content such as quality,
style, and framing.</p>
    </sec>
    <sec id="sec-2">
      <title>BACKGROUND AND RELATED WORK</title>
      <p>The Multimedia Evaluation Benchmark (MediaEval) examines the
intersection of multimodality and recommendation for the fourth
time in 2021. In 2018, the NewsREEL Multimedia1 task ofered data
from several publishers. In 2019, a subtask of the MultimediaRecsys2
features similar data. In 2020, the NewsImages3 task provided data
covering three months of news. In 2021, we have extended the
coverage of previously released data set by another month.</p>
      <p>
        News outlets have introduced personalisation in the form of
recommender systems [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. News recommender systems help users
ifnding relevant content. Still, most recommender systems rely on
information extracted from ratings and texts; the consideration of
image data could help to further improve recommendations. On
the other hand, personalisation has introduced some issues. The
emergence of ‘fake news’ has raised some red flags [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. ‘Fake
news’ often put data or images in an misleading context. Thus
a fine-grained analysis of images and text helps to get a better
understanding of this phenomenon.
      </p>
      <p>
        Research has recognised the importance of topics related to news
personalisation. As a result, more and more venues for discussion
have been created [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Besides, recent years have seen an upward
trend concerning research on multimodal recommender systems.
For instance, Truong and Lauw [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] investigate how to leverage
multimodal user feedback, Salah et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] prepare a framework
for multimodal recommender systems and Oramas et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
examine the use of multimodal data for music recommendation. For a
comprehensive review, we refer to [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>News recommender system research strives to learn about the
positive and negative efects of personalisation. The News Images
task supports the research toward multimodality. We want to learn
more about how images afect news readers’ experience.
3</p>
    </sec>
    <sec id="sec-3">
      <title>TASK DESCRIPTION</title>
      <p>News Images explores the interplay of text and imagery for news
consumption. The task defines subtasks. Participants can take part
in either or both of the subtasks.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Task 1: Image-Text Re-Matching</title>
      <p>Publishers manually equip news articles with images. The content
curators will take images from the event if available. Otherwise,
they can use images from designated databases such as stock images.
Consequently, we encounter pairs of articles and images. For this
1https://www.multimediaeval.org/mediaeval2018/newsreelmm/
2https://www.multimediaeval.org/medaieval2019/mmrecsys/
3https://multimediaeval.github.io/editions/2020/tasks/newsimages/
subtask, we have removed the links between images and articles.
The subtask asks participants to re-establish the correct links.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Task 2: News Click Prediction</title>
      <p>News outlets monitor users’ engagement with their channels. Their
webservers record interactions between readers and news articles.
We expect images to play a role in readers’ complicated decision
making on whether or not to read an article. For this task, we have
concealed the interaction statistics of the test set. Participants must
predict the articles which attracted most user engagement.</p>
      <p>Both subtasks focus on news consumption. We assess
submissions both in terms of quantitative performance and qualitative
value. More qualitative aspects concern the increased
understanding of images’ efect on news consumption.
4</p>
    </sec>
    <sec id="sec-6">
      <title>DATASET</title>
      <p>The task’s underlying data is derived from four months worth of
webserver log files of a German news publisher. The data contains
information related to articles, images, and interactions with users.
Each article and images has a reference number assigned. Articles’
metadata includes the URL, title, and a text snippet of at most 256
characters. Participants have to download the images as we lack the
necessary copyright to distribute them. Users have engaged with
articles in three ways: accessing the article, seeing recommendations,
and clicking on the recommendations. While the system delivers
the recommendations, users explicitly choose to read articles and
click on recommendations.</p>
      <p>The data set comprises five batches. The first three batches
constitute the training data for both subtasks. The training data contains
the links between articles and images as well as the interaction
statistics. The fourth batch contains the test data for subtask
ImageText Re-Matching. Therein, the link between articles and images
has been removed. The fifth batch contains the evaluation data for
subtask News Click Prediction. Therein, the interaction statistics
have been removed.</p>
      <p>Table 1 illustrates the data set. The data have been split
chronologically to guarantee meaningful results. All batches contain
between 1900 and 2700 articles and images. Downloading a batch
can take around 45 minutes with a standard broadband internet
connection.</p>
    </sec>
    <sec id="sec-7">
      <title>EVALUATION</title>
      <p>Participants must re-match articles and images and predict how
much attention articles attract. For each subtask, participants can
submit up to five runs.
5.1</p>
    </sec>
    <sec id="sec-8">
      <title>Task 1: Image Task Re-Matching</title>
      <p>The link between articles and images has been removed for the
evaluation set. Participants deliver a ranked list of at most 100 candidate
images for each article. The evaluation set contains 1915 articles and
images. Participants must provide a file with 101 columns separated
by a tab character. The first column contains the article reference.
The second column contains the most likely matched image. The
third column contains the second best match and so on. For each
article, we compute the precision at a set of cut-of points. For
instance, we can check whether the first items contain the actually
linked image. Averaging over all articles, we obtain the quantitative
evaluation metric: precision at rank  for  ∈ {1, 5, 10, 20, 50, 100}.
5.2</p>
    </sec>
    <sec id="sec-9">
      <title>Task 2: News Click Prediction</title>
      <p>The training partitions reveal how frequently articles have been
read by users. In the evaluation partition, the reading statistics
remain hidden. Participants must estimate the statistics. We are less
interested in precise estimates. Instead, we want to know whether the
estimator can discern the most interesting articles. Consequently,
participants provide a file with article references populating the first
column. The second column shows the estimated reading statistics.
The estimates have to be numeric such that we can define a ranking.
The evaluation takes the ranking and computes the precision at 
for  ∈ {1, 5, 10, 20, 50, 100}. The one hundred articles with most
accesses constitute the ground truth.
5.3</p>
    </sec>
    <sec id="sec-10">
      <title>Run Description</title>
      <p>Participants inform about their ideas and discuss the evaluation
results in working notes. The working notes highlight their
reasoning, qualitative findings, and critical reflections about what can be
deduced from the quantitative results. Participants may compare
the results from diferent runs and analyze the finding with respect
to result quality, computational complexity, and the used resources.
The discussion of the results should take into account the specific
properties of the dataset and explain how the finding can be used
in related scenarios. Ultimately, the participants ought to describe
what they have learned and how their insights can help to move
the research forward.
6</p>
    </sec>
    <sec id="sec-11">
      <title>CONCLUSION</title>
      <p>How users interact with digital content on a larger scale remains
hard to understand. Multimodal news presentation becomes more
important. Publishers aim to retain their readership by informing
and entertaining. Consequently, understanding users’ preferences
toward content can give them competitive advantages. News Images
wants to shed light on the relation between images and text in the
news domain. Better understanding this relationship can help to
prevent harm to the news eco-system, for instance in the form of
increased polarisation and diminished trust in information.</p>
    </sec>
    <sec id="sec-12">
      <title>ACKNOWLEDGMENTS</title>
      <p>We would like to thank plista for kindly providing the real world
data. Further, we thank Martha Larson for her support.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Francesco</given-names>
            <surname>Corsini</surname>
          </string-name>
          and
          <string-name>
            <given-names>Martha</given-names>
            <surname>Larson</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>CLEF NewsREEL 2016: Image based Recommendation</article-title>
          .
          <source>In Working Notes of the 7th International Conference of the CLEF Initiative</source>
          , Evora, Portugal. CEUR Workshop Proceedings.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Deldjoo</surname>
          </string-name>
          , Markus Schedl, Paolo Cremonesi, and
          <string-name>
            <given-names>Gabriella</given-names>
            <surname>Pasi</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Recommender systems leveraging multimedia content</article-title>
          .
          <source>ACM Computing Surveys (CSUR) 53</source>
          ,
          <issue>5</issue>
          (
          <year>2020</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Jia</given-names>
            <surname>Deng</surname>
          </string-name>
          , Wei Dong, Richard Socher,
          <string-name>
            <surname>Li-Jia</surname>
            <given-names>Li</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kai</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>Li</surname>
          </string-name>
          Fei-Fei.
          <year>2009</year>
          .
          <article-title>Imagenet: A large-scale hierarchical image database</article-title>
          .
          <source>In 2009 IEEE Conference on Computer Vision and Pattern Recognition. IEEE</source>
          ,
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Mouzhi</given-names>
            <surname>Ge</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Persia</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Survey of Multimedia Recommender Systems: Challenges and Opportunities</article-title>
          .
          <source>International Journal of Semantic Computing</source>
          <volume>11</volume>
          ,
          <issue>03</issue>
          (
          <year>2017</year>
          ),
          <fpage>411</fpage>
          -
          <lpage>428</lpage>
          . https://doi.org/10.1142/ S1793351X17500039
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>MD. Zakir</given-names>
            <surname>Hossain</surname>
          </string-name>
          , Ferdous Sohel, Mohd Fairuz Shiratuddin, and
          <string-name>
            <given-names>Hamid</given-names>
            <surname>Laga</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A Comprehensive Survey of Deep Learning for Image Captioning</article-title>
          .
          <source>ACM Comput. Surv</source>
          .
          <volume>51</volume>
          ,
          <issue>6</issue>
          ,
          <string-name>
            <surname>Article 118</surname>
          </string-name>
          (
          <issue>Feb</issue>
          .
          <year>2019</year>
          ). https://doi.org/10.1145/3295748
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Mozhgan</given-names>
            <surname>Karimi</surname>
          </string-name>
          , Dietmar Jannach, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Jugovac</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>News recommender systems-Survey and roads ahead</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>54</volume>
          ,
          <issue>6</issue>
          (
          <year>2018</year>
          ),
          <fpage>1203</fpage>
          -
          <lpage>1227</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Lommatzsch</surname>
          </string-name>
          , Benjamin Kille, Frank Hopfgartner, Martha Larson, Torben Brodt, Jonas Seiler, and
          <string-name>
            <given-names>Özlem</given-names>
            <surname>Özgobek</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>CLEF 2017 NewsREEL Overview: A Stream-based Recommender Task for Evaluation and Education</article-title>
          .
          <source>In 8th International Conference of the CLEF Association: Experimental IR Meets Multilinguality</source>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          (CLEF
          <year>2017</year>
          ). Springer.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Nelleke</given-names>
            <surname>Oostdijk</surname>
          </string-name>
          , Hans van Halteren, Erkan Bas, ar, and
          <string-name>
            <given-names>Martha</given-names>
            <surname>Larson</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>The Connection between the Text and Images of News Articles: New Insights for Multimedia Analysis</article-title>
          .
          <source>In Proceedings of The 12th Language Resources and Evaluation Conference</source>
          .
          <volume>4343</volume>
          -
          <fpage>4351</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Sergio</given-names>
            <surname>Oramas</surname>
          </string-name>
          , Oriol Nieto, Mohamed Sordo, and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Serra</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A deep multimodal approach for cold-start music recommendation</article-title>
          .
          <source>In Proceedings of the 2nd workshop on deep learning for recommender systems</source>
          .
          <volume>32</volume>
          -
          <fpage>37</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Özlem</surname>
            <given-names>Özgöbek</given-names>
          </string-name>
          , Benjamin Kille, Jon Atle Gulla, and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Lommatzsch</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>The 7th international workshop on news recommendation and analytics (INRA 2019)</article-title>
          .
          <source>In Proceedings of the 13th ACM Conference on Recommender Systems</source>
          .
          <volume>558</volume>
          -
          <fpage>559</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Aghiles</surname>
            <given-names>Salah</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quoc-Tuan Truong</surname>
          </string-name>
          , and Hady W Lauw.
          <year>2020</year>
          .
          <article-title>Cornac: A Comparative Framework for Multimodal Recommender Systems</article-title>
          .
          <source>J. Mach. Learn. Res</source>
          .
          <volume>21</volume>
          (
          <year>2020</year>
          ),
          <fpage>95</fpage>
          -
          <lpage>1</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Quoc-Tuan Truong</surname>
            and
            <given-names>Hady</given-names>
          </string-name>
          <string-name>
            <surname>Lauw</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Multimodal review generation for recommender systems</article-title>
          .
          <source>In The World Wide Web Conference</source>
          .
          <year>1864</year>
          -
          <fpage>1874</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Xinyi</given-names>
            <surname>Zhou</surname>
          </string-name>
          and
          <string-name>
            <given-names>Reza</given-names>
            <surname>Zafarani</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>A Survey of Fake News</article-title>
          .
          <source>Comput. Surveys</source>
          <volume>53</volume>
          ,
          <issue>5</issue>
          (Sep
          <year>2020</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>40</lpage>
          . https://doi.org/10.1145/3395046
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>