<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Visual Sentiment Analysis: A Natural Disaster Use-case Task at MediaEval 2021</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Syed Zohaib Hassan</string-name>
          <email>syed@simula.no</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kashif Ahmad</string-name>
          <email>kahmad@hbku.edu.qa</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael A. Riegler</string-name>
          <email>michael@simula.no</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Steven Hicks</string-name>
          <email>steven@simula.no</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Conci</string-name>
          <email>nicola.conci@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pål Halvorsen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ala Al-Fuqaha</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>SimulaMet</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Norway</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Information</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Computing Technology (ICT) Division</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>College of Science</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Engineering</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hamad Bin Khalifa</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University</institution>
          ,
          <addr-line>Doha 34110</addr-line>
          ,
          <country country="QA">Qatar</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>The Visual Sentiment Analysis task is being ofered for the first time at MediaEval. The main purpose of the task is to predict the emotional response to images of natural disasters shared on social media. Disaster-related images are generally complex and often evoke an emotional response, making them an ideal use case of visual sentiment analysis. We believe being able to perform meaningful analysis of natural disaster-related data could be of great societal importance, and a joint efort in this regard can open several interesting directions for future research. The task is composed of three sub-tasks, each aiming to explore a diferent aspect of the challenge. In this paper, we provide a detailed overview of the task, the general motivation of the task, and an overview of the dataset and the metrics to be used for the evaluation of the proposed solutions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        An enormous amount of multimedia content is generated and
shared over the internet daily, especially after the introduction
of several social media platforms, such as Twitter, Instagram, and
Facebook. Among other types of content, visual content is used
extensively to deliver a specific message to the viewer, implying
the common phrase “a picture is worth a thousand words” [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Various emotions can be expressed in visual content, where extracting
and analyzing them could be of great value in diferent
application domains, such as education, entertainment, advertisement, and
journalism. However, it is not clear which part of an image evoke
certain emotions and, more importantly, how the underlying
sentiments can be derived from a scene by an automatic algorithm and
how the sentiments can be expressed. This opens an interesting
line of research to interpret emotions and sentiments perceived by
users viewing visual content.
      </p>
      <p>
        Disaster-related images can be challenging to interpret,
especially when analyzing them using current visual analysis techniques.
These images generally contain important details in both the
background and foreground, making no single part the main point of
interest [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Developing a visual sentimental analysis pipeline for
such a use-case is dificult but could greatly serve our society.
Example use cases include having news agencies fact check and provide
their audience with relevant information in case of adverse events.
It can also be used to assess the damage caused by a disaster and
use the relevant imagery to create awareness and raise funds for
humanitarian activities. We note that the resultant solutions are
not intended to manipulate public emotions rather are concerned
with estimating disaster severity and supporting disaster relief. For
instance, humanitarian organizations can benefit from these
solutions to reach a wider audience by communicating visual content
that best demonstrates the evidence of a certain event [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In order to facilitate future work in the domain, a large-scale
dataset is collected, annotated, and made publicly available. For the
annotation of the dataset, a crowd-sourcing activity with a large
number of participants has been conducted.
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        The literature reports several interesting works aiming to minimize
and mitigate the loss of lives and infrastructure damage. Diferent
types of solutions rely on ground sensors and satellite imagery
for efective disaster handling [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. With the emergence of social
media, more human-centric approaches have arisen for efective
response management to natural disasters. There are also some
eforts on sentiment analysis of disaster-related text. For instance,
IN-SPIRE [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and PUlSE [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] are two frameworks dealing with the
sentimental analysis of text data shared on social media.There has
been a great deal of work done on textual sentimental analysis of
disaster-related textual content. However, the concept of extracting
sentiments from the visual content is quite new. There are also
some eforts on visual sentiment analysis of natural disaster-related
multimedia content. For instance, in [
        <xref ref-type="bibr" rid="ref5 ref9">5, 9</xref>
        ] a deep visual sentiment
analysis is introduced along with a benchmark dataset visual
sentiment analysis of natural disaster analysis. We believe the task will
help in developing a task force to further explore this interesting
aspect of natural disaster-related visual content.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>DATASET DETAILS</title>
      <p>
        As a first step, in creating this benchmark dataset [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], we crawled
images related to natural disasters from the social media platforms,
such as Twitter, Google API, and Flicker. During this process, special
consideration was put on the fact that the crawled images meet
the licensing policy of free usage and sharing. A keyword search
was the basic technique used to crawl these images, for example,
searching by type of disaster, such as floods , wildfires , earthquake,
landslides, and storms, or some specific recent natural disasters that
happened in diferent parts of the world, such as floods in Pakistan
etc.
      </p>
      <p>After data collection, the next step was data annotation, which
was crowd-sourced using Microworkers. The challenging part for
crowd-sourcing was designing the task and selecting proper
keywords to represent diferent categories of the annotated data. The
three keywords namely Positive, Negative, and Neutral are the most
widely used sentiment labels in the literature. However
considering the applications and the complexity of sentimental analysis
of disaster-related images, it is quite complicated to interpret and
represent the sentiment and emotion in these three categories.
We rather need a more specific and larger set of sentiment
categories/labels to interpret the emotions associated with natural
disasters. There is an interesting recent work in psychology on
emotions and sentiments representation, where 27 diferent
emotion categories are reported. Based on this study, we set up four
diferent types of annotations, which are expected to better
represent the sentiments and emotions associated with natural disasters
by covering diferent aspects of natural disasters. These sets of
labels are described below.</p>
      <p>(1) Positive, negative, and neutral.
(2) Relax/calm, normal and simulated/excited.
(3) Joy, sadness, fear, disgust, anger, surprise, and neutral.
(4) Anger, anxiety, craving, empathetic pain, fear, horror, joy,
relief, sadness, and surprise.</p>
      <p>As a first step towards visual sentiment analysis of natural
disaster multimedia content, the scope of the task is limited to images
from diferent types of natural disasters. Overall, 4, 003 images were
selected for the crowd-sourcing study. To ensure quality and
consistency, each image was annotated by five diferent persons. The
study was disseminated through multiple channels. A total of 10, 010
responses were received during the study, and 2, 338 participants
participated from 98 diferent countries.
4</p>
    </sec>
    <sec id="sec-4">
      <title>TASK DESCRIPTION</title>
      <p>
        It is the first time we are hosting this task in MediaEval. The task is
closely related to previous tasks namely ”The 2017-2019 Multimedia
Satellite Task: Emergency Response for Flooding Events” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and
”The 2020 Flood-related Multimedia Task” [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However, the goals,
challenges, images, and the types of natural disasters covered in
the dataset are diferent from the previous tasks.
      </p>
      <p>Images from the dataset provided for the task are filtered to
check whether there is a majority consensus among the annotations
provided by the crowd-sourcing participants. The resulted dataset
consists of 3, 631 images. The development set is composed of 2, 432
images, and the test set contains 1, 199 images. The task is divided
into three subtasks, which are described below.
4.1</p>
    </sec>
    <sec id="sec-5">
      <title>Subtask 1</title>
      <p>This is a multi-class single label classification task, where the images
are arranged in three diferent classes, namely positive, negative,
and neutral. There is a strong imbalance towards the negative class,
given the nature of the topic.
4.2</p>
    </sec>
    <sec id="sec-6">
      <title>Subtask 2</title>
      <p>This is a multi-class multi-label image classification task, where
the participants are provided with multi-labeled images. The
multilabel classification strategy, which assigns multiple labels to an
image, better suits our visual sentiment classification problem and
is intended to show the correlation of diferent sentiments. In this
task, seven classes, namely joy, sadness, fear, disgust, anger, surprise,
and neutral, are covered.
4.3</p>
    </sec>
    <sec id="sec-7">
      <title>Subtask 3</title>
      <p>This task is also multi-class multi-label, however, a wider range of
sentiment classes are covered. Going deeper in the sentiment
hierarchy, the complexity of the task increases. The sentiment categories
covered in this task include anger, anxiety, craving, empathetic pain,
fear, horror, joy, relief, sadness, and surprise.
5</p>
    </sec>
    <sec id="sec-8">
      <title>EVALUATION</title>
      <p>The proposed solutions will be evaluated using a weighted 1
score. It is a more preferable metric when dealing with imbalanced
datasets as it is more sensitive to data distribution. It also provides
a good balance between precision and recall as it is a harmonic
mean of these two metrics. The weighted 1 score can be calculated
using the following four equations.
(1)
(2)
(3)
(4)
 =</p>
      <p>True Positives
(True Positives + False Positives )
 =</p>
      <p>True Positives
(True Positives + False Negatives )
1 Score =
 ℎ 1 =
2 ∗  ∗ 
( +  )
Í=1 (1 Score ∗  )
Í=1 ( )</p>
      <p>Equation 1, 2, and 3 shows the calculation of precision, recall
and 1 score for each class, respectively. Equation 4 shows the
calculation weighted 1 score where  is the number of instances
of a particular class.
6</p>
    </sec>
    <sec id="sec-9">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>Though the literature reports some initial eforts on the topic,
several important aspects need to be addressed yet. We believe that
extracting and representing sentiments from disaster-related visual
content will benefit several stakeholders, such as news agencies,
Government and non-government, and other humanitarian
organizations. From a research point of view, visual sentiment analysis
of natural disaster-related visual content is not limited to object
recognition. Instead, it requires a more complex identification
analysis and establishing a connection among salient objects, scenes,
expressions, and color schemes. We hope the task will result in a
joint efort towards this important topic, and it also will open new
research directions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Ahsan</given-names>
            <surname>Adeel</surname>
          </string-name>
          , Mandar Gogate, Saadullah Farooq, Cosimo Ieracitano, Kia Dashtipour, Hadi Larijani, and
          <string-name>
            <given-names>Amir</given-names>
            <surname>Hussain</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A NewEficient Alert Model for Disaster Management</article-title>
          .
          <source>Geological Disaster Monitoring Based on Sensor Networks</source>
          (
          <year>2018</year>
          ),
          <fpage>57</fpage>
          -
          <lpage>66</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Stelios</given-names>
            <surname>Andreadis</surname>
          </string-name>
          , Ilias Gialampoukidis, Anastasios Karakostas, Stefanos Vrochidis, Ioannis Kompatsiaris, Roberto Fiorin, Daniele Norbiato, and
          <string-name>
            <given-names>Michele</given-names>
            <surname>Ferri</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>The flood-related multimedia task at mediaeval 2020</article-title>
          .
          <source>In Proceedings of the MediaEval 2020 Workshop</source>
          , Online.
          <fpage>14</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bischke</surname>
          </string-name>
          , Patrick Helber, Christian Schulze, Venkat Srinivasan, Andreas Dengel, and
          <string-name>
            <given-names>Damian</given-names>
            <surname>Borth</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The Multimedia Satellite Task at MediaEval</article-title>
          <year>2017</year>
          .. In MediaEval.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Michelle</given-names>
            <surname>Gregory</surname>
          </string-name>
          , Deborah Payne,
          <string-name>
            <surname>Dave</surname>
            <given-names>McColgin</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Nick</given-names>
            <surname>Cramer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V</given-names>
            <surname>DouglasLove</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Visual Analysis of Weblog Content</article-title>
          .
          <source>INTERNATIONAL CONFERENCE ON WEB AND SOCIAL MEDIA</source>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Syed</given-names>
            <surname>Zohaib</surname>
          </string-name>
          <string-name>
            <surname>Hassan</surname>
          </string-name>
          , Kashif Ahmad, Steven Hicks, Pål Halvorsen, Ala Al-Fuqaha,
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Conci</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Visual sentiment analysis from disaster images in social media</article-title>
          . arXiv preprint arXiv:
          <year>2009</year>
          .
          <volume>03051</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Imran</given-names>
            <surname>Khan</surname>
          </string-name>
          , Kashif Ahmad, Namra Gul, Talhat Khan, Nasir Ahmad, and
          <string-name>
            <surname>Ala</surname>
          </string-name>
          Al-Fuqaha.
          <year>2021</year>
          .
          <article-title>Explainable Event Recognition</article-title>
          .
          <source>arXiv preprint arXiv:2110.00755</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Naina</given-names>
            <surname>Said</surname>
          </string-name>
          , Kashif Ahmad, Michael Riegler, Konstantin Pogorelov, Laiq Hassan, Nasir Ahmad, and
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Conci</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Natural disasters detection in social media and satellite imagery: a survey</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          <volume>78</volume>
          ,
          <issue>22</issue>
          (
          <year>2019</year>
          ),
          <fpage>31267</fpage>
          -
          <lpage>31302</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Marc</surname>
            <given-names>A</given-names>
          </string-name>
          <string-name>
            <surname>Smith and Andrew T Fiore</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Visualization componentsfor persistent conversations. Proceedings of the SIGCHI conferenceon Human factors in computing systems (</article-title>
          <year>2001</year>
          ),
          <fpage>136</fpage>
          -
          <lpage>143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Conci Syed Zohaib</surname>
          </string-name>
          , Kashif Ahmad and
          <string-name>
            <given-names>Ala</given-names>
            <surname>Al-Fuqaha</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Sentiment Analysis from Images of Natural Disasters</article-title>
          .
          <source>International Conference on Image Analysis and Processing</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>