<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DisasterMM: Multimedia Analysis of Disaster-Related Social Media Data Task at MediaEval 2022</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stelios Andreadis</string-name>
          <email>andreadisst@iti.gr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aristeidis Bozas</string-name>
          <email>arbozas@iti.gr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ilias Gialampoukidis</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anastasia Moumtzidou</string-name>
          <email>moumtzid@iti.gr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Fiorin</string-name>
          <email>roberto.fiorin@distrettoalpiorientali.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesca Lombardo</string-name>
          <email>francesca.lombardo@distrettoalpiorientali.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thanassis Mavropoulos</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniele Norbiato</string-name>
          <email>daniele.norbiato@distrettoalpiorientali.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefanos Vrochidis</string-name>
          <email>stefanos@iti.gr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Ferri</string-name>
          <email>michele.ferri@distrettoalpiorientali.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ioannis Kompatsiaris</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Eastern Alps River Basin District</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Information Technologies Institute - Centre of Research and Technology Hellas</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the “DisasterMM: Multimedia Analysis of Disaster-Related Social Media Data” Task at MediaEval 2022. Social media data have been widely used in disaster management and have been proven valuable to all phases of a disaster: from early warning to nowcasting and from response to recovery. Nevertheless, the large streams of user-generated content can be very noisy and usually lack geoinformation, which is an essential attribute. The goal of DisasterMM is to tackle these challenges with two subtasks: RCTP (Relevance Classification of Twitter Posts), which asks participants to build a classifier that will predict whether a tweet is relevant or not to a disaster, in particular floods, and LETT (Location Extraction from Twitter Texts), which calls for a text analysis technique that detects which words inside a Twitter message refer to a location.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Flooding is the most common disaster occurring worldwide [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and their impact is expected to
grow due to climate change [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Besides loss of lives and property damage, floodwaters pose
immediate dangers to human health [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and long-term efects resulting from displacement
and worsened living conditions. With the prominent rise of the social media, the ability to
get real-time crowdsourced data, including text, photos and videos, becomes integrated into
daily activities. Social media data are already explored by first responders and civil protection
authorities as an alternative source of information [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7">4, 5, 6, 7</xref>
        ], complementary to traditional
means such as telephone, in order to raise the situational awareness and support their operations.
In parallel, the scientific society has been proposing AI and Machine Learning solutions that
improve the quality of the incoming social media data [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
      </p>
      <p>However, the utilization of large and continuous streams from social media platforms comes
with two significant limitations. First, the vast amount of user-generated posts carries lots of
noise, with messages that may contain flood-related words, but are actually irrelevant to floods
(e.g., words used in a metaphorical way). Second, the majority of posts are not geotagged (i.e.
not associated with a geographic position) or their geoinformation is questionable [10].</p>
      <p>The automatic prediction of a post’s relevance could reduce the social media noise and thus
assist the interested parties in receiving only useful information, without spending time on
ifltering out unrelated messages. In addition, recognizing the locations that are mentioned inside
the post’s text could enhance the post with geographic information, which would allow the
automatic positioning of a potential incident. By receiving solely high-quality and geotagged
social data, disaster management practitioners will be able to manage their resources more
eficiently, which could even lead to saving more human lives.</p>
      <p>The above has motivated the organisation of the “DisasterMM: Multimedia Analysis of
Disaster-Related Social Media Data” Task1 at MediaEval 2022, which swaps the focus back to
lfoods from water quality (2021’s WaterMM [ 11]), following the Multimedia Satellite Task
(20172019) [12, 13, 14] and the Flood-related Multimedia Task (2020) [15]. The goal of DisasterMM is
to tackle two individual challenges: to identify posts that are related to floods using textual,
visual and metadata information and to detect mentioned locations inside a post text. For both
subtasks, the datasets are in Italian language, so as to encourage the research community to
move beyond the English language for text analysis.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task Description</title>
      <p>The DisasterMM task concerns the multimedia analysis of social media data, specifically posts
from the popular platform of Twitter, that relate to the disaster of floods. The participants of
this task are provided with a set of Tweet IDs in order to download textual as well as visual
information and other metadata of tweets that have been selected with keyword-based search
that involved words/phrases about flood. DisasterMM includes two subtasks:
• Relevance Classification of Twitter Posts (RCTP) : The objective of this subtask is to build a
binary classification system that will be able to distinguish whether a tweet is relevant or
not to flooding incidents. An example of a relevant tweet is shown in Fig. 1, while an
irrelevant tweet in Fig. 2
• Location Extraction from Twitter Texts (LETT): In this subtask, participants are asked to
develop a named-entity recognition model in order to identify which words (or sequence
of words) inside a tweet’s text refer to locations. An example is shown in Table 1, with
the tags being explained in the next section.
1https://multimediaeval.github.io/editions/2022/tasks/disastermm/</p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset Description</title>
      <p>DisasterMM involves two datasets that have been retrieved from Twitter by searching for
lfood-related keywords (e.g., “alluvione”, “allagamento”, “esondazione” – all translated as flood).
The complete list of keywords can be found along with the datasets for RCTP2 and LETT3 in
the task’s repository, which will be made public after MediaEval 2022.</p>
      <p>The dataset for RCTP is a set of 6,672 tweets collected between May 25, 2020 and June 12,
2020. The ground truth of the RCTP dataset refers to the relevance of a tweet, i.e., relevant (1)
or not relevant (0), and is provided to participants in the form of key-value pairs of Tweet ID
and relevance label. The dataset is separated randomly4 into two sets: the development-set that
contains 5,337 posts and the test-set with 1,335 posts.</p>
      <p>The dataset for LETT consists of 4,992 tweets collected between March 25, 2017 and August
1, 2018. The ground truth of the LETT dataset is a string per tweet that contains one of the
following labels for each word of its text: “B-LOC” for the first word of a sequence that refers to
a location or a single-word location, “I-LOC” for the subsequent word of a sequence that refers
to a location, and “O” for any non-location word (as mentioned before, an example can be seen
in Table 1). Furthermore, in order to have a fairer evaluation, the word tokens deriving from the
processing of the Twitter text (e.g., removal of multiple spaces, new lines, etc.) are also shared.
The development-set of this dataset includes 3,993 posts and the test-set 999 posts (again split
randomly).</p>
      <p>Both datasets (RCTP/LETT) have been manually annotated by native speakers that are
employed by the Eastern Alps River Basin District, which is responsible for the hydrogeological
defense and flood risk management in the Eastern Alps partition of North-East Italy.
Furthermore, only the ground truth for the development-sets is released, since the ground truth for
the test-sets is used in the evaluation stage and will be available only after the completion of
the challenges. Finally, it should be noted that only the IDs of the tweets are distributed to the
participants, in order to be fully compliant with the Twitter Developer Agreement &amp; Policy5.
2https://github.com/multimediaeval/2022-DisasterMM/blob/main/DisasterMM2022_RCTP_keywords.json
3https://github.com/multimediaeval/2022-DisasterMM/blob/main/DisasterMM2022_LETT_keywords.json
4With scikit-learn’s train_test_split(): https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.
train_test_split.html
5https://developer.twitter.com/en/developer-terms/agreement-and-policy</p>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation</title>
      <p>In RCTP, F1-score6 is selected as the oficial evaluation metric for the binary classification of
tweets as relevant (1) or not relevant (0) on the test set.</p>
      <p>In LETT, F1-score will be used too, not in sentence level, but in word level. To further explain,
if a given label for a word matches the label of the annotator for this particular word, then it is
considered as true (true positive if “B-LOC”/“I-LOC”, true negative if “O”). Two scores will be
measured per each run: the exact F1-score, where labels have to fully match, and the partial
F1-score, where either “B-LOC” or “I-LOC” can be considered as true as long as the annotator’s
label concerns location.</p>
      <p>Participants are also greatly encouraged to carry out a failure analysis of their results in order
to gain insight in the mistakes that their models make.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work has been supported by the EU’s Horizon 2020 research and innovation programme
under grant agreements H2020-833435 INGENIOUS, H2020-101004152 CALLISTO, H2020-832876
aqua3S.
[10] A. Kruspe, M. Häberle, E. J. Hofmann, S. Rode-Hasinger, K. Abdulahhad, X. X. Zhu, Changes in
twitter geolocations: Insights and suggestions for future usage, arXiv preprint arXiv:2108.12251
(2021).
[11] S. Andreadis, I. Gialampoukidis, A. Bozas, A. Moumtzidou, R. Fiorin, D. Norbiato, M. Ferri,
S. Vrochidis, I. Kompatsiaris, WaterMM: Water Quality in Social Multimedia Task at MediaEval
2021, in: Proceedings of the MediaEval 2021 Workshop, Online, 2021.
[12] B. Bischke, P. Helber, C. Schulze, V. Srinivasan, A. Dengel, D. Borth, The Multimedia Satellite Task
at MediaEval 2017., in: MediaEval, 2017.
[13] B. Benjamin, H. Patrick, Z. Zhengyu, B. Damian, et al., The Multimedia Satellite Task at MediaEval
2018: Emergency response for flooding events (2018).
[14] B. Bischke, P. Helber, S. Brugman, E. Basar, Z. Zhao, M. Larson, K. Pogorelov, The Multimedia
Satellite Task at MediaEval 2019: Estimation of Flood Severity, in: Proc. of the MediaEval 2019
Workshop, Sophia Antipolis, France, 2019, p. 1621–1627.
[15] S. Andreadis, I. Gialampoukidis, A. Karakostas, S. Vrochidis, I. Kompatsiaris, R. Fiorin, D. Norbiato,
M. Ferri, The Flood-related Multimedia Task at MediaEval 2020, in: Proceedings of the MediaEval
2020 Workshop, Online, 2020, pp. 14–15.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>W.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>FitzGerald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Clark</surname>
          </string-name>
          , X.-Y. Hou, Health impacts of floods,
          <source>Prehospital and disaster medicine 25</source>
          (
          <year>2010</year>
          )
          <fpage>265</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N. W.</given-names>
            <surname>Arnell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. N.</given-names>
            <surname>Gosling</surname>
          </string-name>
          ,
          <article-title>The impacts of climate change on river flood risk at the global scale</article-title>
          ,
          <source>Climatic Change</source>
          <volume>134</volume>
          (
          <year>2016</year>
          )
          <fpage>387</fpage>
          -
          <lpage>401</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Paterson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Harris</surname>
          </string-name>
          ,
          <article-title>Health risks of flood disasters</article-title>
          ,
          <source>Clinical Infectious Diseases</source>
          <volume>67</volume>
          (
          <year>2018</year>
          )
          <fpage>1450</fpage>
          -
          <lpage>1454</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bensi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. B.</given-names>
            <surname>Baecher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Social media crowdsourcing for rapid damage assessment following a sudden-onset natural hazard event</article-title>
          ,
          <source>International Journal of Information Management</source>
          <volume>60</volume>
          (
          <year>2021</year>
          )
          <fpage>102378</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Goncalves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Morreale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bonafide</surname>
          </string-name>
          ,
          <article-title>Crowdsourcing for public safety</article-title>
          ,
          <source>in: 2014 IEEE International Systems Conference Proceedings, IEEE</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>50</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Alexander</surname>
          </string-name>
          ,
          <article-title>Social media in disaster risk reduction and crisis management</article-title>
          ,
          <source>Science and engineering ethics 20</source>
          (
          <year>2014</year>
          )
          <fpage>717</fpage>
          -
          <lpage>733</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. L.</given-names>
            <surname>Austin</surname>
          </string-name>
          ,
          <article-title>Examining the role of social media in efective crisis management: The efects of crisis origin, information form, and source on publics' crisis responses</article-title>
          ,
          <source>Communication research 41</source>
          (
          <year>2014</year>
          )
          <fpage>74</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moumtzidou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Andreadis</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gialampoukidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karakostas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vrochidis</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kompatsiaris</surname>
          </string-name>
          ,
          <article-title>Flood relevance estimation from visual and textual content in social media streams</article-title>
          ,
          <source>in: Companion Proceedings of the The Web Conference</source>
          <year>2018</year>
          ,
          <year>2018</year>
          , pp.
          <fpage>1621</fpage>
          -
          <lpage>1627</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Imran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ofli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Caragea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          ,
          <article-title>Using AI and social media multimodal content for disaster response and management: Opportunities, challenges</article-title>
          , and future directions,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>