<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Multimedia Satellite Task at MediaEval 2018</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Emergency Response for Flooding Events</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Benjamin Bischke</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>German Research Center for Artificial Intelligence (DFKI)</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Radboud University</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>TU Kaiserslautern</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>VU University Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>This paper provides a description of the MediaEval 2018 Multimedia Satellite Task. The primary goal of the task is to extract and fuse content associated with events represent in Satellite Imagery and Social Media. Establishing a link from Satellite Imagery to Social Multimedia can yield to a comprehensive event representation which is vital for numerous applications. Focusing on natural disaster events, the main objective of the task is to leverage the combined event representation within the context of emergency response and environmental monitoring. In particular, our task focuses on flooding events and consists of two subtasks. The first Image Classification from Social Media subtask requires participants to retrieve images from Social Media that show a direct evidence for road passability during flooding events. The second task Flood Detection from Satellite Images aims to extract potentially ofloded road sections from satellite images. The task seeks to go beyond state-of-the-art lfooding map generation by focusing on information about road passability and the accessibility of urban infrastructure. Such information shows a clear potential to complement information from social images with satellite imagery for emergency management.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Recent advances in Earth observation and the access to satellite
imagery at a large scale are opening up a new exciting area for
the applications of remotely sensed data. A proper analysis of this
data source has potential to change how agriculture, urbanization
and environmental monitoring will be done in the future. Hand in
hand with this development, the Multimedia Satellite Task at
MediaEval 2018 addresses natural disaster and environmental monitoring,
allowing to improve situational awareness for such events.</p>
      <p>One challenge when solely relying on remotely sensed data is the
sparsity problem of satellite imagery over time, which often results
in a poor event representation. The larger goal of this task is
therefore to combine the satellite view with the ground-level perspective
represented by images in social media streams in order to obtain a
comprehensive picture of disaster events. Such a multi-modal event
representation from social media and satellite imagery is of vital
importance to achieve situational awareness and to provide support
in emergency response, e.g., helping to coordinate rescuer eforts
in large scale disasters. It is also important for studying disasters
after they have happened, and support planning that will prevent
or mitigate the impact of future disasters.</p>
      <p>
        The Multimedia Satellite Task 2018 continues to focus on
flooding events as in last year’s Task 2017 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], since, among high-impact
natural disasters, flooding events represent according to the United
Nations Ofice for the Coordination of Humanitarian Afairs 1 the
most common type of disaster worldwide. This year the task will
look at passability, namely whether or not it is possible to travel
through a flooded region. Rapid information about road passability
and the accessibility of the urban infrastructure is a critical aspect
in emergency response. Additionally, passability of roads is also an
area in which the information in social images has clear potential
to complement the information in satellite images.
2
The main objective of this year’s task is to quantify the impact of
lfooding events on infrastructure. The task involves two subtasks:
      </p>
      <sec id="sec-1-1">
        <title>Flood Classification from Social Multimedia.</title>
        <p>The goal of the first subtask is to retrieve all images from social
media that provide direct evidence for passability of roads by
conventional means (no boats, of-the-road vehicles, monster trucks,
Hummer, Landrover, farm equipment). The objective is to design a
system/algorithm/method that (in principle) given any collection
of flood related multimedia images and their metadata (e.g., Twitter,
Flickr, YFCC100M) is able to identify those images that (1) provide
evidence for road passability and (2) discriminate between images
showing passable vs. non passable roads. In our context, road
passability is related to the water level visible in the image and the
surrounding context. Participants are allowed to submit 5 runs:
• Required run 1: using visual data only
• General run 2, 3, 4, 5: everything automated allowed,
including using data from external sources (e.g. Twitter, Flickr)</p>
      </sec>
      <sec id="sec-1-2">
        <title>Flood Detection from Satellite Imagery.</title>
        <p>Participants receive high resolution satellite imagery for areas
in Houston, that have been partially flooded during the hurricane
event Harvey in 2017 from DigitalGlobe2. The goal of this subtask
is to move forward the state-of-the-art of flood map generation
by concentrating on road passability. In this regard, the challenge
of this subtask is to identify sections of roads that are potentially
blocked due to high water levels. Participants receive in addition to
the very high resolution satellite patches, two pre-defined points
on the road network depicted in the image. The task is to decide
whether or not it is possible for a vehicle to drive on the road
between the two points using the shortest path without passing
through potentially flooded sections. Fusion of satellite and social
1http://reliefweb.int/disasters
2https://www.digitalglobe.com/opendata</p>
        <sec id="sec-1-2-1">
          <title>Metadata</title>
        </sec>
        <sec id="sec-1-2-2">
          <title>Visual Features</title>
          <p>image_id, image_url, date_taken, date_uploaded, user_nsid, user_nickname, title, text, hashtags,
capture_device, latitude, longitude
AutoColorCorrelogram, EdgeHistogram, Color and Edge Directivity Descriptor (CEDD),
ColorLayout, Fuzzy Color and Texture Histogram (FCTH), Joint Composite Descriptor (JCD), Gabor,
ScalableColor, Tamura
multimedia information is encouraged. Participants are allowed to
submit 5 runs:
• Required run 1, 2: using the provided satellite data only
• General run 3, 4, 5: everything automated allowed,
including using data from external sources (e.g. Open Street Map,</p>
          <p>Elevation Maps, Other Satellite Images, Social Media)</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3 DATA</title>
      <sec id="sec-2-1">
        <title>Flood Classification from Social Multimedia.</title>
        <p>
          The dataset of the first subtask consists of 7,387 Tweet-Ids
(devset) and 3.683 Tweet-Ids (test-set). All tweets with the tags flooding,
lfood and floods in the text and an accompanying image have been
collected during the three big hurricane events in 2017 (named by
Harvey, Irma and Maria) from Twitter. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] In line with previous
research [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], we also observed a large number of (near)-duplicated
images in the collected dataset. Therefore, two pre-processing steps
have been applied in order to de-duplicated such content. In a first
step, perceptual hashing using the pHash function [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] was applied
to remove all duplicated images based on the same hash-value.
In the second step, near duplicates have been excluded based on
the similarity of the deep feature representation of the last fully
connected layer of an ImageNet [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] pre-trained ResNet101 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. As
similarity measure, the cosine distance was used and all image
features with a small distance under an empirically determined
threshold (t=0.1) were grouped to one cluster of image duplicates.
        </p>
        <p>The ground truth labels of the dataset consists of two classes: (1)
one class label for the evidence of road passability for each tweet-Id
with respect to the embedded image (0=no evidence/ 1=evidence).
Those images that are labeled as showing evidence, have a second
class label (2) for the actual road passability (0=not passable/
1=passable). The images accompanying the text of the tweets were labeled
by human annotators in a crowd-sourcing setup on the platform
Figure Eight3.</p>
        <p>Participants were asked multiple questions about the image
content with respect to the road passability and corresponding evidence
for passability. The examples for road passability were available to
the annotators in the interface during the entire process. The
annotation process was not time restricted. The scores were collected
from three annotators and aggregated according to the majority
voting.</p>
        <p>For each image, classical visual feature descriptors are provided
to participants. These features were extracted with the open-source
LIRE library4 using default parameter settings. An overview of the
provided features is given in Table 1. The dataset is separated with
a ratio of 70/30 into the following two sets:
3https://www.figure-eight.com
4LIRE, http://www.lire-project.net/
• Development-Set contains 7,387 tweets, along with
visual and metadata features as well as two class labels for
evidence and road passability
• Test-Set contains 3,683 images and features</p>
      </sec>
      <sec id="sec-2-2">
        <title>Flood Detection from Satellite Imagery.</title>
        <p>The dataset for the second remote sensing subtask consists of
1,664 satellite image patches that were extracted from DigitalGlobe’s
WorldView satellite. The imagery has a ground-sample distance
(GSD) of about 0.5 meters and was collected from the Houston area
during the hurricane event Harvey in 2017. The image patches have
the spatial resolution of 512 x 512 pixels and show flooded as well
as unflooded areas of Houston.</p>
        <p>The satellite imagery comes with additional binary annotations
for the road passability between two given point locations on the
road network. The dataset is separated into the following split:
• Development-Set contains 1,438 image patches. For each
image patch we provide two points on the road network
and an annotation for the passability (1= passable, 0 = non
passable).</p>
        <p>• Test-Set consists of 226 satellite image patches.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 EVALUATION</title>
      <sec id="sec-3-1">
        <title>Flood Classification from Social Multimedia.</title>
        <p>The oficial metric for evaluating the correctness of classified
images from social multimedia is the macro averaged F1-Score. In our
problem definition, the metric has to consider the following three
classes (C1) images with no evidence on passability, (C2) images
with evidence and passable roads as well as (C3) images with
evidence and non passable roads. Since this definition extends the
binary classification to a multi-label problem, the average of two
F1-Scores for class C2 and C3 is computed.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Flood Detection from Satellite Imagery.</title>
        <p>In order to assess the performance of the system for the
classiifcation of satellite patches that depict potentially blocked road
connections between two given points, the metric F1-Score is used.
This metric computes the harmonic mean between precision and
recall for the non passable road class.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>ACKNOWLEDGMENTS</title>
      <p>We would like to thank Martha Larson for the very valuable
feedback and support during the setup of this task. Additionally, we
would like to thank DigitalGlobe for providing us with high-resolution
satellite images for this task. This work was partially funded by the
BMBF Project DeFuseNN (01IW17002).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bischke</surname>
          </string-name>
          , Damian Borth,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Schulze</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Dengel</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Contextual enrichment of remote-sensed events with social media streams</article-title>
          .
          <source>In Proceedings of the 2016 ACM on Multimedia Conference. ACM</source>
          ,
          <volume>1077</volume>
          -
          <fpage>1081</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bischke</surname>
          </string-name>
          , Patrick Helber, Christian Schulze, Srinivasan Venkat, Andreas Dengel, and
          <string-name>
            <given-names>Damian</given-names>
            <surname>Borth</surname>
          </string-name>
          .
          <source>The Multimedia Satellite Task at MediaEval</source>
          <year>2017</year>
          :
          <article-title>Emergency Response for Flooding Events</article-title>
          .
          <source>In Proc. of the MediaEval 2017 Workshop (Sept</source>
          .
          <fpage>13</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2017</year>
          ). Dublin, Ireland.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bischke</surname>
          </string-name>
          , Patrick Helber,
          <string-name>
            <given-names>Zhengyu</given-names>
            <surname>Zhao</surname>
          </string-name>
          , Jens de Bruijn, and
          <string-name>
            <given-names>Damian</given-names>
            <surname>Borth</surname>
          </string-name>
          .
          <source>The Multimedia Satellite Task at MediaEval</source>
          <year>2018</year>
          :
          <article-title>Emergency Response for Flooding Events</article-title>
          .
          <source>In Proc. of the MediaEval 2018</source>
          Workshop (Oct.
          <fpage>29</fpage>
          -
          <lpage>31</lpage>
          ,
          <year>2018</year>
          ). Sophia-Antipolis, France.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Mengjuan</given-names>
            <surname>Fei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jing</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Honghai</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Visual tracking based on improved foreground detection and perceptual hashing</article-title>
          .
          <source>Neurocomputing</source>
          <volume>152</volume>
          (
          <year>2015</year>
          ),
          <fpage>413</fpage>
          -
          <lpage>428</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Kaiming</given-names>
            <surname>He</surname>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          .
          <volume>770</volume>
          -
          <fpage>778</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Olga</given-names>
            <surname>Russakovsky</surname>
          </string-name>
          , Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Bernstein</surname>
          </string-name>
          , and others.
          <source>2015</source>
          .
          <article-title>Imagenet large scale visual recognition challenge</article-title>
          .
          <source>International Journal of Computer Vision</source>
          <volume>115</volume>
          ,
          <issue>3</issue>
          (
          <year>2015</year>
          ),
          <fpage>211</fpage>
          -
          <lpage>252</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>