<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CNN and GAN Based Satellite and Social Media Data Fusion for Disaster Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kashif Ahmad</string-name>
          <email>kashif.ahmad@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pogorelov Konstantin</string-name>
          <email>konstantin@simula.no</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Riegler</string-name>
          <email>michael@simula.no</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Conci</string-name>
          <email>nicola.conci@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pal Holversen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DISI-University of Trento</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Simula Research Labs Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper presents the method proposed by team UTAOS for the Mediaeval 2017 challenge on Multi-media and Satellite. In the first task, we mainly rely on diferent Convolutional Neural Network (CNN) models combined with two diferent late fusion methods. We also utilize the additional information available in the form of meta-data. The average and mean over precision at diferent cut-ofs for our best runs are 84.94% and 95.11%, respectively. For challenge two, we utilize a Generative Adversarial Network (GAN). The mean Intersection-over-Union (IoU) for our best run is 0.8315. Linking social media information to remote sensed data holds large possibilities for society and research [1-3]. The Multimedia and Satellite task in Mediaeval 2017 [4] aims to integrate information from both sources, sensed data and social media, to provide a better overview of a disaster. This paper provides a detailed description of the methods developed by the UTOS team for the Mediaeval 2017 Multimedia Satellite Task. The challenge consists of two sub tasks, (i) Disaster Image Retrieval from Social Media (DIRSM) and (ii) Flood Detection in Satellite Images (FDSI).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>2.1</p>
    </sec>
    <sec id="sec-2">
      <title>PROPOSED APPROACH</title>
    </sec>
    <sec id="sec-3">
      <title>Methodology for DIRSM Task</title>
      <p>
        To tackle challenge (i), we rely on Convolutional Neural Network
(CNN) features. In detail, we first extract CNN features for seven
diferent models from state-of-the-art architectures pre-trained on
the ImageNet [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and places datasets [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. These models include
AlexNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] (pre-trained on both ImageNet and places datasets),
GoogleNet [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] (pre-trained on ImageNet ), VGGNet 19 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
(pretrained on both ImagNet and places datasets) and diferent
conifgurations of ResNet [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] with 50, 101 and 152 layers. For feature
extraction from Alexnet and VGGNet19 we use the Cafe toolbox 1
while in the case of GoogleNet and Resnet we exploited Vlfeat
Matcovnet2.
      </p>
      <p>All in all, we extract eight feature vectors through four
diferent network architectures from the same image. AlexNet and
VGGNet16 provide a feature vector of size 4096 while GoogleNet and
Resnet provide feature vectors of 1024 and 2048, respectively.
Subsequently, the extracted features are fed into ensembles of Support
Vector Machines (SVMs), which provide classification scores in</p>
      <sec id="sec-3-1">
        <title>1http://cafe.berkeleyvision.org/ 2http://www.vlfeat.org/matconvnet/</title>
        <p>2.2</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Methodology for FDSI Task</title>
      <p>For the challenge (ii), we started from the visual analysis of the
provided development set. We observed that it is not possible to
use any already existing open-source framework due to the nature
of the provided satellite data. Furthermore, we observed that the
used four-channel 16-bit TIFF file format is too specific and cannot
be correctly processed and even viewed by existing libraries.</p>
      <p>To perform the visual analysis we developed a conversion code
which provide a conversion from geo-TIFF to a pair of images:
RGB and infrared (IR). For the RGB images we used the
per-threechannels normalization which fits all the R, G and B pixel values of
the input geo-image into standard 0-255 RGB region. Normalization
coeficients are the same for all three channels to achieve real color
balance even in cases of low variations in one of the components.
The normalization of the IR component is performed separately.
rдbmin = min(min ri , min дi , min bi )</p>
      <p>i ∈R i ∈G i ∈B
rдbmax = max(max ri , max дi , max bi )
i ∈R i ∈G i ∈B
irmin = min irk , irmax = max irk</p>
      <p>k ∈I R k ∈I R
∀i ∈ {R |G |B} {r |д|b }i∗ = ({r |д|b }i − rдbmin ) ∗ 255</p>
      <p>rдbmax − rдbmin
∀k ∈ I R iri∗ = (irk − irmin ) ∗ 255</p>
      <p>irmax − irmin</p>
      <p>
        Moreover, we performed a human-expert-driven visual
analysis of the images and found them all to be non-contrast, blurry
and color-range-limited. From our previous experience [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] we
decided to use a generative adversarial network (GAN). GANs3 are
a class of artificial intelligence algorithms used in unsupervised
machine learning, implemented by a system of two neural networks
contesting with each other in a zero-sum game framework.
      </p>
      <p>
        As the basis for our method we selected a neural network
architecture used for retinal vessel segmentation in fundoscopic images
with generative adversarial networks (V-GAN) 4. The V-GAN
architecture is designed [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for processing of retinal images that have
comparable visual properties and provides the required output with
one-class image segmentation masks.
      </p>
      <p>V-GAN is implemented in Python on top of Keras with
Tensorflow GPU-enabled back-end. We have modified the network
architecture by changing the top-layers configuration in order to
support four-channel floating-point geo-image-compatible input.
The final generator network output layer used for creation of
probabilistic output segmentation image was extended by the simple
threshold activation layer to generate the binary segmentation map.</p>
      <p>First, we have performed experiments with the development set
only and found that the modified V-GAN is able to perform the
segmentation of the provided satellite images, but the estimated
performance metrics were below the expected level. Additional
visual analysis of the converted RGB and IR images showed that
sometimes IR component of the sourced geo-images was irrelevant
to the flooding areas that probably caused our GAN to bias during
training process and prevent it from the correct flooding areas
properties extraction. Thus, we have decided to exclude IR component
from the model input and process only the RGB components of the
converted normalized geo-images. This resulted in the significant
performance improvement and correct segmentation most of the
developments set flooding areas except for the some images taken
in not-common lighting and cloudy conditions.
3
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>RESULTS AND ANALYSIS</title>
    </sec>
    <sec id="sec-6">
      <title>Runs Description in DIRSM Task</title>
      <p>For DIRSM, we submitted five diferent runs. Table 1 provides the
oficial results of our methods in terms of average precision at
cutof 480 and mean over precision at diferent cutof ( 50, 100, 250, 480).
Run 1 and run 4 are mainly based on visual information extracted
with seven diferent CNN models and jointly utilized in PSO and
IOWA based fusions, respectively. As it can be seen in Table 1, the
PSO based fusion method outperforms IOWA with a significant
gain of 3.79% and 5.34%. On the other hand, run 2 is based on
metadata achieving the worst results among the all runs. Similarly, run
3 and run 5 represents two diferent variations of our method used
for combining meta-data and visual information. Run 3 is based on
IOWA while run 5 represents our PSO based fusion of meta-data and</p>
      <sec id="sec-6-1">
        <title>3http://en.wikipedia.org/wiki/Generative_adversarial_networks 4https://bitbucket.org/woalsdnd/v-gan</title>
        <p>visual information. Again, PSO based fusion performs better. One
of the main limitations of IOWA based fusion is its mechanism of
assigning more weight to a more confident model. In this particular
case, we noticed that our classifier trained on meta-data provides
more confident decisions with high probabilities causing significant
reduction in the performance. This can also be concluded from the
results on run 2 where the meta-data obtain worst results. The
degradation in the performance due to the inclusion of meta-data
shows that the additional information available are not much useful.
This paper provides a detailed description of the methods proposed
by UTAOS for the Mediaeval 2017 challenge on Multimedia and
Satellite. During the experimental evaluation of sub-task 1 (DIRSM),
we noticed that visual information seems more useful compared to
meta-data for the retrieval of disaster images. For sub-task 2 (FDSI),
we rely on a Generative Adversarial Network where better results
are obtained in 3 and 4. Based on the experiments conducted in this
work we believe that a proper fusion of social media information
and satellite data can provide a better story of a natural disaster.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Kashif</given-names>
            <surname>Ahmad</surname>
          </string-name>
          , Michael Riegler, Konstantin Pogorelov, Nicola Conci, Pål Halvorsen, and Francesco De Natale.
          <year>2017</year>
          .
          <article-title>JORD: A System for Collecting Information and Monitoring Natural Disasters by Linking Social Media with Satellite Imagery</article-title>
          .
          <source>In Proceedings of the 15th International Workshop on Content-Based Multimedia Indexing. ACM</source>
          ,
          <volume>12</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Kashif</given-names>
            <surname>Ahmad</surname>
          </string-name>
          , Michael Riegler, Ans Riaz, Nicola Conci,
          <string-name>
            <surname>Duc-Tien Dang-Nguyen</surname>
            , and
            <given-names>Pål</given-names>
          </string-name>
          <string-name>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The JORD System: Linking Sky and Social Multimedia Data to Natural Disasters</article-title>
          .
          <source>In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval. ACM</source>
          ,
          <volume>461</volume>
          -
          <fpage>465</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bischke</surname>
          </string-name>
          , Damian Borth,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Schulze</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Dengel</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Contextual enrichment of remote-sensed events with social media streams</article-title>
          .
          <source>In Proceedings of the 2016 ACM on Multimedia Conference. ACM</source>
          ,
          <volume>1077</volume>
          -
          <fpage>1081</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bischke</surname>
          </string-name>
          , Patrick Helber, Christian Schulze, Srinivasan Venkat, Andreas Dengel, and
          <string-name>
            <given-names>Damian</given-names>
            <surname>Borth</surname>
          </string-name>
          .
          <source>The Multimedia Satellite Task at MediaEval</source>
          <year>2017</year>
          :
          <article-title>Emergence Response for Flooding Events</article-title>
          .
          <source>In Proc. of the MediaEval 2017 Workshop (Sept</source>
          .
          <fpage>13</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2017</year>
          ). Dublin, Ireland.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Jia</given-names>
            <surname>Deng</surname>
          </string-name>
          , Wei Dong, Richard Socher,
          <string-name>
            <surname>Li-Jia</surname>
            <given-names>Li</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kai</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>Li</surname>
          </string-name>
          Fei-Fei.
          <year>2009</year>
          .
          <article-title>Imagenet: A large-scale hierarchical image database</article-title>
          .
          <source>In Computer Vision and Pattern Recognition</source>
          ,
          <year>2009</year>
          .
          <article-title>CVPR 2009</article-title>
          .
          <article-title>IEEE Conference on</article-title>
          . IEEE,
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Hall</surname>
          </string-name>
          , Eibe Frank, Geofrey Holmes, Bernhard Pfahringer,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Reutemann</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ian H Witten</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The WEKA data mining software: an update</article-title>
          .
          <source>ACM SIGKDD explorations newsletter 11</source>
          ,
          <issue>1</issue>
          (
          <year>2009</year>
          ),
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Kaiming</given-names>
            <surname>He</surname>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          .
          <volume>770</volume>
          -
          <fpage>778</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Alex</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , Ilya Sutskever, and
          <string-name>
            <given-names>Geofrey E</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Imagenet classification with deep convolutional neural networks</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>1097</volume>
          -
          <fpage>1105</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Michael Riegler, Sigrun Losada Eskeland, Thomas de Lange, Dag Johansen, Carsten Griwodz, Peter Thelin Schmidt, and
          <string-name>
            <given-names>Pål</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Eficient disease detection in gastrointestinal videos-global features versus neural networks</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          (
          <year>2017</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Karen</given-names>
            <surname>Simonyan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          .
          <source>arXiv preprint arXiv:1409.1556</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Jaemin</surname>
            <given-names>Son</given-names>
          </string-name>
          , Sang Jun Park, and
          <string-name>
            <surname>Kyu-Hwan Jung</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Retinal Vessel Segmentation in Fundoscopic Images with Generative Adversarial Networks</article-title>
          .
          <source>arXiv preprint arXiv:1706.09318</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Christian</surname>
            <given-names>Szegedy</given-names>
          </string-name>
          , Wei Liu, Yangqing Jia,
          <string-name>
            <given-names>Pierre</given-names>
            <surname>Sermanet</surname>
          </string-name>
          , Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Rabinovich</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Going deeper with convolutions</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition. 1-9.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Ronald R Yager and Dimitar P Filev</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Induced ordered weighted averaging operators</article-title>
          .
          <source>IEEE Transactions on Systems, Man, and Cybernetics</source>
          ,
          <string-name>
            <surname>Part</surname>
            <given-names>B</given-names>
          </string-name>
          (
          <year>Cybernetics</year>
          )
          <volume>29</volume>
          ,
          <issue>2</issue>
          (
          <year>1999</year>
          ),
          <fpage>141</fpage>
          -
          <lpage>150</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Bolei</surname>
            <given-names>Zhou</given-names>
          </string-name>
          , Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Learning deep features for scene recognition using places database</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>487</volume>
          -
          <fpage>495</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>