<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Convolutional Neural Networks for Disaster Images Retrieval</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>DCSE</institution>
          ,
          <addr-line>UET Peshawar</addr-line>
          ,
          <country country="PK">Pakistan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DISI-University of Trento</institution>
          ,
          <addr-line>Trento</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Sheharyar Ahmad</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper presents the method proposed by MRLDCSE team for the disaster image retrieval task in Mediaeval 2017 challenge on Multimedia and Satellite. In the proposed work, for visual information, we rely on Convolutional Neural Networks (CNN) features extracted with two diferent models pre-trained on ImageNet and places datasets. Moreover, a late fusion technique is employed to jointly utilize visual and the additional information available in the form of meta-data for the retrieval of disaster images from social media. The average precision for our three diferent runs with visual information only, meta-data and combination of meta-data and visual information are 95.73%, 18.23% and 92.55%, respectively.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        In recent years, social media emerged as an important source of
information and communication; especially in disaster situations
where usually news agencies are unable to provide information in
time due to unavailability of reporters in the area. For instance, the
authors in [
        <xref ref-type="bibr" rid="ref14 ref9">9, 14</xref>
        ] have proved social network as an efective medium
of mass communication in emergency situations. A rather recent
trend is to infer events from information shared through social
media [
        <xref ref-type="bibr" rid="ref15 ref2">2, 15</xref>
        ]. The analysis of recent literature reveals that social
media platforms, particularly Twitter and Flickr have been heavily
exploited for inferring information about diferent types of events,
such as social and sports events. In this regards, an interesting
application is to collect and analyze information about natural
disasters available on social network. To this aim, a number of
interesting solutions have been proposed to efectively utilize social
media for information collection and analyzing the impact of a
natural disaster [
        <xref ref-type="bibr" rid="ref12 ref4">4, 12</xref>
        ].
      </p>
      <p>
        On the other hand, satellite images have also been proved very
efective to explore and monitor the surface of the earth and its
environment [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In this regards, Joyce et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] provides a
detailed review of diferent techniques developed to eficiently utilize
remote-sensed data for the monitoring and assessment of damage
due to natural hazards and disasters.
      </p>
      <p>
        A rather recent trend is to combine remote-send data with social
media information allowing to provide a better overview of a
disaster [
        <xref ref-type="bibr" rid="ref3 ref5">3, 5</xref>
        ]. For instance, in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], a system called "JORD" is introduced
to automatically collect information from diferent platforms of
social media and link it with remote-sensed data to provide a more
detailed story of a disaster. Similarly, a task to automatically link
social media with satellite images was introduced as a challenge in
ACM MM 20161.
      </p>
      <p>
        This paper provides a detailed description of the method
proposed by team MLRDCSE for the first task of Mediaeval2017
Multimedia and Satellite challenge [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The basic insight of the task
is to jointly utilize satellite imagery and social media as a source
of information to provide a detailed story of the disaster. The
proposed challenge is composed of two sub-tasks namely (i) Disaster
Image Retrieval from Social Media (DIRSM) and (ii) Flood
Detection in Satellite Images (FDSI). Detailed description of the tasks are
provided in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>PROPOSED APPROACH</title>
      <p>Figure 1 provides a block diagram of the proposed methodology.
As can be seen, the proposed approach is composed of three main
phases namely feature extraction, classification and fusion. In the
next sub-sections, we provide a detailed description of the each
phase.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Feature Extraction</title>
      <p>
        DIRSM is composed of three mandatory runs involving (i) Visual
information only (ii) Meta-data only and (iii) combination of
metadata and visual information. For visual information, we extract
Convolutional Neural Network (CNN) features via AlexNet [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
pre-trained on ImageNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and Places dataset [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] from each
image at hand. AlextNet is composed of 8 fully connected layers
including five convolutional and three fully connected layers. The
ultimate insight of the proposed scheme for visual information is
to utilize both object specific and scene-level information for the
representation of the disaster related images. A model pre-trained
on ImageNet corresponds to object specific information while the
one pre-trained on the places dataset is intended to extract
scenelevel information. This scheme has also been proved very efective
in social event detection in single images [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We extract a
4096dimensional feature vector from each model in cafe toolbox 2. On
the other hand, we also consider user tags, title and GPS information
from the available meta-data.
1http://www.acmmm.org/2016/wp-content/uploads/
2http://cafe.berkeleyvision.org/tutorial/
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Classification and Fusion</title>
      <p>
        In the proposed methodology, next steps correspond to
classification and fusion of the classification results obtained in the previous
step. For the classification purposes, we rely on Support Vector
Machines (SVM) based on its proven performances in object
recognition and classification [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We train separate Support Vector
Machines (SVM) classifiers for both CNN models on the complete
development dataset. Subsequently, test images are classified with
the trained classifiers providing results in the form of posterior
probabilities. On the other hand, for meta-data, we rely on Random
Forest classifier in WEKA Machine Learning library 3. The trained
classifier provide the results in terms of posterior probabilities.
      </p>
      <p>In the subsequent phase, we fuse the scores of the individual
classifiers in a late fusion method as shown in Equ. 1 where w1, w2
and w3 represent the weights used for each type of information
while p1, p2 and p3 represent the posterior probabilities obtained
with classifiers trained with features obtained with AlexNet
pretrained on ImageNet, AlexNet pre-trained on Places dataset and
meta-data, respectively.</p>
      <p>S = w1 ∗ p1 + w2 ∗ p2 + w3 ∗ p3 (1)</p>
      <p>In the current implementation, we use equal weights for each
classifier. Results can be further improved if some optimization
techniques, such as Genetic Algorithms (GA), are used. It is
important to mention that fusion method is used in both run 1 (fusion of
two classifiers trained on features extracted with both CNN models)
and run 3 (fusion of all types of information including meta-data
and visual information).
3</p>
    </sec>
    <sec id="sec-5">
      <title>RESULTS AND ANALYSIS</title>
      <p>Table 1 provides the experimental results of our method proposed
for the Mediaeval2017 Multimedia and Satellite task in terms of
average precision at cut-tofs 480. As can be seen, we achieve best
results in run 1, where we use visual information extracted with
two diferent CNN models of AlexNet pre-trained on ImageNet
and Places dataset. On the other hand, in Run 2, which is mainly
based on meta-data, we achieve the worst results among all runs by
having an average precision of just 22.83%. A significant diference
of around 64% can be noticed in the performances of meta-data and
visual information. This huge diference in the performances shows
a clear advantage of visual information over the meta-data in this
particular application. Moreover, run 2 also shows the limitations
of meta-data. Some common problems with meta-data includes
missing time stamps and geo-location information. Moreover, the
ambiguous meaning of user’s tags also afects the performance of
the model.</p>
      <p>The third run requires to combine meta-data and visual
information. In this experiment, our team achieves an average precision
of 83.73%, which is significantly lower than the performance with
visual information only. This is mainly caused by equally treating
the classifiers trained on visual information and meta-data. This
fact can also be concluded from the results of run 2, where the
classifier trained on meta-data achieves very low precision.</p>
      <p>In Table 2, we provide the experimental results of the proposed
method in terms of mean average precision at diferent cutofs
3http://www.cs.waikato.ac.nz/ml/weka/
namely 50, 100, 250 and 480. Again, better results are reported for
run 1 relying on visual information only. Similarly, worst results
are achieved with meta-data. It can also be noticed in Table 2 that
mean average precision at diferent cutofs for meta-data is lower
than the average precision at maximum cutof 480. However, in
the case of run 1 and run 3 diferent behaviour can be noticed by
achieving better performances at lower cutofs, which shows the
strength of visual information in diferentiating among flooded and
non-flooded images.</p>
      <p>Moreover, a significant increase can be noticed in the precision
using lower cutofs (mean of 50,100 and 250, 480), which shows
that increasing the cutof allows false positive to be included in the
threshold.
4</p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>This paper reports the description of the method proposed by team
MRLDCSE along with a detailed description and analysis of the
experimental results. For visual information, we rely on the
combination of object and scene-level information extracted through
two diferent Convolutional Neural Networks (CNN) models
pretrained on ImageNet and Places datasets. On the other hand, we
use user’s tags, title, description and geo-location information from
the available meta-data. Over all, better results are obtained with
visual information only. In contrast, meta-data produce worst results
among all the runs we submitted. We also noticed that the
inclusion of meta-data degrades the performance of the model when
combined with visual information in this particular application.</p>
      <p>In the current implementation, we are relying on a single deep
architecture, in future, we aim to incorporate multiple deep
architectures to better utilize visual information for the retrieval of flooded
images. Moreover, very low performance has been noticed with
meta-data, in future we aim to employ more sophisticated methods
to better utilize the additional information. An other interesting
direction can be using some optimization techniques for learning
weights of each classifier to fuse them, properly.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Kashif</given-names>
            <surname>Ahmad</surname>
          </string-name>
          , Nicola Conci, Giulia Boato, and
          <string-name>
            <surname>Francesco GB De Natale</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>USED: a large-scale social event detection dataset</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Multimedia Systems. ACM</source>
          ,
          <volume>50</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Kashif</given-names>
            <surname>Ahmad</surname>
          </string-name>
          , Francesco De Natale, Giulia Boato, and
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Rosani</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A hierarchical approach to event discovery from single images using MIL framework</article-title>
          .
          <source>In Signal and Information Processing (GlobalSIP)</source>
          ,
          <source>2016 IEEE Global Conference on. IEEE</source>
          ,
          <fpage>1223</fpage>
          -
          <lpage>1227</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Kashif</given-names>
            <surname>Ahmad</surname>
          </string-name>
          , Michael Riegler, Konstantin Pogorelov, Nicola Conci, Pål Halvorsen, and Francesco De Natale.
          <year>2017</year>
          .
          <article-title>JORD: A System for Collecting Information and Monitoring Natural Disasters by Linking Social Media with Satellite Imagery</article-title>
          .
          <source>In Proceedings of the 15th International Workshop on Content-Based Multimedia Indexing. ACM</source>
          ,
          <volume>12</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Kashif</given-names>
            <surname>Ahmad</surname>
          </string-name>
          , Michael Riegler, Ans Riaz, Nicola Conci,
          <string-name>
            <surname>Duc-Tien Dang-Nguyen</surname>
            , and
            <given-names>Pål</given-names>
          </string-name>
          <string-name>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The JORD System: Linking Sky and Social Multimedia Data to Natural Disasters</article-title>
          .
          <source>In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval. ACM</source>
          ,
          <volume>461</volume>
          -
          <fpage>465</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bischke</surname>
          </string-name>
          , Damian Borth,
          <string-name>
            <given-names>Christian</given-names>
            <surname>Schulze</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Dengel</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Contextual enrichment of remote-sensed events with social media streams</article-title>
          .
          <source>In Proceedings of the 2016 ACM on Multimedia Conference. ACM</source>
          ,
          <volume>1077</volume>
          -
          <fpage>1081</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Bischke</surname>
          </string-name>
          , Patrick Helber, Christian Schulze, Srinivasan Venkat, Andreas Dengel, and
          <string-name>
            <given-names>Damian</given-names>
            <surname>Borth</surname>
          </string-name>
          .
          <source>The Multimedia Satellite Task at MediaEval</source>
          <year>2017</year>
          :
          <article-title>Emergence Response for Flooding Events</article-title>
          .
          <source>In Proc. of the MediaEval 2017 Workshop (Sept</source>
          .
          <fpage>13</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2017</year>
          ). Dublin, Ireland.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Hyeran</given-names>
            <surname>Byun</surname>
          </string-name>
          and
          <string-name>
            <surname>Seong-Whan Lee</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Applications of support vector machines for pattern recognition: A survey. Pattern recognition with support vector machines (</article-title>
          <year>2002</year>
          ),
          <fpage>571</fpage>
          -
          <lpage>591</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Jia</given-names>
            <surname>Deng</surname>
          </string-name>
          , Wei Dong, Richard Socher,
          <string-name>
            <surname>Li-Jia</surname>
            <given-names>Li</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kai</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>Li</surname>
          </string-name>
          Fei-Fei.
          <year>2009</year>
          .
          <article-title>Imagenet: A large-scale hierarchical image database</article-title>
          .
          <source>In Computer Vision and Pattern Recognition</source>
          ,
          <year>2009</year>
          .
          <article-title>CVPR 2009</article-title>
          .
          <article-title>IEEE Conference on</article-title>
          . IEEE,
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Amanda</given-names>
            <surname>Lee</surname>
          </string-name>
          Hughes and
          <string-name>
            <given-names>Leysia</given-names>
            <surname>Palen</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Twitter adoption and use in mass convergence and emergency events</article-title>
          .
          <source>IJEM 6</source>
          ,
          <issue>3</issue>
          -
          <fpage>4</fpage>
          (
          <year>2009</year>
          ),
          <fpage>248</fpage>
          -
          <lpage>260</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Karen</surname>
            <given-names>E Joyce</given-names>
          </string-name>
          , Stella E Belliss, Sergey V Samsonov,
          <string-name>
            <surname>Stephen J McNeill</surname>
          </string-name>
          , and
          <string-name>
            <surname>Phil</surname>
          </string-name>
          J Glassey.
          <year>2009</year>
          .
          <article-title>A review of the status of satellite remote sensing and image processing techniques for mapping natural hazards and disasters</article-title>
          .
          <source>Progress in Physical Geography</source>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Alex</surname>
            <given-names>Krizhevsky</given-names>
          </string-name>
          , Ilya Sutskever, and
          <string-name>
            <given-names>Geofrey E</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Imagenet classification with deep convolutional neural networks</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>1097</volume>
          -
          <fpage>1105</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Chenliang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Aixin</given-names>
            <surname>Sun</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Anwitaman</given-names>
            <surname>Datta</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Twevent: segment-based event detection from tweets</article-title>
          .
          <source>In Proc. of ACM IKM. ACM</source>
          ,
          <volume>155</volume>
          -
          <fpage>164</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Ashbindu</given-names>
            <surname>Singh</surname>
          </string-name>
          .
          <year>1989</year>
          .
          <article-title>Review article digital change detection techniques using remotely-sensed data</article-title>
          .
          <source>International journal of remote sensing 10</source>
          ,
          <issue>6</issue>
          (
          <year>1989</year>
          ),
          <fpage>989</fpage>
          -
          <lpage>1003</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Brian</given-names>
            <surname>Stelter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Noam</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Citizen journalists provided glimpses of Mumbai attacks</article-title>
          .
          <source>The New York Times</source>
          <volume>30</volume>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Christos</surname>
            <given-names>Tzelepis</given-names>
          </string-name>
          , Zhigang Ma, Vasileios Mezaris, Bogdan Ionescu, Ioannis Kompatsiaris, Giulia Boato, Nicu Sebe, and
          <string-name>
            <given-names>Shuicheng</given-names>
            <surname>Yan</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Event-based media processing and analysis: A survey of the literature</article-title>
          .
          <source>Image and Vision Computing</source>
          <volume>53</volume>
          (
          <year>2016</year>
          ),
          <fpage>3</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Bolei</surname>
            <given-names>Zhou</given-names>
          </string-name>
          , Agata Lapedriza, Jianxiong Xiao, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Learning deep features for scene recognition using places database</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>487</volume>
          -
          <fpage>495</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>