<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Predicting Sperm Motility and Morphology using Deep Learning and Handcrafted Features</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Steven Hicks</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pål Halvorsen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Trine B. Haugen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorunn M. Andersen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oliwia Witczak</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hugo L. Hammer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Duc-Tien Dang-Nguyen</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mathias Lux</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Riegler</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Alpen-Adria-Universität Klagenfurt</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Oslo Metropolitan University</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>SimulaMet</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Bergen</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>27</fpage>
      <lpage>29</lpage>
      <abstract>
        <p>This paper presents the approach proposed by the organizer team (SimulaMet) for MediaEval 2019 Multimedia for Medicine: The Medico Task. The approach uses a data preparation method which is based on global features extracted from multiple frames within each video and then combines this with information about the patient in order to create a compressed representation of each video. The goal is to create a less hardware expensive data representation that still retains the temporal information of the video and related patient data. Overall, the results need some improvement before being a viable option for clinical use.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        In this paper, we detail the approach of the Medico Task
organization team (SimulaMet) as part of MediaEval 2019. The Medico Task
explores the challenge of using multimedia data to make the daily
work of medical doctors more eficient [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. This year’s task focuses
on automatically predicting sperm quality, in terms of motility and
morphology, based on a microscopic video recordings of human
semen. The task provides a dataset consisting of 85 videos and
associated patient data, which will be used to make predictions on
semen quality. The videos are taken from the VISEM dataset [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
which is an open-source dataset with barely any restrictions
regarding usage. More details on the 2019 Medico Task and the provided
dataset can be found in the overview paper [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and the paper on
VISEM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>The proposed approach is based on handcrafted features
extracted form multiple frames within each video and associated
patient/sensor data which is fed into a deep learning model for the
prediction. This approach is used to solve the prediction of motility
task and the prediction of morphology task, which are the tasks
required in order to participate in the competition. The prediction of
motiltiy task involves predicting three quality metrics (progressive,
non-progressive, and immotile) tied to the movement of the sperm
in a given semen sample. The prediction of morphology task focuses
on predicting more visual features of the sperm, namely, predicting
defects which may be present in the head, tail, or midpiece.</p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>As the organization team, our main goal for this year’s Medico Task
was to present a baseline approach which utilize all the available
data in the dataset to make a prediction on sperm motility and
morphology (as the tasks require). Furthermore, we wanted our
approach to be computationally eficient so to not require expensive
hardware or a complex setup procedure. Therefore, we decided to
base our contribution on handcrafted features which were extracted
from multiple frames within each video in the provided dataset.
These handcrafted features are combined with each category of
associated patient data in order to train a deep learning model
to make a prediction on each of the sperm quality features. In
the following few section, we will describe the approach in more
detail. This includes a description on how the data was prepared,
information about the model architecture, and how each model was
trained.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Data Preparation</title>
      <p>
        To prepare the data for training and evaluation, we start by
extracting the first two frames of each second for 60 seconds for each video.
The result is a sub-sample of each video containing 120 frames from
the first minute of each of the 85 videos in the dataset. From this
step, we extract four diferent types of global features from each of
the frames, namely, Edge Histogram, Tamura, Luminance Layout
and Simple Color Histogram [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. All image features were extracted
using the LIRE library Lire [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which is a popular open-source
library for image retrieval. The global image features are
concatenated row-wise to represent a single frame. The extracted image
feature vectors are then concatenated column-wise into a 225 × 120
matrix, where the columns represent each frame, and the rows
represent the extracted features.
      </p>
      <p>The motivation behind this approach is that each column
contains the visual features of a single frame, while the temporal
information is represented through the change of feature values across
the diferent columns. To add the information about the patient, we
simply concatenate the values to each frame column so that each
patient and sensor value will be the same for each frame in a given
training sample. We create these video representations using each
of the five provided categories of patient information values in the
dataset, namely, sex hormones, patient-related data (age, body mass
index (BMI), and days of sexual abstinence), fatty acids in serum,
fatty acids in spermatozoa, and sperm analysis data. Each of the five
video categories were used to train two deep neural networks, one
for predicting morphology and one for predicting motility. Note
that for the sperm analysis data, we removed the data regarding
motility and morphology when used to predict these values in each
respective task. This is important as to not feed the ground truth
values to the network during training.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Model Architecture and Training</title>
      <p>As explained in the previous section, the data format used to train
our deep neural networks is a matrix of extracted image features
from 120 frames within each video and associated patient and sensor
data. The exact input shape varies depending on the category of
patient information used in the representation (number of variables
vary depending on the category), but rest of the network is kept the
same. We use a novel convolutional neural network architecture to
perform our experiments.</p>
      <p>The architecture itself is modeled to analyze the video
representations using an ensemble of similar networks (network architecture
shown in Figure 1). The network consists of an inital convolution
block, after which the output gets passed through multiple residual
modules. The output of each network is then concatenated before
being passed through two fully-connected layers and then making
the prediction. The network is repeated multiple times in order to
create an ensemble of networks which each analyze the same input.
The number of networks used is a parameter which may be tuned,
from which we decided to use three based on some internal testing.</p>
      <p>
        As previously mentioned, the same model architecture was used
for both the motility tasks and the morphology task. The model was
trained with a learning rate of 0.00001 using Nadam [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to optimize
the weights. We used a mean absolute error (MAE) to calculate loss
on a batch size of 16 diferent samples. Overall, each experiment
was trained for a maximum of 5000 epochs, but stopped training
if the loss of the model did not improve over the last 300 epochs.
Note that the model used for evaluation is the one which achieved
the best MAE on the validation set. The hardware used to train
all models was a desktop computer running Linux with an Nvidia
GTX 1080TI, Intel core i7 processor running at 3.6 gigahertz, and
16 gigabytes of RAM. Models were implemented using the deep
learning library Keras [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] with a TensorFlow [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] back-end. Due to
the small number of training samples, despite training for many
epochs, no experiment took longer than 1 hour to train.
3
      </p>
    </sec>
    <sec id="sec-5">
      <title>RESULTS AND DISCUSSION</title>
      <p>Looking at Table 1 and Table 2, we see the results for the Prediction
of Morphology Task the Prediction of Motility Task. Overall, the
results show that the deep neural networks for both tasks are able
to learn something from the data, but the performance is overall
quite poor when compared to the ZeroR baseline. For future work,
we aim to use features extracted from a deep neural network.
4</p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSION</title>
      <p>
        In this paper, we described the approach submitted by the 2019
Medico organization team (SimulaMet). The presented method
used handcrafted features and associated patient data to train a
deep learning model to predict sperm quality in terms of motility
and morphology. Based on the results, we see that the future of
automatic sperm quality prediction is promising, but requires more
work before being used in any real-world scenario. The way of
representing the video into a single image also allows for future
experiments where we want to use grad cam methods, e.g. [
        <xref ref-type="bibr" rid="ref5 ref7">5, 7</xref>
        ],
to explain the important parts of the video and data.
      </p>
      <sec id="sec-6-1">
        <title>Input</title>
      </sec>
      <sec id="sec-6-2">
        <title>Convolution</title>
      </sec>
      <sec id="sec-6-3">
        <title>Convolution</title>
      </sec>
      <sec id="sec-6-4">
        <title>Average Pooling</title>
      </sec>
      <sec id="sec-6-5">
        <title>Residual Module</title>
      </sec>
      <sec id="sec-6-6">
        <title>Average Pooling</title>
      </sec>
      <sec id="sec-6-7">
        <title>Concatenation</title>
      </sec>
      <sec id="sec-6-8">
        <title>Fully-Connected</title>
      </sec>
      <sec id="sec-6-9">
        <title>Fully-Connected</title>
      </sec>
      <sec id="sec-6-10">
        <title>Prediction</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Martín</given-names>
            <surname>Abadi</surname>
          </string-name>
          , Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis,
          <string-name>
            <given-names>Jefrey</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Matthieu</given-names>
            <surname>Devin</surname>
          </string-name>
          , Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geofrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke,
          <string-name>
            <given-names>Yuan</given-names>
            <surname>Yu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xiaoqiang</given-names>
            <surname>Zheng</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <source>TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems</source>
          . (
          <year>2015</year>
          ). https://www.tensorflow.org/ Software available from tensorflow.
          <source>org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>François</given-names>
            <surname>Chollet</surname>
          </string-name>
          and others.
          <source>2015</source>
          . Keras. https://keras.io. (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Timothy</given-names>
            <surname>Dozat</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Incorporating Nesterov Momentum into adam</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Trine</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Haugen</surname>
          </string-name>
          , Steven A.
          <string-name>
            <surname>Hicks</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jorunn M. Andersen</surname>
            , Oliwia Witczak, Hugo L. Hammer, Rune Borgli, Pål Halvorsen, and
            <given-names>Michael A.</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>VISEM: A Multimodal Video Dataset of Human Spermatozoa</article-title>
          .
          <source>In Proceedings of the 10th ACM on Multimedia Systems Conference (MMSys'19)</source>
          . https://doi.org/10.1145/3304109.3325814
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Steven</given-names>
            <surname>Hicks</surname>
          </string-name>
          , Sigrun Eskeland, Mathias Lux, Thomas de Lange, Kristin Ranheim Randel, Mattis Jeppsson, Konstantin Pogorelov, Pål Halvorsen, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Mimir: An Automatic Reporting and Reasoning System for Deep Learning Based Analysis in the Medical Domain</article-title>
          .
          <source>In Proceedings of the 9th ACM Multimedia Systems Conference (MMSys '18)</source>
          . ACM, New York, NY, USA,
          <fpage>369</fpage>
          -
          <lpage>374</lpage>
          . https://doi.org/10.1145/3204949.3208129
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Steven</given-names>
            <surname>Hicks</surname>
          </string-name>
          , Pål Halvorsen, Trine B Haugen,
          <string-name>
            <surname>Jorunn M Andersen</surname>
            ,
            <given-names>Oliwia</given-names>
          </string-name>
          <string-name>
            <surname>Witczak</surname>
          </string-name>
          , Konstantin Pogorelov, Hugo L Hammer,
          <string-name>
            <surname>Duc-Tien</surname>
            Dang-Nguyen,
            <given-names>Mathias</given-names>
          </string-name>
          <string-name>
            <surname>Lux</surname>
            , and
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Medico Multimedia Task at MediaEval 2019</article-title>
          . In CEUR Workshop Proceedings - Multimedia Benchmark Workshop (MediaEval).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hicks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. V.</given-names>
            <surname>Anonsen</surname>
          </string-name>
          , T. de Lange,
          <string-name>
            <given-names>D.</given-names>
            <surname>Johansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jeppsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ranheim Randel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Losada</given-names>
            <surname>Eskeland</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Dissecting Deep Neural Networks for Better Medical Image Classification and Classification Understanding</article-title>
          .
          <source>In 2018 IEEE 31st International Symposium on Computer-Based Medical Systems (CBMS)</source>
          .
          <volume>363</volume>
          -
          <fpage>368</fpage>
          . https://doi.org/10.1109/CBMS.
          <year>2018</year>
          .00070
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Mathias</given-names>
            <surname>Lux</surname>
          </string-name>
          and
          <string-name>
            <given-names>Savvas A.</given-names>
            <surname>Chatzichristofis</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Lire: Lucene Image Retrieval: An Extensible Java CBIR Library</article-title>
          .
          <source>In Proceedings of the 16th ACM International Conference on Multimedia (MM '08)</source>
          . ACM, New York, NY, USA,
          <fpage>1085</fpage>
          -
          <lpage>1088</lpage>
          . https://doi.org/10.1145/1459359.1459577
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Mathias</given-names>
            <surname>Lux</surname>
          </string-name>
          , Michael Riegler, Pål Halvorsen, Konstantin Pogorelov, and
          <string-name>
            <given-names>Nektarios</given-names>
            <surname>Anagnostopoulos</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>LIRE: open source visual information retrieval</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Multimedia Systems. ACM</source>
          ,
          <volume>30</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Riegler</surname>
          </string-name>
          , Mathias Lux, Carsten Griwodz, Concetto Spampinato, Thomas de Lange, Sigrun L Eskeland, Konstantin Pogorelov, Wallapak Tavanapong,
          <string-name>
            <surname>Peter T Schmidt</surname>
          </string-name>
          ,
          <article-title>Cathal Gurrin, and</article-title>
          <string-name>
            <surname>others.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Multimedia and medicine: Teammates for better disease detection and survival</article-title>
          .
          <source>In Proceedings of the 24th ACM international conference on Multimedia. ACM</source>
          ,
          <volume>968</volume>
          -
          <fpage>977</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>