<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Concept detection based on multi-label classi cation and image captioning approach - DAMO at ImageCLEF 2019</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jing Xu</string-name>
          <email>xujing212@buaa.edu.cn</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wei Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chao Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yu Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ying Chi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xuansong Xie</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiansheng Hua</string-name>
          <email>xiansheng.hxsg@alibaba-inc.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Alibaba Group DAMO Academy AI Center</institution>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Beihang University</institution>
          ,
          <addr-line>Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Medical image captioning is an important and challenging task, which covers computer vision and natural language processing. This ImageCLEF 2019 [6] Caption competition is dedicated to research this eld. The purpose of this year challenge is using radiological images to detect the concepts representing the key information. In this paper, we illustrate the proposed method to address the issue, based on multilabel classi cation model and CNN-LSTM architecture with attention mechanism. We also perform a detailed analysis and processing for the overall dataset and demonstrate performance with the baseline in the caption prediction task. In nal evaluation, we completed 9 submissions and ranked second among 12 participants with our best mean F1-score.</p>
      </abstract>
      <kwd-group>
        <kwd>Radiology Image caption Concept detection Multi-label classi cation Encoder-decoder</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Medical images, such as radiological images, are widely used in hospital
diagnosis and disease treatment. The reading and summarization of medical images is
usually performed by experienced medical professionals, and obtaining
information from radiological medical images is a time-consuming and laborious task.
Therefore, it is essential to automatically and e ciently extract vital
information from medical images. ImageCLEF 2019 Caption [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is the third year of the
challenge, starting in 2017, to analyze and solve the problem of medical image
caption. The organizing committee provided a large corpus of medical radiology
images and UMLS (Uni ed Medical Language System) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] concepts pairs, and
the purpose of this task is to detect the relevant concepts based on the visual
radiology images. Evaluation criteria is conducted in terms of F1-score between
concepts predicted and ground truth concepts.
      </p>
      <p>
        Inspired by the recent successes of convolutional architectures on other
endto-end frameworks [
        <xref ref-type="bibr" rid="ref14 ref16 ref3 ref5">3,5,14, 16</xref>
        ], we study convolutional architectures for the task
of image concept detection. Speci cally, we handle each concept sequence
corresponding to each radiological image as a set of labels, and attempt to build
a multi-label classi cation network to solve the task. Furthermore, increasing
research has been devoted to image captioning, and almost all of the current
proposed methods are under the framework of CNN+RNN [
        <xref ref-type="bibr" rid="ref10 ref15 ref17 ref9">9, 10, 15, 17</xref>
        ]. To
imitate the human visual attention mechanism, the attention module has been
applied. Hence, we adopt the encoder-decoder network, in which a basic CNN is
used for the vision feature extractor, and an LSTM is employed to generate
sentences due to the ability of learning long term dependencies through a memory
cell.
      </p>
      <p>The paper is organized as follow: Section 2 describes the analysis of the
overall data, Section 3 introduces the method for the concept detection task,
Section 4 demonstrates the details of the experiments and results, and Section
5 discusses and concludes our work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Data analysis</title>
      <p>
        In the ImageCLEF 2019 Concept Detection Task, the overall dataset contains
70,786 radiology images of several medical imaging modalities. The images are
collected from open access biomedical journal articles (PubMed Central) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
and the corresponding UMLS concepts that totals 5,528 are extracted from the
original image caption. Training dataset includes 56,629 images, and the number
of associated concepts is 5216. Validation dataset includes 14,157 images, and
the concepts related is 3233. It is worth mentioning that the sequences of the
training dataset does not include the total concepts, and 312 concepts appear
only in the validation dataset.
      </p>
      <p>To further understand the datasets, we performed statistical analysis to
reveal the overall data distribution. The Top-10 concepts descriptions, and the
statistics of the length of concept sequence corresponding to each image and the
distribution of concept frequency are shown in Table 1 and Fig. 1, respectively.
From Table 1, we can see that among the 10 concepts of high frequency, the
C0043299 and C1962945, or the C0040395 and C0040405 have similar meanings.
Thus, we conducted correlation analysis for all pairs of concepts. See section 4.1
for details. We counted the length of the concept sequences corresponding to
each images in dataset. Only one or two concepts of the sequence account for
17.09% of the overall sample. From Fig. 1, the total number of concepts with
a frequency less than 3 is 2,293, accounting for 41.48% of the entire concept
dictionary, while there are only 15 concepts with a frequency of more than 3,000
times.
We design two distinct methods to address the concept detection issue, one is to
transform the issue into a multi-label classi cation problem, the other is to treat
it as an image captioning task, using the encoder-decoder network to generate
the concepts.
Since there is no strong contextual correlation between the concepts of an image,
we transform this task into a multi-label classi cation problem. That is, an image
has several labels. Let li be the label of i-th image, as follows:
li = [ci;1; ci;2; : : : ; ci;j ; : : : ci;n]
(1)
where n is the total number of labels. If the i-th image has the j-th label, the
ci;j is set to 1, else 0.</p>
      <p>
        We utilize the latest deep learning method to solve this problem, which has
achieved great success on the eld of image processing, such as classi cation,
captioning. Empirically the deeper the network is, the richer the features extracted
on di erent levels. While the drawbacks of gradient vanishing and explosion make
it di cult to converge. To overcome this problem, He et al. proposed ResNet [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
which reformulates the layers as learning residual functions with reference to the
layer inputs, instead of learning unreferenced functions, and had won the rst
prize on ImageNet competitions. We choose the pre-trained ResNet-101 model
on the ImageNet dataset [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] as backbone in our multi-label classi cation
experiment. The overall process is shown in Fig. 2. An image is rstly preprocessed
to adapt to the input of the net, and feed forward to the net to get the output
feature vectors. Then passing by a fully connection layer with sigmoid activation
function to calculate the probability of each class. If the probability is greater
than 0.5, we assert the input image belongs to that class. Finally the predicted
labels obtained, which can be re ected back to the original concepts.
Considering that the concept detection task is to generate text information from
the corresponding radiology images, we attempt to address it with the
CNNRNN model framework with attention mechanism. Typically, the model that
generates a sequence of concepts will use an encoder to encode the input into a
xed form and use a decoder to decode it into a sequence verbatim.
      </p>
      <p>
        In our approach, the encoder built upon the pre-trained ResNet-101 is rst
applied to extract visual features from the input images. We resize the input
images normalized by the mean and standard deviation to 224 224 for uniformity,
and then ne-tuned the convolutional blocks on the given medical dataset with
a smaller learning rate. We utilize the visual features captured by the conv 5
convolution block in the ResNet to better describe the local information.
Meanwhile, the model combine soft attention mechanism [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] to dynamically select
spatial characteristic of the input image.
      </p>
      <p>
        In decoder, We apply a long short-term memory (LSTM) network [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] that
produces a caption by generating one word at every time step conditioned on a
context vector capturing the visual information, the previous hidden state and
the previously generated concepts. After extracting visual features in CNN, we
transform the encoded image to create the initial hidden state h and cell state c
for the LSTM decoder. At each decode step, the encoded image and the previous
hidden state is used to generate weights for each pixel in the attention network.
Finally, the previous generated concept and the weighted average of the encoded
image are fed to the LSTM decoder to generate the next concept with the highest
score. In addition, we also perform beam search with di erent beam sizes instead
of sampling the maximum probability words.
4
4.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experiments and Results</title>
      <sec id="sec-3-1">
        <title>Data preprocessing</title>
        <p>Concept association mining It has been found that some concepts have a
certain relevance as they often appear simultaneously in di erent radiological
images. Thus, we lter out the high-correlation concept combinations from the
high-frequency concepts in training dataset. First, we utilize association rule
mining to search for relationships between all concepts de ned as a set of items
C = fc1; c2; c3; : : : cM g, and I = fi1; i2; i3; : : : iN g represents a collection of all
training samples, where im is the concept sets corresponding to each image.
Obviously, iN C. The form of association rule for the concept sets X and Y
can be written as: X ! Y , where X C, Y C, and X \ Y = ;. The support
is the fraction of training set that contain both X and Y , and the conf idence
represents the measure that how often concepts in Y appear in sample sets that
contain X. We suppose represents the frequency of occurrence of an item-set.
Speci cally,</p>
        <p>support(X ! Y ) =
conf idence(X ! Y ) =
(X [ Y )</p>
        <p>N
(X [ Y )
(X)
(2)
(3)</p>
        <p>Second, the concept subsets divided with support &gt; 0:02 is a total of 99,
and from which we select the combinations contain with the most elements with
conf idence &gt; 0:9. Finally, we de ne 9 di erent concept combinations as 9 new
concepts, called concept grouping, as shown in Table 2. During the training
process, we replace the concepts involved with the 9 new concepts and de ne
the dataset changed as Cg, and then map them in predicted results.
Data ltering It is obvious that the dataset is extremely unbalanced through
the statistics above. The low-frequency concepts would not only not be learned,
but bring great bias to the model. Therefore, we lter out the concepts which
are indeed rare. Firstly, we pick out all the concepts which only occurs once and
get the corresponding images. Then checking all the related concepts on each
image one by one, if the frequencies of all related concepts are once either, the
image would be moved out of the dataset. The ltered dataset denotes as Df;1
with 163 concepts and 98 images omitted.</p>
        <p>At the same time, we also roughly lter out the concepts with frequency
less than 3 or 5 times de ned Df;3 and Df;5, to avoid these noises a ecting the
overall dataset distribution.</p>
        <p>Data redivision Since the pre-divided training dataset provided by organizer
dose not contain all the concepts need to be learned, we re-divided all the data
as follows:</p>
        <p>(a) Picking out the images form validation dataset, as mentioned above,
which has the concepts that are not occurred in training dataset.</p>
        <p>(b) Changing these images slightly by random transformation, such as
mirroring, rotation, etc.</p>
        <p>(c) Appending these transformed images to the training dataset, while the
original validation dataset keeps unchanged.</p>
        <p>Image preprocess and normalization In order to make full use of the
provided data, several data augmentation operations are introduced in our
experiment. Speci cally, the image is rstly random ipped horizontally or vertically
with a probability of 0.5, and then resized to di erent scale ranging from 0.6 to
1.2 with bilinear interpolation. Finally, the transformed image is random cropped
into the size of 224 224 before input into the backbone net.
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Training parameters</title>
        <p>
          Multi-label classi cation method Parameters of the last fully connection
layer are initialized by MSRA method [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and the F1-score is utilized as the
criteria. The batch size and the max iteration epochs are set to 64 and 100
respectively. We apply the Adam [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] optimizer to ne-tune the model with an
initial learning rate of 0.001. The training procedure is shown in the Fig. 3.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>CNN-RNN with attention mechanism method With ne-tuning the en</title>
        <p>coder, the model was trained with cross entropy loss for 30 epochs, batch size of
20 and dropout rate of 0.5. In concept generation, we set the dimensions of all
hidden states and word embeddings as 512. We used the Adam optimizer and
the learning rates for the CNN and the LSTM were 1e 4 and 4e 4 respectively.
Early stopping was used to prevent over- tting when performance on a
validation dataset started to degrade. The best model saved was used to predict the
sequence of concepts in the test images.
4.3</p>
      </sec>
      <sec id="sec-3-4">
        <title>Speci c description in each run</title>
        <p>We completed a total of 9 graded submissions before the deadline, the evaluation
results for our submitted runs is shown in Table 3 and the speci c method for
each run is as follows:
Run ID 27103: In this run, we applied multi-label classi cation model
introduced in Section 3.1 and the dataset trained was ltered dataset Df;1. We chose
the pre-trained ResNet-101 model as backbone in our experiment and performed
concept grouping Cg in data preprocessing. Meanwhile, we applied the Adam
optimizer to ne-tune the model with an initial learning rate of 1e 3, and the
max epoch was 100 in training procedure.</p>
        <p>Run ID 27184: This process is similar to ID 27103. We chose the pre-trained
ResNet-101 model and the learning rate is 1e 3. The multi-label classi cation
model was trained with ltered dataset Df;1 except for the concept grouping
strategy and the max epoch was 60 in training procedure.</p>
        <p>Run ID 26786: This run we utilized the CNN-RNN architecture with attention
mechanism, based on pre-trained ResNet-101 and LSTM. We used the Adam
optimizer and the learning rates for the CNN and the LSTM were 1e 4 and 4e 4
respectively. In the training dataset Df;3, concepts occurring less frequently than
3 was ignored. Early stopping was used, and the best model saved was used to
predict the sequence of concepts in the test images.</p>
        <p>Run ID 27107: We combined the predicted results of ID 27103 and ID 26786,
that is, the nal results in test dataset was the union of two methods for each
sample.</p>
        <p>Run ID 27106: Similarly, the nal results in this run was the union of the
predicted results of ID 27184 and ID 26786.</p>
        <p>
          Run ID 27188&amp;26877&amp;27111&amp;27158: These process were based on ID 26786,
the pre-trained ResNet-101 was used for the vision model, and an LSTM was
employed to generate sentences. The dataset trained was ltered dataset Df;5.
Otherwise, we made an attempt to apply reinforcement learning [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] in decoder, and
the experimental results performed well on validation dataset but were poorly
e ective on the test dataset.
The evaluation for the caption detection task is conducted using the mean
F1score. As shown in Table 3, among the 61 results submitted by all participants,
Run ID 27103 based multi-label classi cation model has achieved the better
performance with the mean F1-score of 0.2655. We mitigate the impact of
extreme data imbalance on the model by setting threshold culling noise data, and
utilize association rule mining to search for the high-correlation concept
combinations. For CNN-LSTM network method, the model did not perform well on
the test dataset. Since in the sequences corresponding to the radiology images,
the concept exists independently, although some frequent concepts have slight
correlation.
        </p>
        <p>Overall, we have completed this challenge in the medical image concept
detection task and our group rank second among 12 participants (see Table 4).
The method adopted has achieved preliminary results and we will further
investigate the medical image captioning task based on higher quality datasets and
advanced deep learning algorithms.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The uni ed medical language system (umls): integrating biomedical terminology</article-title>
          .
          <source>Nucleic acids research 32(suppl 1)</source>
          ,
          <source>D267{D270</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Delving deep into recti ers: Surpassing humanlevel performance on imagenet classi cation</article-title>
          .
          <source>In: Proceedings of the IEEE international conference on computer vision</source>
          . pp.
          <volume>1026</volume>
          {
          <issue>1034</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>770</volume>
          {
          <issue>778</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9(8)</source>
          ,
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Der Maaten</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weinberger</surname>
            ,
            <given-names>K.Q.</given-names>
          </string-name>
          :
          <article-title>Densely connected convolutional networks</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>4700</volume>
          {
          <issue>4708</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Peteri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cid</surname>
            ,
            <given-names>Y.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liauchuk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klimuk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarasau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abacha</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Datla</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dang-Nguyen</surname>
            ,
            <given-names>D.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelka</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Herrera</surname>
            ,
            <given-names>A.G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavallieratou</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>del Blanco</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          , Rodr guez, C.C.,
          <string-name>
            <surname>Vasillopoulos</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karampidis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>ImageCLEF 2019: Multimedia retrieval in medicine, lifelogging, security and nature</article-title>
          . In:
          <article-title>Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the 10th International Conference of the CLEF Association (CLEF</source>
          <year>2019</year>
          ),
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Lugano,
          <source>Switzerland (September 9-12</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
          </string-name>
          , J.:
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>arXiv preprint arXiv:1412.6980</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.E.:
          <article-title>Imagenet classi cation with deep convolutional neural networks</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>1097</volume>
          {
          <issue>1105</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kulkarni</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Premraj</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ordonez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dhar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berg</surname>
            ,
            <given-names>A.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berg</surname>
            ,
            <given-names>T.L.</given-names>
          </string-name>
          :
          <article-title>Babytalk: Understanding and generating simple image descriptions</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>35</volume>
          (
          <issue>12</issue>
          ),
          <volume>2891</volume>
          {
          <fpage>2903</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiong</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parikh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
          </string-name>
          , R.:
          <article-title>Knowing when to look: Adaptive attention via a visual sentinel for image captioning</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>375</volume>
          {
          <issue>383</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pelka</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garc</surname>
            a Seco de Herrera,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H.:
          <article-title>Overview of the ImageCLEFmed 2019 concept prediction task</article-title>
          .
          <source>In: CLEF2019 Working Notes. CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Lugano,
          <source>Switzerland (September</source>
          <volume>09</volume>
          -12
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Rennie</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marcheret</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mroueh</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ross</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Self-critical sequence training for image captioning</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <volume>7008</volume>
          {
          <issue>7024</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Roberts</surname>
          </string-name>
          , R.J.:
          <article-title>Pubmed central: The genbank of the published literature (</article-title>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Simonyan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          .
          <source>arXiv preprint arXiv:1409.1556</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toshev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erhan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Show and tell: A neural image caption generator</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>3156</volume>
          {
          <issue>3164</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dollar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Aggregated residual transformations for deep neural networks</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>1492</volume>
          {
          <issue>1500</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiros</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Courville</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakhutdinov</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zemel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Show, attend and tell: Neural image caption generation with visual attention</article-title>
          .
          <source>arXiv preprint arXiv:1502.03044</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>