<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Competitive Deep Neural Network Approach for the ImageCLEFmed Caption 2020 Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marimuthu Kalimuthu</string-name>
          <email>marimuthu.kalimuthu@dfki.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabrizio Nunnari</string-name>
          <email>fabrizio.nunnari@dfki.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Sonntag</string-name>
          <email>daniel.sonntag@dfki.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>German Research Center for Artificial Intelligence (DFKI) Saarland Informatics Campus</institution>
          ,
          <addr-line>66123 Saarbrücken</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The aim of ImageCLEFmed Caption task is to develop a system that automatically labels radiology images with relevant medical concepts. We describe our Deep Neural Network (DNN) based approach for tackling this problem. On the challenge test set of 3,534 radiology images, our system achieves an F1 score of 0.375 and ranks high, 12th among all systems that were successfully submitted to the challenge, whereby we only rely on the provided data sources and do not use any external medical knowledge or ontologies, or pretrained models from other medical image repositories or application domains.</p>
      </abstract>
      <kwd-group>
        <kwd>Medical Imaging</kwd>
        <kwd>Concept Detection</kwd>
        <kwd>Image Labeling</kwd>
        <kwd>MultiLabel Classification</kwd>
        <kwd>Deep Convolutional Neural Networks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        ImageCLEF organises 4 main tasks for the 2020 edition with a global
objective of promoting the evaluation of technologies for annotation, indexing, and
retrieval of visual data with the aim of providing information access to large
collections of images in various usage scenarios and application domains, including
medicine [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Interpreting and summarizing the insights gained from medical images is a
time-consuming task that involves highly trained experts and often represents a
bottleneck in clinical diagnosis pipelines. Consequently, there is a considerable
need for automatic methods that can approximate this mapping from visual
information to condensed textual descriptions. The more image characteristics are
known, the more structured are the radiology scans and hence, the more efficient
are the radiologists regarding interpretation, see https://www.imageclef.org/
2020/medical/caption/ and [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        Recent years have witnessed tremendous advances in deep neural networks in
terms of architectures, optimization algorithms, tooling and techniques for
training large networks, and handling multiple modalities (e.g., text, images, videos,
speech, etc.). In particular, deep convolutional neural networks have proved to
be extremely successful image encoders and have thus become the de facto
standard for visual recognition [
        <xref ref-type="bibr" rid="ref13 ref2 ref3">2,3,13</xref>
        ]. We have been working on machine learning
problems in several medical application domains [
        <xref ref-type="bibr" rid="ref14 ref15 ref16">14,15,16,17</xref>
        ] in our projects, see
https://ai-in-medicine.dfki.de/. In this paper, we describe how we built a
competitive deep neural network approach based on these projects.
      </p>
      <p>The rest of the paper is organized as follows. In Section 2, we formally
describe the challenge and its goal. In Section 3, we present some statistics on
the dataset and explain the approach we adopted to tackle the challenge. In
Section 4, we describe the experiments that we conducted with different
architectures and introduce a new loss function that addresses the sparsity problem
in ground truth labels. In Section 5, results are presented followed by a short
discussion. Finally, we summarize our work in Section 6 and discuss some future
directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The Challenge</title>
      <p>
        The overarching goal of ImageCLEFmed Caption challenge is to assist medical
experts such as radiologists in interpreting and summarising information
contained in medical images. As a first step towards this goal, a simpler task would
be to detect as many key concepts as possible, with the goal that these concepts
can then be composed into comprehensible sentences, and eventually into
medical reports. For full details about the challenge, we refer the reader to Pelka et
al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>The challenge has evolved over the years, since its first edition in 2017, to
focus only on radiology images in this year’s version, and incorporating the lessons
learned from previous years. The aim of 2020 ImageCLEFmed Caption challenge
is to develop a system that would automatically assign medical concepts to
radiology images that were sorted into 7 different categories (see Table 1). More
concretely, given an image (I), the objective is to learn a function F that maps
I to a set of concepts (C1, C2, ..., CvI ) where vI is the number of concepts
associated with I. A peculiarity of this challenge is that v k, where k is the total
number of unique labels.</p>
      <p>F : I ! (C1; C2; :::; CvI )</p>
      <p>On the challenge data, the ground truth concepts are not known for images
in the test set. The only known constraint is that predictions of the test set must
be submitted with a maximum of 100 non-repeating concepts per image.</p>
      <p>The performance of submissions to the challenge are evaluated on a withheld
test set of 3,534 radiology images using the F 1 score evaluation metric, which is
defined as the harmonic mean of precision and recall values.</p>
      <p>Firstly, the instance level F1 scores are computed using the predicted concepts
for images in the test set. Later, an average F1 score is computed over all images
in the test set using scikit-learn1 library’s default binary averaging method. This
yields the final F1 score for an accepted submission.</p>
      <p>All registered teams are allowed for a maximum of 10 submissions. Successful
submissions and teams are then ranked based on the achieved F1 scores and
results are made publicly available.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Method</title>
      <p>We provide information about the dataset and some analysis that we performed
on it during the exploratory phase. Then, we discuss our learning approach and
a suitable data preparation strategy for training our models.
3.1</p>
      <p>
        Dataset
All participants are provided with ImageCLEFmed Caption dataset which is a
subset of the ROCO dataset [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. It is a multi-modal, medical images dataset
containing radiology images that are each labelled with a set of medical concepts,
called as Concept Unique Identifiers (CUIs) in the literature. As common in
many challenges, the provided images are already split into train, validation,
and test sets. Images of the first two sets come with ground truth labels, while
the test set contains only images.
      </p>
      <p>Moreover, images in each of the splits are sorted into one of the 7 categories
as described in tables 1 and 2; such information can be inferred from the name
of the sub-directory containing the images. As we can observe from Table 2, the
images are not equally distributed across categories, indicating an imbalance in
the dataset. The DRCT category, which contains images captured using
Computerized Tomography (CT), has the highest representation, followed by X-ray
images (DRXR), while the least number of images (approximately 410 th of the
images in DRCT) are seen in DRCO category.</p>
      <p>1https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_
score.html</p>
      <p>Split DRATNotaDlRNCuOmbDeRrCofTImDRagMesRinDtRhPeECaDtRegUoSryDoRfXR Total
Train 4,713 487 20,031 11,447 502 8,629 18,944 64,753
Val 1,132 73 4,992 2,848 74 2,134 4,717 15,970
Test 325 49 1,140 562 38 502 918 3,354</p>
      <p>Total 6,170 609 26,163 14,857 614 11,265 24,579 84,077</p>
      <p>
        Despite having these category labels as a meta-information, we could not
leverage them during model training due to time constraints. Thus, the
relationship between CUIs and image category labels is, for us, still uninvestigated.
Possibly, in a follow-up work we will investigate on how to exploit such
metainformation to improve classification performance of predictive models, for
example by using one of the metadata fusion strategies tested by the authors in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
3.2
      </p>
      <p>Data Analysis
Here, we outline some insights that we gained after performing analysis on the
CUIs and the category labels meta-information. Furthermore, this section sheds
light on the imbalance in the dataset, which is a common problem in many
research domains.</p>
      <p>Figure 1 provides a conceptual representation of the input, viz. images paired
with relevant concepts. In this case, both images are labelled with four CUIs,
the descriptions of which are provided by Unified Medical Language System
(UMLS)2 terms. These terms are depicted in Figure 1 merely for the purpose
of understanding since the mapping of CUIs to UMLS terms is not part of the
provided dataset.</p>
      <p>Figure 2 shows the frequencies of top 30 CUIs on the training split. Two CUIs
(C0040398, C0040405 ), which both occur around 20k times in the training set,
dominate this list and represent the concept “computer assisted tomography”.
Similar behavior is observed on the validation set (see Figure 3). However, the
frequencies of top 2 CUIs in this case are only around 4.6k.</p>
      <p>Figure 4 depicts a histogram representation of the number of images and
the CUI counts on the training set. For instance, there are around 5,200 images
that have exactly two CUIs as ground truth labels. On the contrary, there are
only around 100 images that have exactly 50 CUIs as ground truth labels in the
provided training set. This histogram is truncated at CUI count 50 (x-axis) for
clarity and uncluttered representation.</p>
      <p>In a similar manner, Figure 5 shows a histogram representation of the
number of images and the CUI counts on the validation set. For instance, there are
around 1,300 images that have exactly two CUIs as ground truth labels. On the
contrary, there are only around 30 images that have exactly 50 CUIs as ground
2https://www.nlm.nih.gov/research/umls/</p>
      <p>UMLS Term (Concept)
C2951888 set of bones of skull
C0037303 set of bones of cranium
C0032743 tomogr positron emission
C0342952 increased basal metabolic rate
C0022742 knees
C0043299 x-ray procedure
C1260920 kneel
C0030647 bone, patella
truth labels in the validation set. This histogram is truncated at CUI count 50
(x-axis) for clarity and uncluttered representation.</p>
      <p>On the combined training and validation set, there are 80,723 images and
907,718 non-unique CUIs. Among them, we counted 3,047 unique CUIs, which
were used to build the label space in our training objective (see Table 3).</p>
      <p>Following Tsoumakas et al. [18], we compute the Label Cardinality (LC) on
the combined training and validation set, denoted as D, using the formula:
LC(D) =
1 jDj</p>
      <p>X
jDj i=1</p>
      <p>jYij
LD(D) =
1 XjDj jYij</p>
      <p>L
jDj i=1 j j</p>
      <p>In a similar manner, we compute the Label Density (LD) using the following
formula:</p>
      <p>where jLj is the number of unique labels in our multi-label classification
objective. The LC and LD scores on the combined training and validation sets
are 11.24 and 0.0037 respectively.
(3)
(4)
As a first step in formatting the labels, we convert CUIs associated with images
to a format that is suitable as input for neural network learning. Since our
objective here is multi-label classification, we cannot use simple one-hot encoding,
as usually done in classification tasks, hence we apply a multi one-hot encoding.
An illustration of this representation can be found in Table 3. Specifically, we
sort the list of CUIs from the unique label set (k) in alphabetical order and use
the positions of CUIs in the sorted list to mark as 1 if a specific CUI is associated
with the image in question, else as 0. After this conversion step, the label set
for each image (I ) is represented as a single multi one-hot vector of fixed size k,
which is equal to 3,047.</p>
      <p>We used only the images from ImageCLEFmed Caption dataset and did not
use pre-training on external datasets, or utilize other modalities such as text
during model training.</p>
      <p>For our experiments, we divided the validation set via random sampling into
two equally sized subsets, namely val1 and val2. We conducted our internal
evaluations by always training pairs of models, first using val1 for validation
and val2 for testing, and then vice-versa. When results were promising, we then
submitted our predictions on the test set images by using the model trained with
first configuration. Training a third model, based on the validation on the full
development set would have been the ideal solution. However, this could not be
applied in our case because of time constraints.</p>
      <p>Histogram of CUI counts on the Train split
5201
4801
4401
4001
3601
seg3201
a
m
I
fo2801
r
e
bm2401
u
N
2001
1601
1201
801
401
1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50</p>
      <p>Count of Concept Unique Identifiers (CUIs)
The prediction problem for this challenge lays in the category of multi-label
classification [18]. It differs from most common classification problems in the
fact that each sample of the dataset is simultaneously associated with more
than one class from the ground truth label pool.</p>
      <p>Technically, when addressing such a problem with deep neural networks, it
means that the final classification layer relies on multiple sigmoidal units rather
than a single softmax probability distribution. The last layer of the network
contains one sigmoidal unit for each of the target classes, and the association
with a true/false result is performed by thresholding the final sigmoid activation
value (usually at 0.5).</p>
      <p>To address the challenge, we followed a classical transfer learning approach
starting from a Convolutional Neural Network (CNN) model pre-trained on an
image classification problem, namely ImageNet, because the pre-trained network
already offers the ability to detect basic image features, viz. edges, borders, and
corners. Then, the final classification stage of the network (i.e., all the layers
after the last convolutional layer) is substituted with randomly initialized fully
connected layers. Finally, the network is fitted for the new target training set.</p>
      <p>
        In detail, we used VGG16 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], ResNet50 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and DenseNet169 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], all of
them pre-trained on ImageNet data used for ILSVRC [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. An example
configu
      </p>
      <p>Histogram of CUI counts on the Validation split
1201
1101
1001
901
s 801
e
g
Iam701
f
o
r
eb601
m
u
N501
401
301
201
101
1 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50</p>
      <p>Count of Concept Unique Identifiers (CUIs)
ration based on VGG16 is shown in listing 1.1. Layers from block1_conv1 to
block5_pool are pre-trained on ImageNet, and unlocked for further training.
Remaining layers, from flatten to predictions, are newly instantiated and
initialized randomly (where applicable). The final predictions layer is a dense
layer with 3,047 sigmoidal activation units (one per target class).</p>
      <p>
        The system was developed in a Python environment using Keras deep
learning framework [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] with TensorFlow [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] as the backend. For our experiments we
used a desktop machine equipped with an 8-core 9th-gen i7 CPU, 64GB RAM,
and NVIDIA RTX TITAN 24GB GPU memory.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <p>In this section, we describe our experimental procedure, model configuration, and
a variety of deep CNN architectures that we tried for achieving the multi-label
classification task.</p>
      <p>
        Table 4 reports the results of the experiment we conducted throughout the
challenge. Because of time constraints, rather than running a grid search for the
best hyper-parameter values, we started from a reference configuration, already
successfully used in other past works [
        <xref ref-type="bibr" rid="ref10 ref9">9,10</xref>
        ].
Listing 1.1: An excerpt of the VGG16 architecture used for the multi-label
classification task.
      </p>
      <p>Together with the base architecture used for convolution (CNN arch) we
report: (res) the resolution in pixels of the input images of equal height and
width, (aug) the data augmentation strategy, (fc layers) the configuration of
final fully-connected stage of the CNN architectures, (do) the dropout value after
each fully connected layer, (bs) the batch size used for training, (loss func) the
loss function used for optimization, (lr-red) the learning rate reduction strategy
(reduction factor/patience/monitored metric). Additionally, we report the best
training epoch, based on an early stopping criteria by monitoring the F1 score on
the validation set. The last three columns report the F1 scores achieved on the
two internal cross-validation sets and finally on the AIcrowd3 online submission
platform.</p>
      <p>Other training parameters, common to all configurations are: NAdam
optimizer, learning rate = 1e-5, and schedule decay 0.9. All images were scaled to
the input resolution of the CNN using nearest filtering, without any cropping.</p>
      <p>In the following, we report on the evolution of our tests and obtained results.
3https://www.aicrowd.com/challenges/imageclef-2020-caption-conceptExperiments 1-4: VGG16 baseline We started with (1) a VGG16
architecture, pretrained on ImageNet, with the last two fully connected (FC) layers
configured with n=2048 nodes, each followed by a dropout layer with the dropout
probability p set to 0.5.</p>
      <p>(2) We observed an increase in performance by increasing the size of the FC
layers to 4096 nodes.</p>
      <p>From the first two experiments, it was evident how the loss value (based on
binary cross-entropy) could not be effectively used to monitor the validation. The
ground truth of each sample, a vector of size 3,047, contains on average about 11
concepts per image (see Section 3.2), and only a few images contain more than
50 concepts. Hence, the ground truth matrix is very sparse. As a consequence,
the loss function quickly stabilizes into a plateau, as does the accuracy, which
saturates to values above 0.9966 after the first epoch. Hence, to better handle
early stopping, we implemented a training-time computation of the F1 score.</p>
      <p>(3) A further improvement was observed by applying a 2X data augmentation
of the input dataset. Each image is provided to the training procedure both
as-it-is and flipped horizontally. At the same time, learning rate reduction was
applied by monitoring F1 scores on the validation set, rather than the loss values.
However, increasing the learning rate happens only after an overfitting occurs.
This configuration led to an online evaluation score of 0.363. Figure 6 shows the
evolution of loss values and F1 scores over epochs during model training.</p>
      <p>(4) We tried to improve the performance by increasing the size of input
images to 450x450 pixels, which forced a reduction of the batch size to 24. We
could not observe any significant improvement in the accuracy, suggesting that
higher image resolutions do not provide useful details for label selection in our
case.</p>
      <p>Experiments 5-7: more powerful CNN architectures By using deeper
CNN architectures, we could observe a slight improvement in the test accuracy.
Indeed, the ResNet50 architecture (5) led to an F1 score of 0.365 in the online
-5
0
evaluation. The DenseNet169 architecture (6-7) led to higher test values, but
the online evaluation was slightly lower (0.360) than the ResNet50 version.
Experiment 8: more layers In order to increase the overall performance,
we tried to increase the number of FC layers to 3x4k (8). However, taking as
reference the performance of configuration (3), we could not observe a significant
improvement by introducing an additional 4k FC layer to the classification stage.
Experiments 9-10: a new loss function To further improve performance,
we decided to directly optimize for the F1 score evaluation metric. Notice that
the F1 score used in the ImageCLEF challenge is computed as an average F1
over the samples (and not over the labels, as more often found in online code
repositories4,5).</p>
      <p>We implemented a loss function F 1 = 1 sF 1, where sF 1 is called the “soft
F1 score”. The sF 1 is a differentiable version of the F1 function that computes
true positives, false positives, and false negatives as continuous sum of likelihood
values, without applying any thresholding to round the probabilities to 0 or 1.
The implementation of F 1 is shown in listing 1.2.</p>
      <p>Experiments using the F 1 loss function could not converge. Likely, the
problem is due to the fact that the F 1 loss lies in the range [0; 1]. As such, the
gradient search space can be abstractly seen as a huge plateau, just below 1.0,
with a solitary hole in the middle that quickly converges to the global minimum.
At the same level of abstraction, we can visualize the binary cross-entropy bce
search space as a wide bag, with a large flat surface, just above 0. It is easy to
reach the bottom of the bag, i.e., reach a very high binary accuracy due to the
sparsity of the labeling, but then we see a very mild slope towards the global
4https://towardsdatascience.com/the-unknown-benefits-of-using-a-softf1-loss-in-classification-systems-753902c0105d</p>
      <p>5https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-scoremetric
# The following is not differentiable .
# Round the prediction to 0 or 1 (0.5 threshold )
# y_pred = K. round ( y_pred )
# By commenting , we implement what is called soft - F1 .
# Compute F1 and return the loss .
f1 = 2 * p * r / (p + r + K. epsilon () )
f1 = tf . where ( tf . is_nan ( f1 ) , tf . zeros_like ( f1 ) , f1 )
return 1 - K. mean ( f1 )
Listing 1.2: Python implementation of the soft-F1 based loss function for the
Keras environment with TensorFlow backend.
1 def loss_1mf1_by_bce ( y_true , y_pred ):
2 import keras . backend as K
3
4
5
loss_f1 = loss_1_minus_f1 ( y_true , y_pred )
bce = K. binary_crossentropy ( target = y_true , output = y_pred , from_logits =
False )
Listing 1.3: Python implementation of the loss function combining soft-F1 score
with binary cross-entropy.
minimum at its center.</p>
      <p>Our intuition is that by combining (multiplying or adding) F 1 and bce results
in a search space where F 1 does not affect the identification of inside of the
bag, and at the same time helps with the identification of the F 1’s and bce’s
common global minimum. The implementation of the F 1 bce loss function is
straightforward and is presented in listing 1.3.</p>
      <p>(9) An experiment using VGG16 confirms that the loss function F 1
bce
leads to better results with 0.3604/0.3606 on our tests. This is the configuration
that performed best in the online submission (0.374).</p>
      <p>Further experiments, e.g., using ResNet50, could not be submitted to the
challenge due to time constraints. However, (10) an internal test using VGG16
in combination with the F 1 + bce loss function, led to the best performance in
our internal evaluation (0.3636/0.3632).
5</p>
    </sec>
    <sec id="sec-5">
      <title>Results and Discussion</title>
      <p>The results of our experiments for the ImageCLEFmedical 2020 challenge can
be summarized as follows.</p>
      <p>The task of concept detection can be modeled as a multi-labeling problem
and solved by a transfer learning approach where deep CNNs pretrained on
real-world images can be fine-tuned on the target dataset. The multi-labeling is
technically addressed by using a sigmoid activation function on the output layer
and a label selection by thresholding. A good configuration consists of a VGG16
deep CNN architecture followed by two fully connected layers of 4096 nodes,
each followed by a dropout layer with probability p set to 0.5. Augmenting
the training set with horizontally flipped images increases accuracy and also
reduces the number of epochs needed for training. Increasing the resolution of
input images does not prove to be useful, while better results are achieved by
substituting the convolution stage with a deeper CNN architecture (ResNet50).
We noticed that the learning rate reduction has never helped in improving the
results.</p>
      <p>Using the standard binary cross-entropy loss function leads to competitive
results, which significantly increases when it is combined with a soft-F1 score
computation. It is worth noticing that when using solely the soft-F1 score as a
loss function, the network could not converge and this problem needs further
investigation.</p>
      <p>In total, we made five online submissions to the challenge. Table 5 presents
the F1 scores achieved on the withheld test set and the overall ranking of our
team (iml ) out of 47 successful submissions, as reported by the challenge
organizers. In addition, it is worth mentioning that the difference in F1 scores
between our best submission and the system that achieved the highest score in
the challenge is 0.0195. What percentage of test set images on which our model
still needs to achieve correct labels to bridge this gap needs further investigation.</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Directions</title>
      <p>In this work we have proposed a deep convolutional neural network based
approach for concept detection in radiology images. Our best performance (12th
position) is achieved by implementing a new loss function whereby we combined
the widely used binary cross-entropy loss together with a differentiable version
of the F1 score evaluation metric.</p>
      <p>
        Still several aspects could be investigated to improve the achieved results. For
instance, as we can observe from the CUI distribution plots, there is an imbalance
in the dataset. Consequently the model is biased towards predicting the concepts
associated with over-represented samples. Our future work will focus on the
approaches to combat such type of biases. A straightforward approach would
be to undersample the over-represented samples using a query strategy that
maximizes informativeness of chosen samples. Such a strategy showed promising
results for incremental domain adaptation task in neural machine translation [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Furthermore, we did not make use of the existing categorization of the images
in 7 sub-sets. A straightforward idea would be to train 7 different models, one
per category, and rely on their ensemble for a final global classification result.
Alternatively, the class identifier might be used as additional metadata
information, concatenated to the images’ internal features representation in the CNN,
and fed to a further shallow neural network for improved classification (see [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]).
Another promising direction to look would be to consider further trends in the
integration of vision and language research [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
17. Daniel Sonntag, Volker Tresp, Sonja Zillner, Alexander Cavallaro, Matthias
Hammon, André Reis, Peter A. Fasching, Martin Sedlmayr, Thomas Ganslandt,
HansUlrich Prokosch, Klemens Budde, Danilo Schmidt, Carl Hinrichs, Thomas
Wittenberg, Philipp Daumke, and Patricia G. Oppelt. The clinical data intelligence
project - A smart data initiative. Informatik Spektrum, 39(4):290–300, 2016.
18. Grigorios Tsoumakas and Ioannis Katakis. Multi-Label Classification: An
Overview. International Journal of Data Warehousing and Mining, 3(3):1–13, July
2007.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. Keras Special Interest Group. Keras. https://keras.io/.</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kaiming</surname>
            <given-names>He</given-names>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In 2016 IEEE Conference on Computer Vision</source>
          and Pattern Recognition,
          <string-name>
            <surname>CVPR</surname>
          </string-name>
          <year>2016</year>
          ,
          <string-name>
            <surname>Las</surname>
            <given-names>Vegas</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NV</surname>
          </string-name>
          , USA, June 27-30,
          <year>2016</year>
          , pages
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          . IEEE Computer Society,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Gao</surname>
            <given-names>Huang</given-names>
          </string-name>
          , Zhuang Liu, Laurens van der Maaten, and
          <string-name>
            <surname>Kilian</surname>
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Weinberger</surname>
          </string-name>
          .
          <article-title>Densely connected convolutional networks</article-title>
          .
          <source>In 2017 IEEE Conference on Computer Vision</source>
          and Pattern Recognition,
          <string-name>
            <surname>CVPR</surname>
          </string-name>
          <year>2017</year>
          ,
          <article-title>Honolulu</article-title>
          ,
          <string-name>
            <surname>HI</surname>
          </string-name>
          , USA, July
          <volume>21</volume>
          -
          <issue>26</issue>
          ,
          <year>2017</year>
          , pages
          <fpage>2261</fpage>
          -
          <lpage>2269</lpage>
          . IEEE Computer Society,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Bogdan</given-names>
            <surname>Ionescu</surname>
          </string-name>
          , Henning Müller, Renaud Péteri, Asma Ben Abacha, Vivek Datla, Sadid A.
          <string-name>
            <surname>Hasan</surname>
          </string-name>
          , Dina Demner-Fushman, Serge Kozlovski, Vitali Liauchuk, Yashin Dicente Cid, Vassili Kovalev, Obioma Pelka,
          <string-name>
            <surname>Christoph M. Friedrich</surname>
          </string-name>
          , Alba García Seco de Herrera,
          <string-name>
            <surname>Van-Tu</surname>
            <given-names>Ninh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tu-Khiem</surname>
            <given-names>Le</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liting Zhou</surname>
          </string-name>
          , Luca Piras, Michael Riegler, Pål Halvorsen,
          <string-name>
            <surname>Minh-Triet</surname>
            <given-names>Tran</given-names>
          </string-name>
          , Mathias Lux, Cathal Gurrin,
          <string-name>
            <surname>Duc-Tien</surname>
          </string-name>
          Dang-Nguyen, Jon Chamberlain, Adrian Clark, Antonio Campello, Dimitri Fichou, Raul Berari, Paul Brie, Mihai Dogariu, Liviu Daniel Ştefan, and Mihai Gabriel Constantin.
          <source>Overview of the ImageCLEF</source>
          <year>2020</year>
          :
          <article-title>Multimedia retrieval in medical, lifelogging, nature, and internet applications</article-title>
          .
          <source>In Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , volume
          <volume>12260</volume>
          <source>of Proceedings of the 11th International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ), Thessaloniki, Greece,
          <source>September 22-25 2020. LNCS Lecture Notes in Computer Science</source>
          , Springer.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Marimuthu</given-names>
            <surname>Kalimuthu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Barz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Sonntag</surname>
          </string-name>
          .
          <article-title>Incremental domain adaptation for neural machine translation in low-resource settings</article-title>
          . In Wassim ElHajj, Lamia Hadrich Belguith, Fethi Bougares, Walid Magdy, and Imed Zitouni, editors,
          <source>Proceedings of the Fourth Arabic Natural Language Processing Workshop</source>
          , WANLP@ACL 2019, Florence, Italy,
          <source>August</source>
          <volume>1</volume>
          ,
          <year>2019</year>
          , pages
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . Association for Computational Linguistics,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Alex</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , Ilya Sutskever, and Geoffrey E Hinton.
          <article-title>ImageNet Classification with Deep Convolutional Neural Networks</article-title>
          . In F. Pereira,
          <string-name>
            <given-names>C. J. C.</given-names>
            <surname>Burges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bottou</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          K. Q. Weinberger, editors,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>25</volume>
          , pages
          <fpage>1097</fpage>
          -
          <lpage>1105</lpage>
          . Curran Associates, Inc.,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Martín</given-names>
            <surname>Abadi</surname>
          </string-name>
          , Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis,
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Matthieu</given-names>
            <surname>Devin</surname>
          </string-name>
          , Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke,
          <string-name>
            <given-names>Yuan</given-names>
            <surname>Yu</surname>
          </string-name>
          , and Xiaoqiang Zheng.
          <source>TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Mogadala</surname>
          </string-name>
          , Marimuthu Kalimuthu, and
          <string-name>
            <given-names>Dietrich</given-names>
            <surname>Klakow</surname>
          </string-name>
          .
          <article-title>Trends in integration of vision and language research: A survey of tasks, datasets, and methods</article-title>
          . CoRR, abs/
          <year>1907</year>
          .09358,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Fabrizio</given-names>
            <surname>Nunnari</surname>
          </string-name>
          , Chirag Bhuvaneshwara, Abraham Obinwanne Ezema, and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Sonntag</surname>
          </string-name>
          .
          <article-title>A study on the fusion of pixels and patient metadata in CNN-based classification of skin lesion images</article-title>
          .
          <source>In International IFIP Cross Domain Conference for Machine Learning and Knowledge Extraction (CD-MAKE)</source>
          . Springer,
          <year>August 2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. Fabrizio Nunnari and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Sonntag</surname>
          </string-name>
          .
          <article-title>A CNN toolbox for skin cancer classification</article-title>
          . arXiv:
          <year>1908</year>
          .08187 [cs, eess],
          <year>August 2019</year>
          . arXiv:
          <year>1908</year>
          .08187.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Obioma</surname>
            <given-names>Pelka</given-names>
          </string-name>
          , Christoph M Friedrich, Alba García Seco de Herrera, and
          <string-name>
            <given-names>Henning</given-names>
            <surname>Müller</surname>
          </string-name>
          .
          <article-title>Overview of the ImageCLEFmed 2020 concept prediction task: Medical image understanding</article-title>
          .
          <source>In CLEF2020 Working Notes, CEUR Workshop Proceedings</source>
          , Thessaloniki, Greece,
          <source>September</source>
          <volume>22</volume>
          -25
          <year>2020</year>
          .
          <article-title>CEUR-WS.org</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Obioma</surname>
            <given-names>Pelka</given-names>
          </string-name>
          , Sven Koitka, Johannes Rückert, Felix Nensa, and
          <string-name>
            <surname>Christoph M Friedrich.</surname>
          </string-name>
          <article-title>Radiology objects in context (roco): A multimodal image dataset</article-title>
          .
          <source>In Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis</source>
          , pages
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          . Springer,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>Karen</given-names>
            <surname>Simonyan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          .
          <source>In Yoshua Bengio and Yann LeCun</source>
          , editors,
          <source>3rd International Conference on Learning Representations, ICLR</source>
          <year>2015</year>
          , San Diego, CA, USA, May 7-
          <issue>9</issue>
          ,
          <year>2015</year>
          , Conference Track Proceedings,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. Daniel Sonntag.
          <source>AI in Medicine, Covid-19 and Springer Nature's Open Access Agreement. KI</source>
          ,
          <volume>34</volume>
          (
          <issue>2</issue>
          ):
          <fpage>123</fpage>
          -
          <lpage>125</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. Daniel Sonntag, Fabrizio Nunnari, and
          <string-name>
            <surname>Hans-Jürgen Profitlich</surname>
          </string-name>
          .
          <article-title>The Skincare project, an interactive deep learning system for differential diagnosis of malignant skin lesions</article-title>
          .
          <source>Technical Report</source>
          . CoRR, abs/
          <year>2005</year>
          .09448,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. Daniel Sonntag and
          <string-name>
            <surname>Hans-Jürgen Profitlich</surname>
          </string-name>
          .
          <article-title>An architecture of open-source tools to combine textual information extraction, faceted search and information visualisation</article-title>
          .
          <source>Artificial Intelligence in Medicine</source>
          ,
          <volume>93</volume>
          :
          <fpage>13</fpage>
          -
          <lpage>28</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>