<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AUEB NLP Group at ImageCLEFmed Caption 2020</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Basil Karatzas</string-name>
          <email>karatzas.basil@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John Pavlopoulos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasiliki Kougia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ion Androutsopoulos</string-name>
          <email>iong@aueb.gr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer and Systems Sciences, Stockholm University</institution>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Informatics, Athens University of Economics and Business</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article concerns the participation of AUEB's NLP Group in the ImageCLEFmed Caption task of 2020. The goal of the task was to identify medical terms that best describe each image, in order to accelerate and improve the interpretation of medical images by experts and systems. The systems we implemented extend our previous work [7,8,9] on models that employ CNN image encoders combined with an image retrieval method or a feed-forward neural network. Our systems were ranked 1st, 2nd and 6th.</p>
      </abstract>
      <kwd-group>
        <kwd>Medical Images</kwd>
        <kwd>Concept Detection</kwd>
        <kwd>Image Retrieval</kwd>
        <kwd>Image Captioning</kwd>
        <kwd>Multi-label Classification</kwd>
        <kwd>Multimodal</kwd>
        <kwd>Ensemble</kwd>
        <kwd>Convolutional Neural Network (CNN)</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Deep Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        ImageCLEF [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is an evaluation campaign held annually since 2003 as part of CLEF3,
and revolves around image analysis and retrieval tasks. ImageCLEFmedical [11] is a
collection of ImageCLEF tasks that are associated with the study of medical images.
In 2020, it consisted of 3 tasks: VQA-Med, Caption and Tuberculosis.4 The
ImageCLEFmed Caption task concerns the automatic assignment of medical terms (called
concepts) to medical images. The dataset of ImageCLEFmed Caption 2020 consisted
of medical images, which were split to 7 categories according to their radiology
modality (see Table 1). Writing a diagnostic report for a medical image is a demanding and
very time-consuming task that needs to be handled by medical experts [
        <xref ref-type="bibr" rid="ref2">2,14</xref>
        ]. One of
the main goals of the ImageCLEFmed Caption task is to assist the development of
efficient, multi-label, medical-image tagging models, which could be used to assist the
medical experts and reduce the time needed for the diagnosis as well as to reduce
potential medical errors.
      </p>
      <p>
        In this paper we describe the medical image tagging systems of the AUEB NLP
Group that were submitted to ImageCLEFmed Caption 2020. Following our last year’s
success [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], our 3 submissions were ranked 1st, 2nd and 6th.5
      </p>
      <p>
        Overall, our submissions were based on two methods. The first method was based
on the Mean@k-NN system of [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] that assigns concepts to each test image using the
k nearest neighbors from the training dataset. The second method extends the
ConceptCXN system of [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and uses a DenseNet-121 CNN to encode any test image and a
Feed Forward Neural Network classifier on top. The remaining of the paper describes
the data, methods, submitted systems and our results, followed by conclusions and
future directions.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>
        Initially, the ImageCLEFmed Caption datasets comprised a broad variety of clinical
images, which were extracted from figures of scientific articles found in the open-access
biomedical literature database PubMed Central.6 Each image was assigned medical
terms from the Unified Medical Language System (UMLS) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These terms, called
concepts, were extracted from the processed text of the respective figure caption. Since
2019, in order to discard compound or non-radiology images from the initial datasets,
the organisers applied filters and also performed a manual revision of their data. A
subset of the resulting dataset, which is called the extended Radiology Objects in COntext
(ROCO) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], was chosen to be used as the dataset of the competition this year (see Fig.
1). Additionally, this year, images were classified into 7 mutually exclusive categories,
as shown in Table 1, depending on the type of the radiology exam.
      </p>
      <p>
        The number of possible concepts was reduced compared to previous years, by
removing concepts with few occurrences, since the large number of concepts in the
previous years resulted in the task being difficult for models [15]. There were 111,156
5Our best performing system will become available in the bioCaption PyPi package.
6https://www.ncbi.nlm.nih.gov/pmc/
possible concepts in 2018 [16], 5,528 in 2019 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and 3,047 in 2020. We also observed
that there were concepts that appeared in every single image of a specific category, but
rarely or never appeared in other categories. The Concept Unique Identifiers (CUIs) of
these concepts are shown in Table 1.7 These concepts rather describe the modality of the
respective category. For example, DRPE images are always assigned with C0032743,
whose UMLS term is POSITRON-EMISSION TOMOGRAPHY.
      </p>
      <p>The dataset was split by the organisers to a training set of 64,753 images, a
validation set of 15,970 images, and a test set of 3,534 images. For our experiments, we
merged the provided training and validation sets and used 10% of the merged data as
a development set. We will refer to the remaining 90% of the merged dataset as the
training set for the rest of the paper.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <p>This section describes the systems that were used in our submissions.
3.1</p>
      <p>
        System 1: 2xCNN+FFNN
CNN+FNNN (a.k.a. ConceptCXN or DenseNet121+FFNN) [
        <xref ref-type="bibr" rid="ref8 ref9">9,8</xref>
        ] is the system that we
submitted last year (and was ranked 1st) for the same task. It is a variation of CheXNet
[13] that uses DenseNet-121 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which is a stack of 120 CNN layers, followed by
7We used UMLS Metathesaurus (uts.nlm.nih.gov/home.html) to map each CUI to
its term.
a feed-forward Neural Network (FFNN) that acts as a classifier layer on top. In the
ImageCLEFmed Caption task of 2019, we changed the original FFNN to comprise
5,528 outputs (instead of 14), one per available concept. For the task of 2020, which also
comprises different image categories (see Table 1), we followed the same approach per
model (i.e., we employed one model per category). For example, the respective FFNN
for the model of category C generates NC outputs, which is the number of all possible
concepts in category C.8 The red bars of Fig. 2 depict the NC values per category C,
computed on the training set.
      </p>
      <p>
        We trained the model by minimizing the binary cross entropy loss. We used Adam
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] as our optimizer and decreased the learning rate by a factor of 10 when the loss
showed no improvement, following the work of [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. We used a batch size of 16 and
early stopping with a patience of 3 epochs. For each CNN+FFNN of a specific category
(i.e., for each category, we fine-tuned a CNN and an FFNN on top), a classification
threshold for all the concepts of the respective category was tuned by optimising the
F1 score. Any concepts for which the respective output values exceeded that threshold
were assigned to the corresponding image.
      </p>
      <p>Two of our 2020 submissions consisted of ensembles of CNN+FFNN models,
constructed in the following way. We trained 5 models per category, and kept the 2 best
performing ones, according to their F1 score. We then created two ensembles, using the
UNION and the INTERSECTION of the concepts returned by these two models.
Hereafter, these two ensembles will be called 2xCNN+FFNN@U and 2xCNN+FFNN@I,
respectively.</p>
      <p>
        8We did not use image augmentation this year due to time restrictions, because of the large
number of models we needed to train.
Following our previous work [
        <xref ref-type="bibr" rid="ref7 ref8">8,7</xref>
        ], the goal of our CNN+k-NN model for each test
image was to retrieve similar images from the training set. The encoder of this model
stemmed from our fine-tuned CNN+FFNN system, hence it is a CNN per category. We
employed the output of the last average pooling layer of the CNN to represent each
encoded image.9 The encoded test image was then compared to all the images in the
training set (encoded offline), using cosine similarity, and the k nearest images were
returned.
      </p>
      <p>After retrieving the k nearest images, CNN+k-NN returned the r concepts that were
most frequently assigned to the k images. We tuned k and r for each category,
investigating values from 1 to 200 for k, and values from 1 to 10 for r. During tuning, we also
considered two other functions for r. First, we used the average number of concepts in
the k images:
k
r = 1 X ni
k
i=1
(1)
9Each image is rescaled to 224x224 and normalised with the mean and standard deviation of
ImageNet.
Second, we used a weighting based on cosine similarity to weigh the concepts:
k
r = X
i=1</p>
      <p>cos(g; gi)
Pk
j=1 cos(g; gj ))
ni
(2)
where ni is the number of concepts of the i-th retrieved image, g is the test image, gi is
the i-th closest training image, and cos(g; gi) the cosine similarity between g and gi.</p>
      <p>
        As with CNN+FFNN, our DenseNet-121 CNN was pretrained on ImageNet and
fine-tuned on the ImageCLEFmed Caption dataset. However, we experimented also
with adding an attention layer10 [12] to our CNN. We call this model CNN+kNN@att.
We also experimented with fine-tuning our CNN on a large dataset of radiography
images called MIMIC-CXR [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (before fine-tuning it further on the ImageCLEFmed
Caption dataset), but we did not obtain any improvements.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Submissions and Results</title>
      <p>In order to decide what models to use for the final submissions, we evaluated all
models on our development set. Since images were separated into seven categories, each
submission consisted of a model per category, resulting in seven models per
submission. Two out of our three submissions employed different instances of the same
system (2xCNN+FFNN@U &amp; 2xCNN+FFNN@I), each for a different category, while the
third one (called BEST@CATEGORY) combined results from different types of
systems for each category.</p>
      <p>The official measure of the competition was F1, macro-averaged over the images,
without taking into account the different categories. To generate the predictions for the
test set, we merged the training with the development set. We used a held-out set (20%
of the merged data) to tune the hyper-parameters of the CNN+FFNN and CNN+k-NN
models (see Table 4 and Table 4 for the final values). As shown in Fig. 5, the best score
improves every year. This is probably because the task was indeed simplified each year
(see Section 2).</p>
      <p>Table 4 presents the scores of our systems on the development and the official test
set, along with the official rankings. 2xCNN+FFNN@I was the best. On the other hand,
2xCNN+FFNN@U, which returns the union (instead of the intersection) of the
predicted concepts of the models in the ensemble, was ranked much lower. It is worth
mentioning that a baseline, which simply returns the concepts always shown per
category, achieves very high F1 on the development set. The submission that combined the
best model per category (see Table 4) was ranked 2nd.</p>
      <p>We noticed that models tended to predict only the concepts that always appear in
each category, thus we also show statistics regarding the diversity of our submissions
for the test set (Fig. 6). We define diversity as the total number of distinct concepts the
models predicted for a specific category divided by the total number of concepts found
in the training set of that category. Observing the output of our models, we noticed that
the ones with very low diversities only predicted the concepts that appeared in every
image (as shown in Table 1) of the category they were trained on.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>
        This article described the submissions of AUEB’s NLP Group to the 2020
ImageCLEFmed Caption task. One of our submissions was ranked 1st, while the other two
were ranked 2nd and 6th. All of our systems were based on a DenseNet-121 CNN [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
to encode images. A retrieval-based method achieved the best results in three out of
seven categories. However, an ensemble of two CNN+FFNN multi-label classifiers was
ranked 1st overall. Future work includes the assessment of our models on more datasets
and improving retrieval-based methods, which are still under-explored.
11. Pelka, O., Friedrich, C.M., Garc´ıa Seco de Herrera, A., Mu¨ller, H.: Overview of the
ImageCLEFmed 2020 concept prediction task: Medical image understanding. In: CLEF2020
Working Notes. CEUR Workshop Proceedings, CEUR-WS.org $$, Thessaloniki, Greece
(September 22-25 2020)
12. Raffel, C., Ellis, D.P.W.: Feed-forward networks with attention can solve some long-term
memory problems (2015)
13. Rajpurkar, P., Irvin, J., Zhu, K., Yang, B., Mehta, H., et al.: CheXNet: Radiologist-Level
      </p>
      <p>Pneumonia Detection on Chest X-rays with Deep Learning. arXiv:1711.05225 (2017)
14. Rimmer, A.: Radiologist shortage leaves patient care at risk, warns royal college. British</p>
      <p>Medical Journal 359 (2017)
15. S. Singh, S. Karimi, K.H.S., Hamey, L.: Biomedical Concept Detection in Medical Images:
MQ-CSIRO at 2019 ImageCLEFmed Caption Task. In: CLEF2019 Working Notes. CEUR
Workshop Proceedings. p. 15. Lugano, Switzerland (2019)
16. Zhang, Y., Wang, X., Guo, Z., Li, J.: ImageSem at ImageCLEF 2018 Caption Task: Image
Retrieval and Transfer Learning. In: CLEF2018 Working Notes. CEUR Workshop
Proceedings. Avignon, France (2018)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The unified medical language system (umls): integrating biomedical terminology</article-title>
          .
          <source>Nucleic acids research 32(Database issue)</source>
          ,
          <source>4 (Jan</source>
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Chokshi</surname>
            ,
            <given-names>F.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hughes</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mullins</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hawkins</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jr</surname>
          </string-name>
          , R.D.:
          <article-title>Diagnostic radiology resident and fellow workloads: a 12-year longitudinal trend analysis using national medicare aggregate claims data</article-title>
          .
          <source>Journal of the American College of Radiology</source>
          <volume>12</volume>
          (
          <issue>7</issue>
          ),
          <fpage>664</fpage>
          -
          <lpage>669</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>van der Maaten</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>Weinberger</surname>
            ,
            <given-names>K.Q.</given-names>
          </string-name>
          :
          <article-title>Densely Connected Convolutional Networks</article-title>
          .
          <source>In: Proceedings of the IEEE conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <fpage>4700</fpage>
          -
          <lpage>4708</lpage>
          . Honolulu,
          <string-name>
            <surname>HI</surname>
          </string-name>
          , USA (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Mu¨ller, H., Pe´teri, R.,
          <string-name>
            <surname>Abacha</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Datla</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kozlovski</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liauchuk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cid</surname>
            ,
            <given-names>Y.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelka</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Herrera</surname>
            ,
            <given-names>A.G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ninh</surname>
            ,
            <given-names>V.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , l Halvorsen,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.T.</given-names>
            ,
            <surname>Lux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gurrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Dang-Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.T.</given-names>
            ,
            <surname>Chamberlain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Campello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Fichou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Berari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Brie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Dogariu</surname>
          </string-name>
          , M.,
          <string-name>
            <given-names>S</given-names>
            ¸ tefan, L.D.,
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.G.</surname>
          </string-name>
          :
          <article-title>Overview of the ImageCLEF 2020: Multimedia retrieval in medical, lifelogging, nature, and internet applications</article-title>
          .
          <source>In: Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the 11th International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ), vol.
          <volume>12260</volume>
          .
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Thessaloniki,
          <source>Greece (September</source>
          <volume>22</volume>
          -25
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>A.E.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pollard</surname>
            ,
            <given-names>T.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greenbaum</surname>
            ,
            <given-names>N.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lungren</surname>
            ,
            <given-names>M.P.</given-names>
          </string-name>
          , ying Deng,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Mark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.G.</given-names>
            ,
            <surname>Berkowitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.J.</given-names>
            ,
            <surname>Horng</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs (</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
          </string-name>
          , J.:
          <article-title>Adam: A Method for Stochastic Optimization</article-title>
          .
          <source>arXiv:1412.6980</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kougia</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavlopoulos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Androutsopoulos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>A Survey on Biomedical Image Captioning</article-title>
          . In: Workshop on Shortcomings in
          <article-title>Vision and Language of the Annual Conference of the North American Chapter of the Association for Computational Linguistics</article-title>
          . pp.
          <fpage>26</fpage>
          -
          <lpage>36</lpage>
          . Minneapolis, MN, USA (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kougia</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavlopoulos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Androutsopoulos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>AUEB NLP Group at ImageCLEFmed Caption 2019</article-title>
          .
          <source>In: CLEF2019 Working Notes. CEUR Workshop Proceedings</source>
          . pp.
          <fpage>9</fpage>
          -
          <lpage>12</lpage>
          . Lugano,
          <string-name>
            <surname>Switzerland</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kougia</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavlopoulos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Androutsopoulos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Medical Image Tagging by Deep Learning and Retrieval</article-title>
          . In: Experimental IR Meets Multilinguality, Multimodality, and
          <source>Interaction Proceedings of the Eleventh International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ). Thessaloniki,
          <string-name>
            <surname>Greece</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Pelka</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koitka</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Ru¨ckert, J.,
          <string-name>
            <surname>Nensa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          :
          <article-title>Radiology Objects in COntext (ROCO): A Multimodal Image Dataset</article-title>
          .
          <source>In: MICCAI Workshop on Large-scale Annotation of Biomedical data and Expert Label Synthesis</source>
          . pp.
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          . Granada,
          <string-name>
            <surname>Spain</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>