<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Attention-based CNN-GRU Model For Automatic Medical Images Captioning: ImageCLEF 2021</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Djamila-Romaissa</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Beddiar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mourad</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oussalah</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tapio</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Seppänen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for machine vision and signal analysis, University of Oulu</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>MIPT, Faculty of Medicine, University of Oulu</institution>
          ,
          <addr-line>Oulu</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>2</volume>
      <fpage>1</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>The action of understanding and interpretation of medical images is a very important task in the medical diagnosis generation. However, manual description of medical content is a major bottleneck in clinical diagnosis. Many research studies were devoted to develop automated alternatives to this process, which would have enormous impact in terms of eficiency, cost and accuracy in the clinical workflows. Diferent approaches and techniques have been presented in the literature ranging from traditional machine learning methods to deep learning based models. Inspired by the outperforming results of the later techniques, we present in the current paper, our team participation (RomiBed) to the ImageCLEF medical caption prediction task. We addressed the challenge of medical image captioning by combining a CNN encoder model with an attention-based GRU language generator model whereas a multi-label CNN classifier is used for the concept detection task. Using the provided data in the training, validation and test subsets, we obtain an average F_measure of 14.3% and a BLEU score of 0.243 on the ImageCLEF concept detection and the caption prediction challenges, respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Automatic image captioning</kwd>
        <kwd>Medical images</kwd>
        <kwd>Concept detection</kwd>
        <kwd>Radiology</kwd>
        <kwd>Multi-label Classification</kwd>
        <kwd>Encoder-decoder</kwd>
        <kwd>Attention Mechanism</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        With the increasing number of medical images generated worldwide from diferent
modalities in hospitals and health centers, the need to analyse and discover their content is crucial.
Indeed, medical images ofer a safe environment to explore patient’s health state without the
need for a surgery or any other invasive procedures [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Besides, this also helps clinicians in
their daily routine by expediting clinical workflows and trigger automated alerts associated
to potentially dangerous diseases. Recently, many research was devoted to the process of
automatically generating clinically sound interpretations of medical images. Roughly speaking,
generating clinically explainable and understandable analysis for medical images may enrich
medical knowledge systems and facilitate the human-machine interactive diagnosis practice [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Therefore, automatic medical image captioning is one of the main focus of the interdisciplinary
research in medical imaging field [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Especially, medical image captioning uses visual features
of images to generate a concise textual description of the content of the medical image by
highlighting the clinically important observations. It represents a convergence of computer
vision and natural language processing (NLP) with an emphasis on medical image processing
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this regard, the ImageCLEFmedical task [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ] is organized each year as part of the CLEF
initiative labs aiming at developing machine learning methods for medical image understanding
and description. It includes two sub-tasks: the concept detection sub-task which aims to identify
the UMLS Concept Unique Identifiers (CUIs) for a given medical image. Whereas the caption
prediction sub-task aims to generate coherent caption based on the clinical concept vocabulary
created in the first sub-task and the visual content of the image.
      </p>
      <p>Motivated by the recent advances made in deep neural networks in diferent tasks of computer
vision and NLP, especially due to their promising results in the machine language translation
models, we present in this paper our contribution to the ImageCLEF 2021 medical task under
the team name ’RomiBed’. We proposed a multi-label classification CNN model for the first
sub-task after applying an augmentation technique based on the center cropping of the medical
images. Features were extracted using a pre-trained model, while the classification is performed
using a CNN network. For the second sub-task, we proposed an encoder-decoder model with
an attention layer where the encoder is based on a CNN feature extractor and the decoder is
composed of a GRU network with an attention mechanism.</p>
      <p>This paper is organized as follows. First, we briefly review the related medical image
captioning studies from the literature in Section. 2. In Section. 3, we provide a brief description
of the ImageCLEF dataset used in this study. Next, we detail the methodology we followed to
construct the concept detection model as well as the caption prediction model in Section. 4. We
discuss each step of the process and deliver the results in terms of F_measure for the concept
detection and BLEU for the caption prediction. Finally, we finish with a conclusion where we
highlight some key insights and future directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Automatic image captioning (AIC) in the medical field has gained a particular attention from
researchers due to its importance and its huge impact on health care centers by allowing
instantaneous understanding of medical images for doctors as well as patients. In addition,
the significant progress made to date in artificial intelligence due to deep learning models
contributed greatly to the AIC task [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Therefore, diferent techniques ranging from traditional
template-based and/or retrieval-based systems to generative models based on deep-neural
networks passing through various hybrid models that combine diferent techniques [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] emerged.
During the last years, many systems have been proposed to compete for the ImageCLEF
medical challenge. For the first step towards medical image captioning, which consists of
concept detection in ImageCLEF medical task, the multi-label classifications is found to play
a leading role. For instance, [
        <xref ref-type="bibr" rid="ref2 ref7">2, 7</xref>
        ] exploited the transfer learning to perform a multi-label
classification by extracting significant features from medical images using pre-trained models
such as the Resnet50, InceptionV3 . . . . In addition, Wang et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] explored a retrieval-based
topic modelling method to extract the most relevant clinical concepts from images similar to
the input image. Encoder-decoder (CNN-RNN) architectures were explored by many studies to
generate appropriate captions. In many cases, attention-mechanism is added to the baseline
encoder-decoder model as in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] who contributed to the ImageCLEF 2017 edition. Similarly,
Hasan et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] enriched the soft attention-based encoder-decoder model by inputting, to the
decoder, the output of the classification model on image modalities. This allowed them to
supplement the decoder with more fine grained details on the data to make the generation
process more focused. Indeed, supplementary information such as image modalities or body
parts could enhance the classification model as reported in Lyndon et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Furthermore,
the RNN decoder [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is replaced by diferent variants such as LSTM [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] Xu et al.[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], or GRU
Ambati and Dudyala [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] who used the captioning module to resolve the task of visual question
answering. Likewise, Benzarti et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] employed the captioning model to medical retrieval
systems in order to obtain the query terms. From another perspective, Rahman [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] proposed to
extract textual and visual information using RNN-based and CNN-based networks respectively,
and then merge the outputs of both models to generate relevant captions. Similarly, Mishra
and Banerjee [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] adopted the same technique aiming to detect retinal diseases and to generate
appropriate medical reports. In other techniques, generative models are combined with retrieval
systems for AIC. For example, Kougia et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] proposed to exploit the image visual features to
retrieve similar images with their known concepts and then combine them to predict enhanced
captions of the input image. The prediction is performed using an encoder-decoder generative
model. Likewise, [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ] suggested to use a retrieval policy module that makes a choice between
generating new sentence or retrieving a template sentence.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Data analysis</title>
      <p>The data used for this year edition is shared between both the ImageCLEFmed Caption and
the ImageCLEF-VQAMed tasks. The dataset include three sets: the training set composed of
the VQA-Med 2020 training data with 2756 medical images; the validation set consisting of
500 radiology images and the test set consisting of 444 radiology images. In addition, for the
concept detection task, an excel file containing the medical image ID and the corresponding
concepts CUIs is given to map each medical image onto its related concepts. Similarly, an excel
ifle containing the captions of each medical image is provided for the caption prediction task.
We present in Fig. 1 sample images with their underlying concepts and in Fig. 2 samples of
images with their captions.</p>
      <p>On further analysis of the datsaet, we illustrate in Fig. 3 the number of images per each
CUI concept. It is obvious that the most frequent CUI is the ’C0040398’ corresponding to
’Tomography, Emission-Computed’ with 1159 images. Moreover, Fig. 4 presents the number of
images by the number of CUIs associated to each image. We can see from this image that most
of the medical images are attached to 2 to 3 concepts whereas the maximum number of concept
CUIs per image is 10.</p>
      <p>Moreover, we noticed that the maximum number of sentences per image caption is 5 and the
maximum length of any caption is 47 words before pre-processing whereas it is 33 words after
the pre-processing.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>
        For this year edition of the ImageCLEFmed Caption challenge [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], two subtasks are put forward:
the concept detection and the caption prediction. To resolve each of these challenges, we present
in the current paper a model for the concept detection based on a multi-label classification
using a CNN architecture [20] and an encoder-decoder-based model for caption prediction.
First, data is pre-processed to transform the text into understandable units. Images are as well
pre-processed by performing a data augmentation technique on the training set. The detail is
provided in the next subsections.
      </p>
      <sec id="sec-4-1">
        <title>4.1. Data Pre-processing</title>
        <p>Text Pre-processing: We apply a pre-processing scheme on the image to concept matching
ifle to organize the concepts of each medical image into a list and create data-frames from
image IDs and their underlying CUIs. In addition, we pre-process the captions using the NLTK
package, by performing tokenization, punctuation and stop-words removal (using the default
NLTK’s "english" stopword list), lower casing each token and finally applying the stemming to
obtain filtered sentences for each medical image using the NLTK’s Snowball stemmer. Morover,
for the RNN decoder, we add two tokens: ’&lt;start&gt;’ and ’&lt;end&gt;’ to identify the beginning and
the end of each caption.</p>
        <p>Image Pre-processing: Image data generators are created for the three sets of data to
preprocess the images before feature extraction. These generators iterate over the data subsets
and normalize the images to facilitate the features calculation. Then we apply horizontal and
vertical flip in addition to a crop-center based data augmentation technique that we implement
with a fraction of 87.5%. The implemented data augmentation approach allowed us to expand
the training data without altering the visual content of the image. Then, the images are resized
to fit in the feature extractor size which is 224 × 224 × 3 in our case.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Concept Detection</title>
        <p>As we mentioned before, the first task of concept detection aims at identifying and localizing
the relevant concepts present in each medical image. Therefore, we exploit the visual image
content to extract significant visual features that allow us to distinguish the underlying concepts.
These concepts are used further to construct image captions and could as well be utilized for
the context-based images and information retrieval purposes.</p>
        <p>To achieve the first step towards caption prediction, we performed a multi label classification.
In the first run, we extracted image features using the pre-trained MobileNet-V2 and then
performed the classification using a GRU network [ 21]. In the second run, we performed feature
extraction using the pre-trained Inception-V3 model followed by a classification using a CNN
network. ImageNet weights were used for both models and features were extracted from the
last convolutional layer.</p>
        <p>As illustrated by Fig. 5, the features extracted from the radiology images using the pre-trained
MobileNet-V2 model are passed through a GRU layer, a flatten layer and then a fully connected
layer with a Relu activation function and dropout. Finally, the labels of each medical image are
predicted using a fully connected layer with a Sigmoid activation function.</p>
        <p>Likewise, the features extracted from the radiology images using the pre-trained Inception-V3
model are passed through a flatten layer, a fully connected layer with batch normalization, Relu
activation function, and dropout. Then, the probability of each class is calculated using a fully
connected layer with a Sigmoid activation function (as shown by Fig. 6). If the probability is
greater than 20%, we assert the input image belongs to that class. This probability was fixed to
20% after experimenting with diferent thresholds. If this threshold is fixed to a higher value, we
would get a lot of false negatives where many images are not classified in their correct classes.
However, if it is fixed to a smaller value, we would get a lot of false positives where many
images are categorized into incorrect classes. Finally, we map the labels to their corresponding
concepts.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Caption Prediction</title>
        <p>The second sub-task relies on the concept vocabulary detected in the first sub-task in addition
to the visual features extracted from the medical images to establish relationships between them
and predict descriptive caption for each medical image. We attempted to address the issue of
caption generation using an encoder-decoder architecture with attention mechanism.</p>
        <p>The visual features are extracted from the medical images using a pre-trained model
’InceptionV3’ where weights from the ImageNet were employed. Then, these features are passed
through a CNN model that is composed of a fully connected layer to flatten the feature vector.
Next, an attention mechanism is employed to focus on the most important parts of the image
and a context vector is constructed. Captions are pre-processed as we mentioned before and
passed to an embedding layer. A concatenation layer is farther used to merge the context
vector with the resulting embedding vector and the output is passed to a GRU layer. A flatten
layer and two fully connected layers with dropout and a Relu for the first one and a Sigmoid
activation function for the second were employed. Finally, relevant captions are generated word
by word until the ’&lt;end&gt;’ token is met. Figure. 7 illustrates the attention-based encoder-decoder
architecture we used to construct new captions for the medical images. Moreover, we exploit
the teacher forcing during the training by using the ground truth sequences at every step rather
than the sequence of newly generated words at previous steps.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiment and Results</title>
      <p>We used the data provided by the ImageCLEF medical task to evaluate the performance of our
models. Three subsets are used for the training, the validation and the test respectively. We
report in this section, the performance metrics calculated as well as the results we obtained for
both models.</p>
      <sec id="sec-5-1">
        <title>5.1. Performance metrics</title>
        <p>We calculate the F_Measure, using the default ’binary’ averaging method, for the concept
detection task as follows:
∑︁</p>
        <p>∑︁
 =
∈{} _∈
∑︁ ∑︁
(_)</p>
        <p>(_′)
′∈{} _′∈′</p>
        <p>=  · exp(∑︁  · log )
=1</p>
        <p>Where BP refers to the brevity penalty, N refers to the number of n_grams (uni-gram, bi-gram,
3-gram and 4-gram),  refers to the weight of each modified precision and  refers to the
modified precision. By default N=4 and = 1/N = 1/4.</p>
        <p>Brevity Penalty (BP) allows us to pick the candidate caption which is most likely close in
length, word choice and word order to the reference caption. It is an exponential decay and is
calculated as follows:
 =
{︃1</p>
        <p>&gt; 
exp(1− /)  ⩽</p>
        <p>Where r refers to the count of words in the reference caption and c refers to the count of
words in the candidate caption.</p>
        <p>Modified precision is computed for each n_gram as the sum of clipped n_gram counts of the
candidate sentences in the corpus divided by the number of candidate n_grams as shows "(6)"
[22]. It allows us to compute the adequacy and the fluency of the candidate translation to the
reference translation.</p>
        <p>· 
 _  = 2 ·   +</p>
        <p>Where the recall and the precision are calculated as follow and TP, FN, TN, FP correspond to
true positive, false negative, true negative and false positive respectively.</p>
        <p>=</p>
        <p>+  
  =   (3)</p>
        <p>+</p>
        <p>For the caption prediction task, we calculate the BLEU score by assuming that each caption
is a single sentence even if it is actually composed of several sentences. For that we use the
default implementation of the Python NLTK based on [22]:
(1)
(2)
(4)
(5)
(6)</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Results</title>
        <p>We obtained an average F_measure of 24.49% during the training process and a value of 14.3%
during the inference process for the concept detection task using the MobileNetV2 as a feature
extractor and the GRU network as a classifier.</p>
        <p>Similarly, we obtained an average F_measure of 23.28% during the training process and a value
of 13.7% during the inference process using the InceptionV3 model as a feature extractor and a
CNN network as a classifier. However, our F_measure results are comparatively lower compared
to the leading group (IALab_PUC) with a score of 50.5 %. Results are illustrated by Table. 1
where the run ID 136011 corresponds to the first configuration and the 136025 corresponds to
the second configuration. Figure. 8 shows the evolution of the multi-label classification model
accuracy across the epochs.</p>
        <p>For the caption prediction task, we obtained a BLEU score of 0.287 during the training process
and a value of 0.243 during the inference. Results are illustrated by Table. 2. In addition, Figure. 9
shows the loss calculated during the training process of the encoder-decoder model. We noticed
the decrease of the cross-entropy loss across the epochs.</p>
        <p>Finally, we show an example of a random medical image from the validation set with its real
caption and the newly generated caption in Fig. 10. We observed a BLEU score of 0.339 for this
image, where 8 words were correctly generated but the order of the words in the generated
caption is diferent.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and Future Work</title>
      <p>We presented in this paper our contribution to the ImageCLEF 2021 medical task where we
proposed a CNN based multi-label classification model for the concept detection task and an
attention-based encoder-decoder model for the caption prediction task. For both models, a
transfer learning is used to extract signicfiant features from the real radiology images and a
data augmentation based on center-cropping is applied to expand the used training subset. The
evaluation of the caption detection task is conducted using the mean F_measure for which
we obtained a score of 14.3%. Furthermore, BLEU score is used to evaluate the reliability of
the generated captions for the caption prediction task for which we obtained a score of 0.243.
We believe that we did not obtain promising results due to the small amount of data used
and the fact that we did not explore more fine-tuned parameters for both models for time
constraints. In addition, we did not include the textual features substituted by the medical
concepts to generate the new captions. In future work, we will integrate the textual features of
the images to the visual information to obtain more relevant captions. We will also investigate
more advanced deep learning algorithms inline with more fine-tuned parameters. In addition,
we will investigate our model performance on larger scale dataset for medical image captioning.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments References</title>
      <p>This work is supported by the Academy of Finland Profi5 DigiHealth project ( #326291), which
is gratefully acknowledged.
[20] Y. LeCun, Y. Bengio, et al., Convolutional networks for images, speech, and time series,</p>
      <p>The handbook of brain theory and neural networks 3361 (1995) 1995.
[21] K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, Y.
Bengio, Learning phrase representations using RNN encoder-decoder for statistical machine
translation, arXiv preprint arXiv:1406.1078 (2014).
[22] K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, Bleu: a Method for Automatic Evaluation of
Machine Translation, in: Proceedings of the 40th annual meeting of the Association for
Computational Linguistics, 2002, pp. 311–318.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I.</given-names>
            <surname>Allaouzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Ben</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Benamrou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ouardouz</surname>
          </string-name>
          ,
          <article-title>Automatic Caption Generation for Medical Images</article-title>
          ,
          <source>in: Proceedings of the 3rd International Conference on Smart City Applications</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Medical Image Labelling and Semantic Understanding for Clinical Applications</article-title>
          , in: International
          <source>Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>260</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Park</surname>
          </string-name>
          , J. Choi,
          <article-title>Feature Diference Makes Sense: A Medical Image Captioning Model Exploiting Feature Diference and Tag Information</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>95</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jacutprakart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEFmed 2021 Concept &amp; Caption Prediction Task</article-title>
          , in: CLEF2021 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bucharest, Romania,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Peteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sarrouti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kozlovski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Liauchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dicente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Jacutprakart</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Berari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Tauteanu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Fichou</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Brie</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Dogariu</surname>
            ,
            <given-names>L. D.</given-names>
          </string-name>
          <string-name>
            <surname>Ştefan</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>T. A.</given-names>
          </string-name>
          <string-name>
            <surname>Oliver</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Moustahfid</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Deshayes-Chossart</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2021: Multimedia retrieval in medical, nature, internet and social media applications, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 12th International Conference of the CLEF Association (CLEF</source>
          <year>2021</year>
          ),
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Bucharest, Romania,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alsharid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Drukker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chatelain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Papageorghiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Noble</surname>
          </string-name>
          , Captioning Ultrasound Images Automatically, in: International Conference on Medical Image Computing and
          <string-name>
            <surname>Computer-Assisted</surname>
            <given-names>Intervention</given-names>
          </string-name>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>338</fpage>
          -
          <lpage>346</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          , W. Liu, C. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.-S.</given-names>
            <surname>Hua</surname>
          </string-name>
          ,
          <article-title>Concept Detection based on Multi-label Classification and Image Captioning Approach-</article-title>
          DAMO at ImageCLEF
          <year>2019</year>
          ., in: CLEF (Working Notes),
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ling</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sreenivasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Anand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. R.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. V.</given-names>
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Qadir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Swisher</surname>
          </string-name>
          , et al.,
          <source>PRNA at ImageCLEF</source>
          <year>2017</year>
          <article-title>Caption Prediction and Concept Detection Tasks</article-title>
          .,
          <source>in: CLEF (Working Notes)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ling</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sreenivasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Anand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. R.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Qadir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Swisher</surname>
          </string-name>
          , et al.,
          <article-title>Attention-based Medical Caption Generation with Image Modality Classification and Clinical Concept Mapping</article-title>
          , in: International
          <source>Conference of the CrossLanguage Evaluation Forum for European Languages</source>
          , Springer,
          <year>2018</year>
          , pp.
          <fpage>224</fpage>
          -
          <lpage>230</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lyndon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Neural Captioning for the ImageCLEF 2017 Medical Image Challenges</article-title>
          .,
          <source>in: CLEF (Working Notes)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Rumelhart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <article-title>Learning representations by back-propagating errors</article-title>
          ,
          <source>nature</source>
          <volume>323</volume>
          (
          <year>1986</year>
          )
          <fpage>533</fpage>
          -
          <lpage>536</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
          </string-name>
          Short-term
          <string-name>
            <surname>Memory</surname>
          </string-name>
          ,
          <source>Neural computation 9</source>
          (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ambati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Dudyala</surname>
          </string-name>
          ,
          <article-title>A Sequence-to-Sequence Model Approach for ImageCLEF 2018 Medical Domain Visual Question Answering</article-title>
          ,
          <source>in: 2018 15th IEEE India Council International Conference (INDICON)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Benzarti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. B. A.</given-names>
            <surname>Karaa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. H. B.</given-names>
            <surname>Ghezala</surname>
          </string-name>
          ,
          <string-name>
            <surname>Cross-Model Retrieval Via Automatic Medical Image Diagnosis Generation</surname>
          </string-name>
          ,
          <source>in: International Conference on Intelligent Systems Design and Applications</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>561</fpage>
          -
          <lpage>571</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Cross</given-names>
            <surname>Modal</surname>
          </string-name>
          <article-title>Deep Learning Based Approach for Caption Prediction and Concept Detection by CS Morgan State</article-title>
          .,
          <source>in: CLEF (Working Notes)</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <article-title>Automatic Caption Generation of Retinal Diseases with Self-trained RNN Merge Model</article-title>
          ,
          <source>in: Advanced Computing and Systems for Security</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V.</given-names>
            <surname>Kougia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pavlopoulos</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Androutsopoulos</surname>
          </string-name>
          , AUEB NLP Group at ImageCLEFmed Caption
          <year>2019</year>
          ., in: CLEF (Working Notes),
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C. Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <article-title>Hybrid Retrieval-Generation Reinforced Agent for Medical Image Report Generation</article-title>
          , in
          <source>: Proceedings of the 32nd International Conference on Neural Information Processing Systems (NIPS</source>
          <year>2018</year>
          ), NIPS'18, Curran Associates Inc.,
          <year>2018</year>
          , p.
          <fpage>1537</fpage>
          -
          <lpage>1547</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Unifying Relational Sentence Generation and Retrieval for Medical Image Report Composition</article-title>
          ,
          <source>IEEE Transactions on Cybernetics</source>
          (
          <year>2020</year>
          ). doi:
          <volume>10</volume>
          . 1109/TCYB.
          <year>2020</year>
          .
          <volume>3026098</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>