<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>IUST_NLPLAB at ImageCLEFmedical Caption Tasks 2023</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yasaman Lotfollahi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Melika Nobakhtian</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Malihe Hajihosseini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sauleh Eetemadi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Assistant Professor of Computer Science, School of Computer Engineering, Iran University of Science and Technology</institution>
          ,
          <addr-line>Tehran</addr-line>
          ,
          <country>Islamic Republic Of Iran</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Student at School of Computer Engineering, Iran University of Science and Technology</institution>
          ,
          <addr-line>Tehran</addr-line>
          ,
          <country>Islamic Republic Of Iran</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>We present models implemented by the IUST_NLPLAB group for ImageCLEFmedical Caption Task 2023. This task contains two subtasks: Concept Detection and Caption Prediction. Under the first subtask, the model should extract medical concepts contained in radiology images. These concepts can be used for context-based image and information retrieval. Under the second subtask, the model predicts the caption for a medical image. This can be used for improving the diagnosis and treatment of diseases by saving time, money and helping physicians. This was our second experience to participate in this competition. We used difrent models for both subtasks. We were able to get the 4th rank in the concepts detection subtask with a score of 0.49. Also, in the caption prediction subtask, we were able to get the 12th rank based on the BERTScore evaluation metric. This is despite the fact that our model has won the first rank based on ROUGE, BLEU and METEOR. From this, it can be concluded that the type of evaluation metric determined has an important efect on the results of this subtask.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Medical Image Captioning</kwd>
        <kwd>Concept Detection</kwd>
        <kwd>Caption Prediction</kwd>
        <kwd>Deep Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In 2022, imageclef used the AIcrowd2 platform to publish contest data and receive submissions
from participating groups[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In that platform, groups could see the score earned after each
submission and plan to improve their models’score. However, the score obtained by other
groups could not be seen.
      </p>
      <p>
        In ImageCLEF 2023[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the contest data was made available to participating groups via a
private GitHub link. Also, sciebo3 system was used to receive the results sent by the groups.
Participating groups could have a maximum of 10 successful submissions in each subtask. In
each run, in addition to the test data results in csv format, a txt file containing a brief description
of the method should be attached. Unfortunately, unlike the AIcrowd platform, in the sciebo
system, the scores obtained after each submission were not presented, and it was not possible
to improve the models and analyze them by comparing the obtained results.
      </p>
      <p>In ImageCLEFmedical 2023, four tasks were proposed
1. Image Captioning.
2. Controlling the Quality of Synthetic Medical Images created via GANs.
3. Visual Question Answering for Colonoscopy Images.
4. Medical Dialogue Summarization.</p>
      <p>
        We selected the Image Captioning task from the ImageCLEFmedical section to participate in
the competition. ImageCLEF medical Image Captioning task in 2023[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], like last year, contained
two subtasks: Concepts Detection and Caption Prediction. Each group could participate in
one or both subtasks. In this paper, we present the methods our group, IUST_NLPLAB, from
the Iran University of Science and Technology4, School of Computer Engineering5, Natural
Language Pocessing Laboratory6 used in both subtasks. This is our second time participating
in the ImageCLEF competition. We participated in both subtasks and registered 7 successful
submissions in the concept detection subtask and 10 successful submissions in the caption
prediction subtasks.
      </p>
      <p>
        In the concept detection subtask, we were able to win the fourth place in the competition
with a gap of about 2 percent in F1-score from the first ranked group. Also, in the subtask of
caption prediction, we were able to get the 12th rank of the competition based on BERTScore[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
which was the main evaluation metric of the competition. But based on other evaluation metrics
such as ROUGE[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], BLEU[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and METEOR[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], our group was able to win the first rank among
other participating groups.
      </p>
      <p>In the following sections, we will describe the task, datasets, models developed and the results
we achieved in detail.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task description</title>
      <p>
        This year the ImageCLEF evaluation campaign hosted the 7th edition of the medical image
caption task. Unlike some of the previous editions which only contained the caption prediction
2https://www.aicrowd.com/ (last accessed: 2023-07-05)
3https://hochschulcloud.nrw/en/ (last accessed: 2023-07-05)
4http://www.iust.ac.ir/en (last accessed: 2023-07-05)
5http://ce-inter.iust.ac.ir/ (last accessed: 2023-07-05)
6https://nlplab.iust.ac.ir (last accessed: 2023-07-05)
task (e.g., 2016[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) or only the concept detection task (e.g., 2019[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]), the 7th edition, as like as
last year, contained both subtasks as described below.
      </p>
      <sec id="sec-2-1">
        <title>2.1. Concept Detection</title>
        <p>In this subtask, the goal is to extract medical concepts in medical images. These concepts
are selected from UMLS7[11] Concept Unique Identifiers (CUIs) specified in the dataset. The
extraction of these concepts can be used for image retrieval and context-based information
purposes.</p>
        <p>
          The 2023 dataset contains 2,125 medical concepts, which has decreased compared to last
year’s dataset. Table 1 shows a list of the 15 most frequent concepts in the training collection
based on their frequency. According to the table published in the [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], most of the most frequent
concepts of the 2023 dataset are in common with the most frequent concepts of the 2022 dataset,
but their frequency has decreased compared to last year. The lowest rate frequency of concepts
in the training set is related to six concepts, each of them was repeated only 2 times.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Caption Prediction</title>
        <p>In this subtask, the goal is to generate a suitable caption for the input medical images. Extracting
medical concepts can help in producing a more appropriate captions. This subtask consists of a
combination of text and image processing and is more complicated than the previous subtask.
7Unified Medical Language System ®</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Data</title>
      <p>The dataset introduced for the ImageCLEFmedical Caption 2023 is a subset of the Radiology
Objects in COntext (ROCO)[12] dataset. The dataset published in 2023 was structurally similar
to the dataset of 2022. In this year’s dataset, there were 60,918 training data, which was reduced
compared to last year’s dataset, but the validation and testing datasets included 10,437 and
10,473 data, respectively, which increased compared to last year. For each image in the training
and validation dataset, the concepts in the image and a suitable caption of it were provided.</p>
      <p>In the following, more details of the data of each subtask are provided.</p>
      <sec id="sec-3-1">
        <title>3.1. Image Concepts</title>
        <p>In this subtask, each image in the dataset has several related concepts. These concepts have
originated from the Unified Medical Language System (UMLS)[ 11] Concept Unique Identifiers
(CUIs). The generated concepts are based on a reduced subset of the UMLS 2022 AB release8
this year. Filtering images according to their semantic type was performed to reach a higher
possibility of recognizing concepts in images. Concepts with a low occurrence were removed
based on recommendations from previous years.</p>
        <p>Each image has a diferent number of concepts. The overall number of concepts is 2125. An
image has at least one related concept and at most 24 concepts. Most images in the dataset have
three concepts.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Image Captions</title>
        <p>A caption is provided for each image in the training and validation sets in this subtask. Last
year, the provided captions were pre-processed in four stages, but according to the explanations
provided by the organizers, in this year only one pre-processing step, removal of links from the
captions, was done on the captions.</p>
        <p>Based on the analysis performed on the training dataset, 63 images have one-word captions,
which is the shortest caption length in the dataset. The maximum length of the caption is 410
words, which is related to one image. Also, the average number of words in captions is 20
words. We also calculated the TTR9 for this annotation dataset. TTR is obtained by dividing
the number of unique words by the text size and is a simple measure of lexical diversity[13].
Considering the stop words, the TTR value in this dataset is 0.07 and without considering the
stop words, it is 0.05, both of them have increased compared to last year. Figure 1 displays
the frequently recurring words in the captions of the training set, along with their frequency
including and excluding stop words.
8https://www.nlm.nih.gov/pubs/techbull/nd22/nd22_umls_2022ab_release_available.html (last accessed: 2023-07-05)
9Type-token ratio
• Contrast-enhanced CT scan
of the lower abdomen and
pelvis showing a single lobe
of a presumed, bilobed
pseudoaneurysm (a) as well as
a 3.5 × 5.5 × 6 cm
rimenhancing, lobular
collection of the superior right
gluteal subcutaneous tissues,
just superior to the right
iliac crest and lateral to the
paraspinal musculature,
consistent with a hematoma (b)
(a) Ten most frequent words in the training set.</p>
        <p>(b) Ten most frequent words in the training set
without stop words.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <sec id="sec-4-1">
        <title>4.1. Concept Detection</title>
        <p>In this section, we present our methods for concept detection and caption prediction subtask.
In concept detection subtask, we used diferent image preprocessing methods. When we used
CLIP[18] and PubMedCLIP[19] models, we used their preprocess method. When used other
pretrained models such as Resnet[20] and Eficientnet[ 21], we used CLAHE[22]. CLAHE is the
one of ways to increase quality of image.</p>
        <p>In the following, we explain our developed models for concept detection subtask. We used
two methods: Ensemble Models as v1 and Multi-Label Classification Method as v2. Table 3
shows the details of the all developed models in concept detection subtask.</p>
        <sec id="sec-4-1-1">
          <title>4.1.1. Ensemble Models</title>
          <p>One of the systems that we designed was based on ensemble systems. We adopted this method
according to the winner of last year, the AUEB-NLP group[23]. We utilized two instances
of EficientNetV2B0[ 24] for this model. All layers of base models were frozen until the last
convolutional layer during the training process. A dense layer was added to each model to
predict concepts. Models were trained for 50 epochs with a batch size of 256. We considered
diferent thresholds for each concept to find out if a concept is related to an image or not. We
tried certain thresholds on validation data and found the best one regarding F1-score. Every
ifve-epoch model weights and best thresholds were saved. After training, the best weights
and thresholds were chosen for each model. To predict concepts, if both models assigned a
concept to an image, we concluded that this image has this concept, in other words, we used an
intersection of concepts.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Multi-Label Classification Method</title>
          <p>In this approach, we built a multi-label classification model to predict the correct concepts for
each input image. We used CNNs with pre-trained weights from ImageNet[25]. These networks
were modified by removing their final layer and adding a classification layer. Then, they were
ifne-tuned on the target dataset. We tried fine-tuning diferent pre-trained models and applied
various thresholds to find the best results. In this year, we also used the vision-language models
of CLIP[18] and its medical version, PubMedCLIP[19], which had achieved good results in many
tasks.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Caption Prediction</title>
        <p>In the caption prediction subtask, we utilized the approach we developed for the previous year’s
challenge, where we achieved first place. This methodology treats each word in a caption as a
label corresponding to the associated image. We trained a multi-label classification model to
predict the words that will ultimately form a caption for the given image.</p>
        <p>To extract image features for the subtask, we used a pre-trained CNN on ImageNet[25]
and fine-tuned it on the training set. During fine-tuning, we excluded the last layer of the
CNN and added a dropout layer and a dense layer. We tried various CNN models to explore
diferent possibilities. Similar to the concept detection subtask, in this subtask, we used the
vision-language models of CLIP[18] and its medical version, PubMedCLIP[19].</p>
        <p>The model generates captions for images by predicting the corresponding words. The
probability of each word is computed in the output layer using the sigmoid activation function.
Two methods are used to select the candidate words:
1. The top  words with the highest probability are chosen.  is a hyper-parameter that
will define the length of captions.
2. A threshold is applied to the model output. Words with probabilities higher than the
threshold are chosen to create the caption.</p>
        <p>After extracting the correct words, we need to sort them to create the full caption. Two
methods are used to arrange the words:
1. Words are arranged from highest to lowest probability.
2. Words are ordered based on their statistical occurrence within the training set. Each word
is assigned to its most common position in the caption.</p>
        <p>
          Diferent values of  and threshold were applied to the output. Although BERTScore[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] is
the primary score in this year’s competition, we were not able to use this metric to evaluate
our models because of our resource limits. Therefore, we used the BLEU[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] score to evaluate
our models and find the best hyperparameters. Details of each submission and their results are
described in table 4 and 5.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>In the previous parts, the details of the models implemented by our group and the results
obtained by each one were explained.</p>
      <p>In the concept detection subtask, two metrics, F1-Score and F1-Score Manual, which was
calculated using a subset of manually validated concepts, were used to evaluate the models,
but the results of the competition was based on the F1-Score metric. In 2023, 9 groups from all
over the world participated in the concept detection subtask and managed to register successful
submissions. The details of the results announced by the organizers of the competition in this
subtask are presented in Table 6. Among the submissions of our group, Run ID 7, which used
the basic model of PubMedCLIP ViT-B/32[19] to extract the features of the images, was able
to get the best result with a diference of about 2 percent from the first group and achived 4th
rank in this competition, which has increased four ranks compared to last year’s results.</p>
      <p>
        In the caption prediction subtask, seven metrics: BERTScore[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], ROUGE[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], BLEURT[26],
BLEU[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], METEOR[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], CIDEr[27] and CLIPScore[28], were used to evaluate the models, but the
results of the competition was based on the BERTScore metric. In 2023, 13 groups from all over
the world participated in the caption prediction subtask and managed to register successful
submissions. The details of the results announced by the organizers of the competition in this
subtask are presented in Table 7. Among the submissions of our group, Run ID 6, which used
the basic model of PubMedCLIP RN50x4[19] to extract the features of the images, was able to
get the best result with a diference of about 7 Percent from the first group and achived 12th
rank in this competition.
      </p>
      <p>The noteworthy point is that based on three metrics ROUGE, BLEU and METEOR, our group
has been able to get the first rank among the participating groups.</p>
      <p>Diferences in rankings based on diferent metrics can show challenges in evaluating generated
captions. This is due to the diferences in how these metrics evaluate the quality of generated
captions. For example, the BLEU score measures n-gram overlap between the generated and
reference captions, rewarding precision, and the presence of matching n-grams. In contrast,
BERTScore used contextualized embeddings from BERT to capture semantic similarity, taking
into account both word order and correctness. Consequently, a result with higher n-gram
overlap but potential issues in word order or overall fluency could receive a better BLEU score
and a lower BERTScore.</p>
      <p>While our models can generate relevant words to describe the input image, they struggle to
shape them into actual sentences which are semantically similar to the original caption. This
issue can be one of the reasons why our results have low BERTScores, while having high BLEU
scores.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This paper describes the participation of IUST_NLPLAB at Iran University of Science and
Technology at ImageCLEFmedical caption 2023 task.</p>
      <p>In the concept detection subtask, we ranked 4 among 9 participating teams. We used MLC
and ensemble models in this subtask. Our MLC methods with PubMedCLIP ViT-B/32 as a base
model had better overall score.</p>
      <p>In the caption prediction subtask, last year, our group won the first place in the competition
based on the BLEU evaluation metric, but this year it won the 12th place in the competition
based on the BERTScore metric. Based on the published results, our group was able to win
ifrst place in all 10 of its submissions based on the three metric ROUGE, BLEU and METEOR.
Based on these results, it can be concluded that the selection of evaluation metric in the analysis
of the models presented for this subtask is very important and the results based on diferent
evaluation metric can have significant diferences from each other.</p>
      <p>This year was our second experience of participating in this competition and we hope to be
able to participate in these competitions in the coming years and gain new experiences.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work has been supported by School of Computer Engineering of Iran University of Science
and Technology and we are very grateful for their support.
medicine, lifelogging, security and nature, in: Experimental IR Meets Multilinguality,
Multimodality, and Interaction, Proceedings of the Tenth International Conference of
the CLEF Association (CLEF 2019), LNCS Lecture Notes in Computer Science, Springer,
Lugano, Switzerland, 2019.
[11] O. Bodenreider, The unified medical language system (umls): integrating biomedical
terminology, Nucleic acids research 32 (2004) D267–D270.
[12] O. Pelka, S. Koitka, J. Rückert, F. Nensa, C. M. Friedrich, Radiology objects in context
(roco): a multimodal image dataset, in: Intravascular Imaging and Computer Assisted
Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis:
7th Joint International Workshop, CVII-STENT 2018 and Third International Workshop,
LABELS 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 16,
2018, Proceedings 3, Springer, 2018, pp. 180–189.
[13] K. Kettunen, Can type-token ratio be used to show morphological complexity of languages?,</p>
      <p>Journal of Quantitative Linguistics 21 (2014) 223–245.
[14] N. Unterstell, A. L. Bressan, L. A. Serpa, P. P. d. Fonseca e Castro, A. C. Gripp, Systemic
sarcoidosis induced by etanercept: first brazilian case report, An. Bras. Dermatol. 88 (2013)
197–199.
[15] A. Zbiciak, T. Markiewicz, A new extraordinary means of appeal in the polish criminal
procedure: the basic principles of a fair trial and a complaint against a cassatory judgment,
Access to Justice in Eastern Europe 6 (2023) 1–18.
[16] S. Ananthi Kumarasamy, B. S. Kannadath, S. Soundamourthy, A. Subramanian, S. P.
Sinhasan, R. V. Bhat, Semimembranosus ganglion cyst, Anat. Cell Biol. 47 (2014) 207–209.
[17] M. Counihan, M. E. Pontell, B. Selvan, A. Trebelev, A. Nunez, Delayed presentation of a
lumbar artery pseudoaneurysm resulting from isolated penetrating trauma, J. Surg. Case
Rep. 2015 (2015) rjv083.
[18] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell,
P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language
supervision, in: International conference on machine learning, PMLR, 2021, pp. 8748–8763.
[19] S. Eslami, G. de Melo, C. Meinel, Does clip benefit visual question answering in the medical
domain as much as it does in the general domain?, arXiv preprint arXiv:2112.13906 (2021).
[20] M. Talo, Convolutional neural networks for multi-class histopathology image classification,</p>
      <p>ArXiv abs/1903.10035 (2019).
[21] M. Tan, Q. Le, Eficientnet: Rethinking model scaling for convolutional neural networks,
in: International conference on machine learning, PMLR, 2019, pp. 6105–6114.
[22] K. Zuiderveld, Contrast limited adaptive histogram equalization, Graphics gems (1994)
474–485.
[23] V. Kougia, J. Pavlopoulos, I. Androutsopoulos, Aueb nlp group at imageclefmed caption
2022, in: Conference and Labs of the Evaluation Forum, 2019.
[24] M. Tan, Q. V. Le, Eficientnetv2: Smaller models and faster training, 2021.</p>
      <p>arXiv:2104.00298.
[25] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical
image database, in: 2009 IEEE conference on computer vision and pattern recognition,
Ieee, 2009, pp. 248–255.
[26] T. Sellam, D. Das, A. P. Parikh, Bleurt: Learning robust metrics for text generation, arXiv
preprint arXiv:2004.04696 (2020).
[27] R. Vedantam, C. Lawrence Zitnick, D. Parikh, Cider: Consensus-based image description
evaluation, in: Proceedings of the IEEE conference on computer vision and pattern
recognition, 2015, pp. 4566–4575.
[28] J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, Y. Choi, Clipscore: A reference-free evaluation
metric for image captioning, arXiv preprint arXiv:2104.08718 (2021).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Peteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bloch</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Brüngel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Idrissi-Yaghir</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Schäfer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kozlovski</surname>
            ,
            <given-names>Y. D.</given-names>
          </string-name>
          <string-name>
            <surname>Cid</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Kovalev</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-D. Ştefan</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Dogariu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Deshayes-Chossart</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Schindler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2022: Multimedia retrieval in medical, social media and nature applications, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 13th International Conference of the CLEF Association (CLEF</source>
          <year>2022</year>
          ),
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hajihosseini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lotfollahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nobakhtian</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Javid</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Omidi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Eetemadi</surname>
          </string-name>
          , Iust_nlplab at imageclefmedical caption tasks (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Drăgulinescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Snider</surname>
          </string-name>
          , G. Adams,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yetisgen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Garcıa Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bloch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brüngel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Idrissi-Yaghir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hicks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Riegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Thambawita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Storås</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Papachrysos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schöler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Andrei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radzhabov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Coman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stan</surname>
          </string-name>
          , G. Ioannidis,
          <string-name>
            <given-names>H.</given-names>
            <surname>Manguinhas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ştefan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dogariu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Deshayes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Popescu</surname>
          </string-name>
          , Overview of ImageCLEF 2023:
          <article-title>Multimedia retrieval in medical, socialmedia and recommender systems applications</article-title>
          , in: Experimental IR Meets Multilinguality, Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 14th International Conference of the CLEF Association (CLEF</source>
          <year>2023</year>
          ), Springer Lecture Notes in Computer Science LNCS, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Seco de Herrera</surname>
          </string-name>
          , L. Bloch,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brüngel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Idrissi-Yaghir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          , Overview of ImageCLEFmedical 2023 -
          <article-title>Caption Prediction and Concept Detection</article-title>
          , in: CLEF2023 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kishore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Artzi</surname>
          </string-name>
          , Bertscore:
          <article-title>Evaluating text generation with bert</article-title>
          , arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>09675</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.-Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Rouge: A package for automatic evaluation of summaries</article-title>
          , in: Text summarization branches out,
          <year>2004</year>
          , pp.
          <fpage>74</fpage>
          -
          <lpage>81</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Papineni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roukos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ward</surname>
          </string-name>
          , W.-J. Zhu,
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          ,
          <source>in: Proceedings of the 40th annual meeting of the Association for Computational Linguistics</source>
          ,
          <year>2002</year>
          , pp.
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Denkowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lavie</surname>
          </string-name>
          ,
          <article-title>Meteor universal: Language specific translation evaluation for any target language</article-title>
          ,
          <source>in: Proceedings of the ninth workshop on statistical machine translation</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>376</fpage>
          -
          <lpage>380</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schaer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bromuri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2016 medical task</article-title>
          ,
          <source>in: Working Notes of CLEF 2016 (Cross Language Evaluation Forum)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Péteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. D.</given-names>
            <surname>Cid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Liauchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Klimuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tarasau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          , D.-
          <string-name>
            <surname>T.</surname>
            Dang-Nguyen,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
            , M.-
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Pelka</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Kavallieratou</surname>
            ,
            <given-names>C. R.</given-names>
          </string-name>
          <string-name>
            <surname>del Blanco</surname>
            ,
            <given-names>C. C.</given-names>
          </string-name>
          <string-name>
            <surname>Rodríguez</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Vasillopoulos</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Karampidis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Campello,
          <string-name>
            <surname>ImageCLEF</surname>
          </string-name>
          <year>2019</year>
          :
          <article-title>Multimedia retrieval in</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>