<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Clinical Medicine 9 (2020) 3992. doi:10.3390/jcm9123992.
[24] R. J. Roberts</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.3390/jcm9123992</article-id>
      <title-group>
        <article-title>Overview of ImageCLEFmedical 2022 - Caption Prediction and Concept Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Johannes Rückert</string-name>
          <email>johannes.rueckert@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asma Ben Abacha</string-name>
          <email>abenabacha@microsoft.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alba G. Seco de Herrera</string-name>
          <email>alba.garcia@essex.ac.uk</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Louise Bloch</string-name>
          <email>louise.bloch@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raphael Brüngel</string-name>
          <email>raphael.bruengel@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ahmad Idrissi-Yaghir</string-name>
          <email>ahmad.idrissi-yaghir@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henning Schäfer</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henning Müller</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph M. Friedrich</string-name>
          <email>christoph.friedrich@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Applied Sciences and Arts Dortmund</institution>
          ,
          <addr-line>Dortmund</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Medical Informatics, Biometry and Epidemiology (IMIBE), University Hospital Essen</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Transfusion Medicine, University Hospital Essen</institution>
          ,
          <addr-line>Essen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Microsoft</institution>
          ,
          <addr-line>Redmond, Washington</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Applied Sciences Western Switzerland (HES-SO)</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Essex</institution>
          ,
          <addr-line>Wivenhoe Park, Colchester CO4 3SQ</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University of Geneva</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <volume>37</volume>
      <fpage>381</fpage>
      <lpage>382</lpage>
      <abstract>
        <p>The 2022 ImageCLEFmedical caption prediction and concept detection tasks follow similar challenges that were already run from 2017-2021. The objective is to extract Unified Medical Language System (UMLS) concept annotations and/or captions from the image data that are then compared against the original text captions of the images. The images used for both tasks are a subset of the extended Radiology Objects in COntext (ROCO) data set which was used in ImageCLEFmedical 2020. In the caption prediction task, lexical similarity with the original image captions is evaluated with the BiLingual Evaluation Understudy (BLEU) score. In the concept detection task, UMLS terms are extracted from the original text captions, combined with manually curated concepts for image modality and anatomy, and compared against the predicted concepts in a multi-label way. The F1-score was used to assess the performance. The task attracted a strong participation with 20 registered teams. In the end, 12 teams submitted 157 graded runs for the two subtasks. Results show that there is a variety of techniques that can lead to good prediction results for the two tasks. Participants used image retrieval systems for both tasks, while multi-label classification systems were used mainly for the concept detection, and Transformer-based architectures primarily for the caption prediction subtask.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Concept Detection</kwd>
        <kwd>Computer Vision</kwd>
        <kwd>ImageCLEF 2022</kwd>
        <kwd>Image Understanding</kwd>
        <kwd>Image Modality</kwd>
        <kwd>Radiology</kwd>
        <kwd>Caption Prediction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The caption task was first proposed as part of the ImageCLEFmedical [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] in 2016. In 2017 and
2018 [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] the ImageCLEFmedical caption task comprised two subtasks: concept detection and
caption prediction. In 2019 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and 2020 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the task concentrated on extracting Unified Medical
Language System® (UMLS) Concept Unique Identifiers (CUIs) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] from radiology images.
      </p>
      <p>
        In 2021 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], both subtasks, concept detection and caption prediction, were running again due
to participants demands. The focus in 2021 was on making the task more realistic by using
fewer images which were all manually annotated by medical doctors. As additional data of
similar quality is hard to acquire, the 2022 ImageCLEFmedical caption task continues with both
subtasks albeit with an extended version of the Radiology Objects in COntext (ROCO) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] data
set used for both subtasks, which was already used in 2020 and 2019.
      </p>
      <p>
        This paper sets forth the approaches for the caption task: automated cross-referencing of
medical images and captions into predicted coherent captions implying UMLS concept detection
in radiology images as a first step. This task is a part of the ImageCLEF benchmarking campaign,
which has proposed medical image understanding tasks since 2003; a new suite of tasks is
generated each subsequent year. Further information on the other proposed tasks at ImageCLEF
2022 can be found in Ionescu et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        This is the 6th edition of the ImageCLEFmedical caption task. Just like in 2016 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], 2017 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
2018 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and 2021 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], both subtasks of concept detection and caption prediction are included
in ImageCLEFmedical Caption 2022. Like in 2020, an extended subset of the ROCO [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] data set
is used to provide a much larger data set compared to 2021.
      </p>
      <p>Manual generation of the knowledge of medical images is a time-consuming process prone to
human error. As this process requires assistance for the better and easier diagnoses of diseases
that are susceptible to radiology screening, it is important that we better understand and refine
automatic systems that aid in the broad task of radiology-image metadata generation. The
purpose of the ImageCLEFmedical 2022 caption prediction and concept detection tasks is the
continued evaluation of such systems. Concept detection and caption prediction information is
applicable to unlabelled and unstructured data sets and medical data sets that do not have textual
metadata. The ImageCLEFmedical caption task focuses on the medical image understanding in
the biomedical literature and specifically on concept extraction and caption prediction based on
the visual perception of the medical images and medical text data such as medical caption or
UMLS CUIs paired with each image (see Figure 1).</p>
      <p>
        For the development data, an extended subset of the ROCO [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] data set from 2020 was used,
with new images from the same source added for the validation and test sets.
      </p>
      <p>This paper presents an overview of the ImageCLEFmedical caption task 2022 including the
task and participation in Section 2, the data creation in Section 3, and the evaluation methodology
in Section 4. The results are described in Section 5, followed by conclusion in Sections 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task and Participation</title>
      <p>In 2022, the ImageCLEFmedical caption task consisted of two subtasks: concept detection and
caption prediction.</p>
      <p>
        The concept detection subtask follows the same format proposed since the start of the task
in 2017. Participants are asked to predict a set of concepts defined by the UMLS CUIs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] based
on the visual information provided by the radiology images.
      </p>
      <p>The caption prediction subtask follows the original format of the subtask used between 2017
and 2018. The task is running again since 2021 because of participant demand. This subtask
aims to automatically generate captions for the radiology images provided.</p>
      <p>In 2022, 20 teams registered and signed the End-User-Agreement that is needed to download
the development data. 12 teams submitted 157 runs for evaluation (all 12 teams submitted
working notes) attracting more attention than in 2021. Each of the groups was allowed a
maximum of 10 graded runs per subtask.</p>
      <p>Table 1 shows all the teams who participated in the task and their submitted runs. 11 teams
participated in the concept detection subtask this year, 3 of those teams also participated in 2021.
10 teams submitted runs to the caption prediction subtask, 4 of those teams also participated in
2021. Overall, 9 teams participated in both subtasks, two teams participated only in the concept
detection subtask and one team participated only in the caption prediction subtask.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Data Creation</title>
      <p>In the previous edition, in an attempt to make the task more realistic, the data set contained
a smaller number of real radiology images annotated by medical doctors which resulted in
high-quality concepts.</p>
      <p>
        Additional data of similar quality is hard to acquire and so it was decided to return to the
data set already used in 2020 and 2019, which originates from biomedical articles of the PMC
PoliMiImageClef [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]
SDVA-UCSD [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
SSNSheerinKavitha
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
vcmi [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
      </p>
      <p>San Diego VA HCS, San Diego, CA, USA
Department of CSE, Sri Sivasubramaniya
Nadar College of Engineering, India
University of Porto, Porto, Portugal and
INESC TEC, Porto, Portugal
10
10
5
10
10
–
10
8
Open Access Subset1 [24] and was extended with new images added since the last time the data
set was updated.</p>
      <p>All captions were pre-processed by removing punctuation, numbers and words containing</p>
      <sec id="sec-3-1">
        <title>1https://www.ncbi.nlm.nih.gov/pmc/tools/openftlist/ [last accessed: 28.06.2022]</title>
        <p>numbers. Additionally, lemmatization was applied using spaCy2 and the pre-trained model
en_core_web_lg. Finally, all captions were converted to lower-case.</p>
        <p>From the resulting captions, UMLS concepts were generated using a reduced subset of the
UMLS 2020 AB release3, which includes the sections (restriction levels) 0, 1, 2, and 9. To improve
the feasibility of recognizing concepts from the images, concepts were filtered based on their
semantic type. Concepts with very low frequency were also removed, based on suggestions
from previous years.</p>
        <p>
          Additional concepts were assigned to all images addressing their image modality. Six modality
concepts were covered: x-ray, computer tomography (CT), magnetic resonance imaging (MRI),
ultrasound, and positron emission tomography (PET) as well as modality combinations (e.g.,
PET/CT) as standalone concept. For images of the x-ray modality further concepts on the
represented anatomy were assigned, covering specific anatomical body regions of the Image
Retrieval in Medical Application (IRMA) [25] classification: cranium, spine, upper extremity/arm,
chest, breast/mamma, abdomen, pelvis, and lower extremity/leg. Both of the described concept
extensions were created performing a two-stage process, each. In the first stage predictions
via classification models were created and assigned as annotations. For modality prediction
for all images a model trained on the ROCO dataset [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], and for anatomy prediction for x-ray
modality images a model trained on an existing IRMA-annotated image dataset [26] was used.
In the second stage, these annotations underwent manual quality control measures, involving
correction of faulty predictions and filtering of images that did not represent one of the minded
modality or anatomy concepts. Three annotators were involved. Each individual modality
concept was processed by a single annotator due to the low complexity of this task part. Anatomy
concepts of x-ray modality images were each, too, processed by a single annotator per concept.
However, due to the complexity/ambiguity of this task, the one annotator most-experienced
in anatomy classification re-evaluated the assessments of the other two. This re-evaluation
resulted in very few adjustments, indicating high agreement between annotators.
        </p>
        <p>The following subsets were distributed to the participants where each image has one caption
and multiple concepts (UMLS-CUI):
• Training set including 83,275 radiology images and associated captions and concepts.
• Validation set including 7,645 radiology images and associated captions and concepts.
• Test set including 7,645 radiology images.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation Methodology</title>
      <p>In this year’s edition, the performance evaluation is carried out in the same way as last year,
with both subtasks being evaluated separately.</p>
      <p>For the concept detection subtask, the balanced precision and recall trade-of were measured
in terms of F1-scores. In addition, a secondary F1-score was introduced in this edition, where
the score is computed using a subset of concepts that was manually curated and only contains
x-ray anatomy and image modality concepts.</p>
      <sec id="sec-4-1">
        <title>2https://spacy.io/api/lemmatizer/ [last accessed: 28.06.2022] 3https://www.nlm.nih.gov/pubs/techbull/nd20/nd20_umls_release.html [last accessed: 28.06.2022]</title>
        <p>Caption prediction performance is evaluated based on the BiLingual Evaluation Understudy
(BLEU) scores [27], which a geometric mean of n-gram scores from 1 to 4. As a preprocessing
step for the evaluation, all captions were lowercased and stripped of all punctuation and English
stop words. Additionally, to increase coverage, lemmatization was applied using spaCy and the
pre-trained model en_core_web_lg. BLEU values are then computed for each test image, treating
the entire caption as one sentence, even though it may contain multiple sentences. The average
of the BLEU values for all images is reported as the primary ranking score. Since evaluating
generated text and image captioning is very challenging and should be based on a single metric,
additional evaluation metrics were explored in this year’s edition in order to find the metric
that correlate well with human judgements for this task. First, the Recall-Oriented Understudy
for Gisting Evaluation (ROUGE) [28] score was adopted as a secondary metric that counts the
number of overlapping units such as n-grams, word sequences, and word pairs between the
generated text and the reference. Specifically, the ROUGE-1 (F-measure) score was calculated,
which measures the number of matching unigrams between the model-generated text and
a reference. All individual scores for each caption are then summed and averaged over the
number of captions, resulting in the final score. In addition to ROUGE, the Metric for Evaluation
of Translation with Explicit ORdering (METEOR) [29] was explored, which is a metric that
evaluates the generated text by aligning it to reference and calculating a sentence-level similarity
score. Furthermore, the Consensus-based Image Description Evaluation (CIDEr) [30] metric was
also adopted. CIDEr is an automatic evaluation metric that calculates the weights of n-grams
in the generated text and the reference text based on term frequency and inverse document
frequency (TF-IDF), and then compares them based on cosine similarity. Another used metric is
the Semantic Propositional Image Caption Evaluation (SPICE) [31], which maps the reference
and generated captions to semantic scene graphs through dependency parse trees and measures
the similarity between the scene graphs for the evaluation. Finally, BERTScore [32] was used,
which is a metric that computes a similarity score for each token in the generated text with
each token in the reference text. It leverages the pre-trained contextual embeddings from
BERT-based models and matches words by cosine similarity. In this work, the pre-trained model
microsoft/deberta-xlarge-mnli4 was utilized, since it is the model that correlates best with human
evaluation according to the authors5.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>For the concept detection and caption prediction subtasks, Tables 2 and 3 show the best results
from each of the participating teams. The results will be discussed in this section.</p>
      <sec id="sec-5-1">
        <title>5.1. Results for the Concept Detection subtask</title>
        <p>In 2022, 11 teams participated in the concept prediction subtask, submitting 85 runs. Table 2
presents the results achieved in the submissions.</p>
        <sec id="sec-5-1-1">
          <title>4https://huggingface.co/microsoft/deberta-xlarge-mnli [last accessed: 28.06.2022] 5https://github.com/Tiiiger/bert_score [last accessed: 28.06.2022]</title>
          <p>
            AUEB-NLP-Group Like in previous years, the AUEB-NLP-Group submitted the best
performing result with a primary F1-score of 0.4511 [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] and a secondary F1-score of 0.7907.
The winning approach was an ensemble of two EficientNetV2-B0 backbones followed by
a single classification layer where the union of predicted concepts was used to form the
ensemble. This solution outperformed their retrieval-based system which won last year’s
concept detection subtask [33].
fdallaserra The second best system, with an only slightly worse primary F1-score of 0.4505
and a better secondary F1-score of 0.8222 [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] was proposed by CMRE-UoG (fdallaserra).
Their best approach consisted of an image retrieval system which used an ensemble of
ifve DenseNet-201, each of which retrieves 100 diferent images. Then CUIs appearing in
at least 30% of the images are taken, and finally a union of each model’s predicted CUIs
is assigned to each image.
          </p>
          <p>
            CSIRO The CSIRO group reached a primary F1-score of 0.4471 [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] and a secondary
F1score of 0.7936. They experimented with a range of diferent backbones for multi-label
classification system, and their best approach is an ensemble of 43 DenseNet-161 with
top-1% threshold optimisation.
eecs-kth The eecs-kth team reached a primary F1 score of 0.4360 [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ] and a secondary F1
score of 0.8546. Their best approach utilized a multi-label classification system based on
DenseNet161 with a single classification layer.
vcmi The VCMI (vcmi) team reached a primary F1-score of 0.4329 [
            <xref ref-type="bibr" rid="ref22">22</xref>
            ] and the best overall
secondary F1-score of 0.8634. They combined a multi-label classification system based
on DenseNet-121 with an information retrieval approach for their best approach, where
the retrieval system is used if the classification did not assign any labels.
          </p>
          <p>
            PoliMi The PoliMi team reached a primary F1-score of 0.4320 [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ] and a secondary F1-score
of 0.8512. They used a ResNext50-based multi-label classification system.
          </p>
          <p>
            SSNSheerinKavitha The SSN MLRG (SSNSheerinKavitha) team reached a primary F1-score of
0.4184 [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ] and a secondary F1-score of 0.6544. They employed DenseNet for multi-label
classification and an information retrieval system.
          </p>
          <p>
            IUST_NLPLAB The IUST_NLPLAB team reached a primary F1-score of 0.3981 [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ] and a
secondary F1-score of 0.6732. They used a multi-label classification model based on
ResNet for their best results.
          </p>
          <p>
            Morgan_CS The CS_Morgan (Morgan_CS) team from Morgan State University (USA) reached
a primary F1-score of 0.3520 [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] and a secondary F1-score of 0.6280. They used a fusion
of Vision Transformers for their best approach, which outperformed their multi-label
classification systems.
          </p>
          <p>
            Kdelab The Kdelab team reached a primary F1-score of 0.3104 [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ] and a secondary F1-score
of 0.4120. They exclusively experimented with image retrieval systems and their best
approach consisted of an ensemble of diferent backbone networks (DenseNet, EficientNet,
ResNet) using simple majority voting.
          </p>
          <p>
            SDVA-UCSD The SDVA-UCSD team reached a primary F1-score of 0.3079 [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ] and a
secondary F1-score of 0.5524. They used a multi-label classification system with ResNet and
DenseNet backbones.
          </p>
          <p>
            To summarize, in the concept detection subtasks, the groups used primarily multi-label
classification systems and image retrieval systems, much like in the 2021 challenge.
Multilabel classification systems outperformed retrieval-based systems for most of the teams who
experimented with both, and while the winner was a multi-label classification approach, the
second placing team with an F1-score only 0.0006 less than the winning team, used a
retrievalbased system for which they took last year’s winning approach and tuned it to include more
CUIs by reducing the threshold for the percentage of retrieved images in which the CUI had to
appear from 50% to 30% [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ].
          </p>
          <p>This year’s models for concept detection do not show an increased F1-score compared to
last year, however due to the much larger data set and number of concepts used in this year’s
challenge, this is not surprising. Comparing it to the 2020 results, where a data set of similar
size was used, the F1-scores show a clear improvement. There are no radically new approaches
used in this year’s concept detection subtask, but the teams experimented with, optimised and
re-combined many diferent existing techniques and created competitive solutions using both
multi-label classification systems and image retrieval systems.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Results for the Caption Prediction subtask</title>
        <p>
          In this sixth edition, the caption prediction subtask attracted 10 teams which submitted 72 runs.
Table 3 presents the results of the submissions.
IUST_NLPLAB The IUST_NLPLAB team presented the best model for the caption prediction
subtask. They reached a BLEU score of 0.4828, outperforming the competition by a large
margin, and a ROUGE score of 0.1422 [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Additionally, they reached the overall best
METEOR score of 0.0928. For their best run, they employed a multi-label classification
system based on ResNet50 which treats every word as a label and assigns 26 words in the
order of their probability to each image.
        </p>
        <p>
          AUEB-NLP-Group The AUEB-NLP-Group submitted the second best performing result with a
BLEU score of 0.3222 [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and a ROUGE score of 0.1664. Their best approach utilizes the
Show &amp; Tell model [34] consisting of a CNN-RNN encoder-decoder with an EficientNetB0
backbone. While they were clearly behind the BLEU score of the winners, they outscore
them in most of the other scores.
        </p>
        <p>
          CSIRO The CSIRO group reached a BLEU score of 0.3114 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and a ROUGE score of 0.1974.
        </p>
        <p>
          Additionally, they reached the overall best BERTScore of 0.6234. They experimented
with diferent encoder-to-decoder models and achieved their best scores with CvT-21 as
the encoder and DistilGPT2 as the decoder, warm-started with a MIMIC-CXR checkpoint
with a penalty for n-grams of size 3 that are repeated.
vcmi The VCMI (vcmi) team reached a BLEU score of 0.3058 [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] and a ROUGE score of 0.1738.
        </p>
        <p>
          They used a vision encoder-to-decoder system for the best results.
eecs-kth The eecs-kth team reached a BLEU score of 0.2917 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and a ROUGE score of 0.1157.
        </p>
        <p>
          They employed an information retrieval system based on AlexNet which summarizes the
captions of a number of similar images using Pegasus.
fdallaserra The CMRE-UoG (fdallaserra) group reached a BLEU score of 0.2913 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] and the
overall best ROUGE score of 0.2012. They used a CNN Transformer approach with
multi-modal (image + CUIs) input for their best results.
        </p>
        <p>
          Kdelab The Kdelab team reached a BLEU score of 0.2782 [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and a ROUGE score of 0.1584.
        </p>
        <p>Additionally, they reached the overall best CIDEr score of 0.4114 and overall best SPICE
score of 0.0512. They used an image retrieval approach with an ensemble of diferent
backbone networks for their best submission results.</p>
        <p>
          Morgan_CS The CS_Morgan (Morgan_CS) team reached a BLEU score of 0.2549 [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] and a
ROUGE score of 0.1441. They used a very similar approach as for the concept detection,
namely a fusion of Vision Transformers.
        </p>
        <p>
          MAI_ImageSem The MAI_ImageSem team reached a BLEU score of 0.2211 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] and a ROUGE
score of 0.1847. For the best results, they use pre-trained BLIP (Bootstrapping
LanguageImage Pre-training), a pre-training framework for vision-language understanding
consisting of a multi-modal encoder-decoder and a captioning and filtering module.
SSNSheerinKavitha The SSN MLRG (SSNSheerinKavitha) team reached a BLEU score of
0.1595 [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] and a ROUGE score of 0.0425. For their best run, they employed a Sparse
Auto Encoder (SAE) with a Multi-Layer Perceptron (MLP) and a Gated Recurrent Unit
(GRU).
        </p>
        <p>To summarize, in the caption prediction subtask most teams experimented with
Transformerbased architectures and image retrieval systems. Only one team used a multi-label classification
approach, and it achieved by far the best BLEU score. However, it did not score as well on most
of the other employed metrics, with the second placing team outscoring the winners in all but
the BLEU and METEOR metrics, which highlights the dificulty of evaluating caption similarity.
One metric to highlight especially is SPICE, which is specifically designed for the evaluation of
image captions. The winners scored a value of 0.0072 in this metric with the rest of the field
(except the last placing team) scoring between 0.0218 and 0.0512.</p>
        <p>Transfer Learning has frequently been used for pre-training, from a variety of diferent data
sets. As in the previous years, simpler architectures ended up yielding better results compared
to more complex ones in many instances.</p>
        <p>Similar to the concept detection, the BLEU scores in the caption prediction subtask are overall
lower compared to last year, which can be explained by the larger and more complex data set
and more varied captions. Since there was no caption prediction subtask running in 2020, no
comparable scores for a similar data set exist.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This year’s caption task of ImageCLEFmedical once again ran with both subtasks, concept
detection and caption prediction. It returned to a larger, ROCO-based data set for both challenges
after a smaller, manually annotated data set was used last year. It attracted 12 teams who
submitted 157 runs overall, a stronger participation compared to last year. For the concept
detection subtask, a secondary F1-score was introduced to distinguish manually curated concepts
from automatically generated ones. For the caption prediction, a number of additional scores
were added to better illustrate the dificulty of evaluating the quality of predicted captions. All
but one team participated in the concept detection subtask, with only two teams choosing not
to participate in the caption prediction subtask as well. Only one team used the generated
concepts as the input for the caption prediction model, most teams approached the subtasks
with separate systems. For the concept detection challenge, most teams employed multi-label
classification systems or image retrieval systems, while the caption prediction challenge was
predominantly approached using Transformer-based architectures and image retrieval systems,
with only the winning team using a multi-label classification system.</p>
      <p>The scores for both subtasks have not improved compared to the 2021 edition. However, the
larger and more complex ROCO-based data set with more concepts and more varied captions
make the scores dificult to compare. Looking at the 2020 edition, which used a similar data set,
the concept detection scores have clearly increased (there was no caption prediction subtask).</p>
      <p>For next year’s ImageCLEFmedical Caption challenge, some possible improvements include
adding more manually validated concepts like increased anatomical coverage and directionality
information, reducing recurring captions, more fine-grained CUI filters, improving the caption
pre-processing, and using a diferent primary score for the caption prediction challenge, since
the BLEU score has some disadvantages which were highlighted by this year’s caption prediction
results.</p>
      <p>What should also be addressed is how to deal with models that were pre-trained on PMC
data, because strictly speaking they have seen the real captions and can have an advantage
when some of these images appear in test data.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was partially supported by the University of Essex GCRF QR Engagement Fund
provided by Research England (grant number G026). The work of Louise Bloch and Raphael
Brüngel was partially funded by a PhD grant from the University of Applied Sciences and
Arts Dortmund (FH Dortmund), Germany. The work of Ahmad Idrissi-Yaghir and Henning
Schäfer was funded by a PhD grant from the DFG Research Training Group 2535
Knowledgeand data-based personalisation of medicine at the point of care (WisPerMed).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schaer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bromuri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2016 medical task</article-title>
          ,
          <source>in: Working Notes of CLEF 2016 (Cross Language Evaluation Forum)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>219</fpage>
          -
          <lpage>232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Eickhof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Schwall</given-names>
            ,
            <surname>A. G. S. de Herrera</surname>
          </string-name>
          , H. Müller, Overview of ImageCLEFcaption 2017 -
          <article-title>Image Caption Prediction and Concept Detection for Biomedical Images</article-title>
          , in: Working Notes of CLEF 2017 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Dublin, Ireland,
          <source>September 11-14</source>
          ,
          <year>2017</year>
          .,
          <year>2017</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-1866/invited_paper_7.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eickhof</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Andrearczyk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2018 Caption Prediction Tasks</article-title>
          , in: Working Notes of CLEF 2018 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Avignon, France,
          <source>September 10-14</source>
          ,
          <year>2018</year>
          .,
          <year>2018</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2125</volume>
          /invited_paper_4.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
          </string-name>
          , H. Müller,
          <article-title>Overview of the ImageCLEFmed 2019 Concept Detection Task</article-title>
          , in: L.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>D. E.</given-names>
          </string-name>
          <string-name>
            <surname>Losada</surname>
          </string-name>
          , H. Müller (Eds.),
          <source>Working Notes of CLEF 2019 - Conference and Labs of the Evaluation Forum, Lugano, Switzerland, September</source>
          <volume>9</volume>
          -
          <issue>12</issue>
          ,
          <year>2019</year>
          , volume
          <volume>2380</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEURWS.org,
          <year>2019</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2380</volume>
          /paper_245.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , H. Müller,
          <article-title>Overview of the ImageCLEFmed 2020 concept prediction task: Medical image understanding</article-title>
          ,
          <source>in: CLEF2020 Working Notes</source>
          , volume
          <volume>1166</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Thessaloniki, Greece,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>O.</given-names>
            <surname>Bodenreider</surname>
          </string-name>
          ,
          <article-title>The Unified Medical Language System (UMLS): integrating biomedical terminology</article-title>
          ,
          <source>Nucleic Acids Research</source>
          <volume>32</volume>
          (
          <year>2004</year>
          )
          <fpage>267</fpage>
          -
          <lpage>270</lpage>
          . doi:
          <volume>10</volume>
          .1093/nar/gkh061.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jacutprakart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEFmed 2021 concept &amp; caption prediction task</article-title>
          ,
          <source>in: CLEF2021 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Bucharest, Romania,
          <year>2021</year>
          , pp.
          <fpage>1101</fpage>
          -
          <lpage>1112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Koitka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nensa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <article-title>Radiology Objects in COntext (ROCO): A Multimodal Image Dataset</article-title>
          , in: Intravascular Imaging and Computer Assisted Stenting - and
          <string-name>
            <surname>-</surname>
          </string-name>
          Large-Scale
          <source>Annotation of Biomedical Data and Expert Label Synthesis - 7th Joint International Workshop</source>
          , CVII-STENT 2018 and Third International Workshop, LABELS 2018,
          <article-title>Held in Conjunction with MICCAI 2018, Granada</article-title>
          , Spain,
          <year>September 16</year>
          ,
          <year>2018</year>
          , Proceedings,
          <year>2018</year>
          , pp.
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -01364-6\_
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Péteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bloch</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Brüngel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Idrissi-Yaghir</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Schäfer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kozlovski</surname>
            ,
            <given-names>Y. D.</given-names>
          </string-name>
          <string-name>
            <surname>Cid</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Kovalev</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-D. Ştefan</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Dogariu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Deshayes-Chossart</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Schindler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chamberlain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2022: Multimedia retrieval in medical, social media and nature applications, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 13th International Conference of the CLEF Association (CLEF</source>
          <year>2022</year>
          ),
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Charalampakos</surname>
          </string-name>
          , G. Zachariadis,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pavlopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Karatzas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Trakas</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Androutsopoulos</surname>
          </string-name>
          , AUEB NLP group at ImageCLEFmed caption
          <year>2022</year>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lebrat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nicolson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Cruz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Belous</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          , J. Dowling, CSIRO at ImageCLEFmed caption
          <year>2022</year>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Moschovis</surname>
          </string-name>
          , E. Fransén, Neuraldynamicslab at ImageCLEF medical
          <year>2022</year>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F. D.</given-names>
            <surname>Serra1</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Deligianni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dalton</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Q.</surname>
          </string-name>
          <article-title>O'Neil, CMRE-UoG team at ImageCLEFmed caption 2022 task: Concept detection and image captioning</article-title>
          ,
          <source>in: CLEF2022 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hajihosseini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lotfollahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nobakhtian</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Javid</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Omidi</surname>
          </string-name>
          , S. Eetemadi, IUST_NLPLAB at ImageCLEFmed caption tasks
          <year>2022</year>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>R.</given-names>
            <surname>Tsuneda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Asakawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shimizu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Komoda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aono</surname>
          </string-name>
          , Kdelab at ImageCLEF 2022:
          <article-title>Medical concept detection with image retrieval and code ensemble</article-title>
          ,
          <source>in: CLEF2022 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Tsuneda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Asakawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shimizu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Komoda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aono</surname>
          </string-name>
          ,
          <article-title>Kdelab at ImageCLEF2022 medical caption prediction task</article-title>
          ,
          <source>in: CLEF2022 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          , ImageSem Group at ImageCLEFmed Caption 2022 Task:
          <article-title>Generating Medical Image Descriptions based on Visual- Language Pre-training</article-title>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>M. M. Rahman</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <string-name>
            <surname>Layode</surname>
          </string-name>
          , CS_Morgan at ImageCLEFmed caption 2022:
          <article-title>Deep learning based multilabel classification and transformers for concept detection &amp; caption prediction</article-title>
          ,
          <source>in: CLEF2022 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S. A. M.</given-names>
            <surname>Ghayyomnia</surname>
          </string-name>
          , K. de Gast,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Carmana</surname>
          </string-name>
          , Polimi-imageclef group at ImageCLEFmed caption
          <year>2022</year>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gentili</surname>
          </string-name>
          ,
          <article-title>ImageCLEFmed concept detection, finding duplicates</article-title>
          ,
          <source>in: CLEF2022 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>N. M. S.</given-names>
            <surname>Sitara</surname>
          </string-name>
          , S. Kavitha, SSN MLRG at ImageCLEF 2022:
          <article-title>Medical concept detection and caption prediction using transfer learning and transformer based learning approaches</article-title>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>I.</given-names>
            <surname>Rio-Torto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Patrício</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Montenegro</surname>
          </string-name>
          , T. Gonçalves,
          <article-title>Detecting Concepts and Generating Captions from Medical Images: Contributions of the VCMI Team to ImageCLEFmed Caption 2022</article-title>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Andrzejowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. K.</given-names>
            <surname>Kanakaris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. V.</given-names>
            <surname>Giannoudis</surname>
          </string-name>
          , Pelvic Girdle Pain, Hypermo-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>