<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Overview of ImageCLEFmedical 2023 - Caption Prediction and Concept Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Johannes Rückert</string-name>
          <email>johannes.rueckert@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asma Ben Abacha</string-name>
          <email>abenabacha@microsoft.com</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alba G. Seco de Herrera</string-name>
          <email>alba.garcia@essex.ac.uk</email>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Louise Bloch</string-name>
          <email>louise.bloch@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raphael Brüngel</string-name>
          <email>raphael.bruengel@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ahmad Idrissi-Yaghir</string-name>
          <email>ahmad.idrissi-yaghir@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henning Schäfer</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henning Müller</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
          <xref ref-type="aff" rid="aff7">7</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph M. Friedrich</string-name>
          <email>christoph.friedrich@fh-dortmund.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Applied Sciences and Arts Dortmund</institution>
          ,
          <addr-line>Dortmund</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Artificial Intelligence in Medicine (IKIM), University Hospital Essen</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Medical Informatics, Biometry and Epidemiology (IMIBE), University Hospital Essen</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Institute for Transfusion Medicine, University Hospital Essen</institution>
          ,
          <addr-line>Essen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Microsoft</institution>
          ,
          <addr-line>Redmond, Washington</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Applied Sciences Western Switzerland (HES-SO)</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University of Essex</institution>
          ,
          <addr-line>Wivenhoe Park, Colchester CO4 3SQ</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff7">
          <label>7</label>
          <institution>University of Geneva</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>The ImageCLEFmedical 2023 Caption task on caption prediction and concept detection follows similar challenges held from 2017-2022. The goal is to extract Unified Medical Language System (UMLS) concept annotations and/or define captions from image data. Predictions are compared to original image captions. Images for both tasks are part of the Radiology Objects in COntext version 2 (ROCOv2) dataset. For concept detection, multi-label predictions are compared against UMLS terms extracted from the original captions with additional manually curated concepts via the F1-score. For caption prediction, the semantic similarity of the predictions to the original captions is evaluated using the BERTScore. The task attracted strong participation with 27 registered teams, 13 teams submitted 116 graded runs for the two subtasks. Participants mainly used multi-label classification systems for the concept detection subtask, the winning team AUEB-NLP-Group used an ensemble of three CNNs. For the caption prediction subtask, most teams used encoder-decoder architectures, with the winning team CSIRO using an encoder-decoder framework with an additional reinforcement learning optimization step.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;ImageCLEF</kwd>
        <kwd>Computer Vision</kwd>
        <kwd>Multi-Label Classification</kwd>
        <kwd>Image Captioning</kwd>
        <kwd>Image Understanding</kwd>
        <kwd>Radiology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>ImageCLEF1 is the image retrieval and classification lab of the CLEF (Conference and Labs
of the Evaluation Forum) conference. ImageCLEF 2023 consists of the ImageCLEFmedical,
ImageCLEFfusion and ImageCLEFaware labs, with the ImageCLEFmedical lab being divided into
the subtasks MEDIQA-Sum (natural language semantic retrieval), Caption, GANs (generation
of medical images), and MEDVQA-GI (gastrointestinal visual question answering).</p>
      <p>
        The Caption task was first proposed as part of the ImageCLEFmedical [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] in 2016. In 2017
and 2018 [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] the ImageCLEFmedical caption task comprised two subtasks: concept detection
and caption prediction. In 2019 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and 2020 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the task concentrated on the concept detection
subtask extracting Unified Medical Language System ® (UMLS) Concept Unique Identifiers
(CUIs) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] from radiology images.
      </p>
      <p>
        In 2021 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], both subtasks, concept detection and caption prediction, were running again due
to participants demands. The focus in 2021 was on making the task more realistic by using
fewer images which were all manually annotated by medical doctors. As additional data of
similar quality is hard to acquire, the 2022 ImageCLEFmedical caption task [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] continued with
both subtasks albeit with an extended version of the Radiology Objects in COntext (ROCO) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
dataset used for both subtasks, which was already used in 2020 and 2019. The 2023 edition of
ImageCLEFmedical caption continues in the same vein, once again using a ROCO-based dataset
for both subtasks but switching from BLEU to BERTScore as the primary evaluation metric for
caption prediction.
      </p>
      <p>
        This paper sets forth the approaches for the caption task: automated cross-referencing of
medical images and captions into predicted coherent captions and UMLS concept detection in
radiology images as a separate subtask. This task is a part of the ImageCLEF benchmarking
campaign, which has proposed medical image understanding tasks since 2003; a new suite of
tasks is generated each subsequent year. Further information on the other proposed tasks at
ImageCLEF 2023 can be found in Ionescu et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        This is the 7th edition of the ImageCLEFmedical caption task. Just like in 2016 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], 2017 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
2018 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], 2021 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and 2022 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] both subtasks of concept detection and caption prediction are
included in ImageCLEFmedical Caption 2023. Like in 2022, an extended subset of the ROCO [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
dataset is used, with images that are not licensed CC BY or CC BY-NC removed.
      </p>
      <p>Manual generation of the knowledge of medical images is a time-consuming process prone to
human error. As this process requires assistance for the better and easier diagnoses of diseases
that are susceptible to radiology screening, it is important that we better understand and refine
automatic systems that aid in the broad task of radiology-image metadata generation. The
purpose of the ImageCLEFmedical 2023 caption prediction and concept detection tasks is the
continued evaluation of such systems. Concept detection and caption prediction information is
applicable to unlabelled and unstructured datasets and medical datasets that do not have textual
metadata. The ImageCLEFmedical caption task focuses on the medical image understanding in
the biomedical literature and specifically on concept extraction and caption prediction based on
the visual perception of the medical images and medical text data such as medical caption or
UMLS CUIs paired with each image (see Figure 1).</p>
      <sec id="sec-1-1">
        <title>1https://www.imageclef.org/ [last accessed: 2023-06-28]</title>
        <p>
          In 2023, for the development data, an extended subset of the ROCO [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] dataset from 2022
was used, with new images from the same source added for the validation and test sets, while
images from articles with licenses other than CC BY and CC BY-NC were removed.
        </p>
        <p>This paper presents an overview of the ImageCLEFmedical caption task 2023 including the
task and participation in Section 2, the data creation in Section 3, and the evaluation methodology
in Section 4. The results are described in Section 5, followed by conclusion in Sections 6.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Task and Participation</title>
      <p>In 2023, the ImageCLEFmedical caption task consisted of two subtasks: concept detection and
caption prediction.</p>
      <p>
        The concept detection subtask follows the same format proposed since the start of the task
in 2017 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Participants are asked to predict a set of concepts defined by the UMLS CUIs [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
based on the visual information provided by the radiology images.
      </p>
      <p>
        The caption prediction subtask follows the original format of the subtask used between
2017 and 2018 [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. This subtask was paused and it is running again since 2021 because of
participant demand. This subtask aims to automatically generate captions for the radiology
images provided.
      </p>
      <p>
        In 2023, 27 teams registered and signed the End-User-Agreement that is needed to download
the development data. 13 teams submitted 116 graded runs for evaluation (12 teams submitted
working notes) attracting similar attention than in 2022 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Each of the groups was allowed a
maximum of 10 graded runs per subtask.
      </p>
      <p>
        Table 1 shows all the teams who participated in the task and their submitted runs. 9 teams
participated in the concept detection subtask this year, 6 of those teams also participated in 2022
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. 13 teams submitted runs to the caption prediction subtask, 7 of those teams also participated
in 2022. Three of the teams participated also in 2021. Overall, 9 teams participated in both
subtasks, and four teams participated only in the caption prediction subtask. Unlike in 2022, no
teams participated only in the concept detection subtask.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Data Creation</title>
      <p>Figure 1 shows an example from the dataset provided by the task.</p>
      <p>Like last year, a dataset that originates from biomedical articles of the PMC Open Access
Subset2 [22] was used and was extended with new images added since the last time the dataset
was updated in October 2021. The overall lower number of images is due to the removal of
non-CC BY images (including CC BY-SA and CC BY-ND).</p>
      <p>Unlike last year, no extensive caption pre-processing beyond the removal of links was
performed to keep the captions as realistic as possible. Captions in languages other than English
were also removed.</p>
      <p>From the resulting captions, concepts were extracted using the Medical Concept Annotation
Toolkit (MedCAT) [23]. MedCAT, which is capable of extracting biomedical concepts from</p>
      <sec id="sec-3-1">
        <title>2https://www.ncbi.nlm.nih.gov/pmc/tools/openftlist/ [last accessed: 2023-06-17]</title>
        <p>Institution
Department of Informatics, Athens University of
Economics and Business, Athens, Greece
Toyohashi University of Technology, Aichi, Japan
and Toyohashi Heart Center, Aichi, Japan
SSN College Of Engineering, Chennai, India</p>
        <p>Baidu Intelligent Health Unit, Beijing, China and</p>
        <p>
          Peng Cheng Laboratory, Shenzhen, China
CS_Morgan* [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] Computer Science Department, Morgan State
        </p>
        <p>
          University, Baltimore, Maryland
CSIRO* [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] Australian e-Health Research Centre,
        </p>
        <p>Commonwealth Scientific and Industrial Research
Organisation, Herston, Queensland, Australia and
CSIRO Data61, Imaging and Computer Vision
Group, Pullenvale, Queensland, Australia and
Queensland University of Technology, Brisbane,</p>
        <p>
          Queensland, Australia
IUST_NLPLAB* [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] School of Computer Engineering, Iran University of
        </p>
        <p>Science and Technology, Tehran, Islamic Republic</p>
        <p>
          Of Iran
KDE-Lab_Med* [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] KDE Laboratory, Department of Computer Science
and Engineering, Toyohashi University of
        </p>
        <p>
          Technology, Aichi, Japan
PCLmed [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] Peng Cheng Laboratory, Shenzhen, China and
        </p>
        <p>ADSPLAB, School of Electronic and Computer</p>
        <p>
          Engineering, Peking University, Shenzhen, China
SSN_MLRG [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] Department of CSE, Sri Sivasubramaniya Nadar
        </p>
        <p>
          College of Engineering, India
SSNSheerinKavitha*[
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] Department of CSE, Sri Sivasubramaniya Nadar
        </p>
        <p>College of Engineering, India
VCMI* [21] University of Porto, Porto, Portugal and INESC TEC,</p>
        <p>Porto, Portugal
–
1
3
5
–
7
10
–
5
2
10
2
1
8
10
4
10
10
5
1
3
7
unstructured text, was trained on the MIMIC-III dataset [24] and links to SNOMED CT IDs,
which were later mapped to CUIs and TUIs of the UMLS2022AB release3. During concept
extraction, concepts were retained only if they exceeded a frequency threshold of 10 occurences,
and semantic filters were applied to focus on visually observable and interpretable concepts.
For example, concepts of semantic type T029 (Body Location or Region) or T060 (Diagnostic
Procedure) are relevant, while concepts of semantic type T054 (Social Behavior) cannot be
derived from the image if it would appear in the caption. In addition, manual filtering was
3https://www.nlm.nih.gov/pubs/techbull/nd22/nd22_umls_2022ab_release_available.html [last accessed: 2023-06-17]
UMLS CUI
C1306645
Caption: Anteroposterior pelvic radiograph of a 30-year-old female diagnosed with Ehlers-Danlos Syndrome
demonstrating fusion of pubic symphysis and both sacroiliac joints (anterior plating, bone grafting and
sacroiliac screw insertion)
performed to exclude UMLS concepts that were either incorrectly detected by the pipeline
or were still not related to the image content in any way after semantic filtering. Blacklisted
concepts often include qualifiers that would divert actual interest to, for example, anatomical
localization or a pathological process, and would also introduce bias, since qualifiers are used
in a highly individual and variable manner. Entity linking systems tend to link concepts with
ambiguous synonyms incorrectly, e.g. C0994894 (Patch Dosage Form) may be linked if the
caption refers to a region that is patchy. In case of high frequency occurrence of such concepts,
they were merged to the correct concept via mapping. Due to the diferent filtering approach,
this year’s dataset contains 2,125 concepts compared to 8,374 last year.</p>
        <p>
          Additional concepts were assigned to all images addressing their image modality. Six medical
image modalities of concepts were covered: X-ray, Computer Tomography (CT), Magnetic
Resonance Imaging (MRI), ultrasound, and Positron Emission Tomography (PET) as well as
modality combinations (e.g., PET/CT) as standalone concept. For images of the X-ray modality
further concepts on the represented anatomy were assigned, covering specific anatomical body
regions of the Image Retrieval in Medical Application (IRMA) [25] classification: cranium, spine,
upper extremity/arm, chest, breast/mamma, abdomen, pelvis, and lower extremity/leg. New
for this year’s dataset is the addition of manually validated directionality concepts for x-ray
images. Directionality refers to the x-ray imaging orientation according to IRMA: coronal
posteroanterior (PA), coronal anteroposterior (AP), sagittal, or transversal. Each of the described
concept extensions were created performing a two-stage process. In the first stage, predictions
via classification models were created and assigned as annotations. For modality prediction
for all images a model trained on the ROCO dataset [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], and for anatomy prediction for X-ray
modality images a model trained on an existing IRMA-annotated image dataset [26] was used.
For directionality, roughly 20,000 images were manually annotated to train an initial classifier.
In the second stage, these annotations underwent manual quality control measures, involving
correction of faulty predictions and filtering of images that did not represent one of the minded
modality or anatomy concepts.
        </p>
        <p>The following subsets were distributed to the participants where each image has one caption
and one or more concepts (UMLS-CUI):
• Training set including 60,918 radiology images and associated captions and concepts, with
a total of 263,091 concept occurrences and 2,125 unique concepts.
• Validation set including 10,437 radiology images and associated captions and concepts,
with a total of 46,584 concept occurrences and 1,945 unique concepts.
• Test set including 10,473 radiology images, with a total of 46,955 concept occurrences and
1,936 unique concepts.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation Methodology</title>
      <p>In this year’s edition, the performance evaluation for the concept detection subtask is carried
out in the same way as last year, while the primary evaluation metric for the caption prediction
subtask is changed from BLEU to BERTScore. Both tasks are evaluated separately. AIcrowd
was not used as a challenge platform this year, instead participants were asked to upload their
submissions to a cloud share file drop, with information about whether each submission was
successfully evaluated and announced on a website that was regularly updated. An important
diference to the last years was the fact that participants were unaware of their own scores on
the test set until after the submission deadline. This was done to avoid teams optimizing their
approaches based on test set results, which would amount to information leakage.</p>
      <p>For the concept detection subtask, the balanced precision and recall trade-of were measured
in terms of F1-scores. Like last year, a secondary F1-score is computed using a subset of concepts
that was manually curated. On the one hand, this involves the diferent image modalities (X-ray,
Angiography, Ultrasound, CT, MRI, PET, and Combined such as PET/CT). On the other hand, if
applicable, for X-ray also the most prominently depicted body region (cranium, chest, upper
extremity, spine, abdomen, pelvis, and lower extremity), and the capture directionality (coronal
anteroposterior, coronal posteroanterior, sagittal, and transversal) were involved.</p>
      <p>As a pre-processing step for evaluating the second task, all captions were lowercased,
punctuation was removed, and numbers were replaced by the token “number”. This step ensures
uniformity and focuses the evaluation on the linguistic content. The performance of caption
prediction is evaluated based on BERTScore [27], which is a metric that computes a similarity score
for each token in the generated text with each token in the reference text. It uses the pre-trained
contextual embeddings from BERT-based models and matches words by cosine similarity. In this
work, the pre-trained model microsoft/deberta-xlarge-mnli4 was used because it is the model that
correlates best with human scoring according to the authors5. Since evaluating generated text
and image captioning is very challenging and should not be based on a single metric, additional
evaluation metrics were explored in this year’s edition in order to find the metrics that correlate</p>
      <sec id="sec-4-1">
        <title>4https://huggingface.co/microsoft/deberta-xlarge-mnli [last accessed: 2023-06-17]</title>
        <p>5https://github.com/Tiiiger/bert_score [last accessed: 2023-06-17]
well with human judgments for this task. First, the Recall-Oriented Understudy for Gisting
Evaluation (ROUGE) [28] score was adopted as a secondary metric that counts the number of
overlapping units such as n-grams, word sequences, and word pairs between the generated text
and the reference. Specifically, the ROUGE-1 (F-measure) score was calculated, which measures
the number of matching unigrams between the model-generated text and a reference. All
individual scores for each caption are then summed and averaged over the number of captions,
resulting in the final score. In addition to ROUGE, the Metric for Evaluation of Translation with
Explicit ORdering (METEOR) [29] was explored, which is a metric that evaluates the generated
text by aligning it to reference and calculating a sentence-level similarity score. Furthermore,
the Consensus-based Image Description Evaluation (CIDEr) [30] metric was also adopted. CIDEr
is an automatic evaluation metric that calculates the weights of n-grams in the generated text,
and the reference text based on term frequency and inverse document frequency (TF-IDF) and
then compares them based on cosine similarity. Another metric used is the BiLingual Evaluation
Understudy (BLEU) score [31], which is a geometric mean of n-gram scores from 1 to 4. For this
task, the focus was on the BLEU-1 score, which takes into account unigram precision. Compared
to last year, BLEURT and CLIPScore were newly introduced. BLEURT (Bilingual Evaluation
Understudy with Representations from Transformers.) [32] is specifically designed to evaluate
natural language generation in English. It uses a pre-trained model that has been fine-tuned to
emulate human judgments about the quality of the generated text. The strength of BLEURT lies
in its end-to-end training, which enables it to model human judgments efectively and makes
it robust to domain and quality variations. For this evaluation, the BLEURT-20 model was
used. CLIPScore [33] is an innovative metric that diverges from the traditional reference-based
evaluations of image captions. Instead, it aligns with the human approach of evaluating caption
quality without references by evaluating the alignment between text and image content. The
metric employs CLIP (Contrastive Language-Image Pretraining) [34], a cross-modal model that
has been pre-trained on a massive dataset of 400 million image-caption pairs sourced from the
web. The model is used to compute similarity scores between images and text. The introduction
of BLEURT and CLIPScore in this edition aims to further align the evaluation process with
human judgment.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>For the concept detection and caption prediction subtasks, Tables 2 and 3 show the best results
from each of the participating teams. The results will be discussed in this section. The full list
of results are shown in appendix A in tables 5, 6 and 7.</p>
      <sec id="sec-5-1">
        <title>5.1. Results for the Concept Detection subtask</title>
        <p>In 2023, 9 teams participated in the concept prediction subtask, submitting 47 graded runs.
Table 2 presents the results achieved in the submissions.</p>
        <p>
          AUEB-NLP-Group Like in previous years, the AUEB-NLP-Group submitted the best
performing result with a primary F1-score of 0.5223 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and a secondary F1-score of 0.9258. The
winning approach was an ensemble of three CNNs (EficientNetB0, DenseNet121, and
EficientNetB0v2) followed by a feed-forward neural network (FFNN) classification head,
which is a very similar approach as last year [35], where an ensemble of two such models
won the concept detection subtask. They also experimented with training separate models
for the diferent modalities, which did not lead to better results.
        </p>
        <p>
          KDE-Lab_Med The KDE-Lab_Med team submitted the second best performing approach,
with a primary F1-score of 0.5074 [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] and a secondary F1-score of 0.9321, which was
the highest overall secondary F1-score. Their best approach is a single CNN+FFNN model
with an EficientNetV2-M backbone. They experimented with image pre-processing by
either converting color to grayscale or colorization of grayscale images by stacking color
channels. The latter approach performed better.
        </p>
        <p>VCMI The VCMI team achieved the third place in the concept detection subtask with a primary
F1-score of 0.4998 [21] and a secondary F1-score of 0.9162. Their best approach utilizes
an autoregressive multi-label classification system with a VGG16 network pre-trained
on ImageNet, which instead of using a single classification layer at the end, uses 17
classification layers each predicting 125 concepts. For any images that are not assigned
any concepts using this model, an image retrieval system assigns concepts appearing in
at least two of the four most similar images in the training data.</p>
        <p>
          IUST_NLPLAB The IUST_NLPLAB team reached a primary F1-score of 0.4959 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] and a
secondary F1-score of 0.8804. They used a multi-label classification system based on the
vision-language model PubMedCLIP for their best results.
        </p>
        <p>
          Clef-CSE-GAN-Team The Clef-CSE-GAN-Team achieved a primary F1-score of 0.4957 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]
and a secondary F1-score of 0.9106. They employed a multi-label classification system
with a DenseNet121 backbone.
        </p>
        <p>
          CS_Morgan The CS_Morgan team reached a primary F1-score of 0.4834 [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and a secondary
F1-score of 0.8902. Their best approach used a multi-label classification system with a
DenseNet121 backbone using CheXNet pre-trained weights.
        </p>
        <p>
          SSN_MLRG and SSNSheerinKavitha The team SSN_MLRG achieved a primary F1-score of
0.4649 [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and a secondary F1-score of 0.8603. They employed a multi-label classification
system using ConceptNet.
        </p>
        <p>
          To summarize, in the concept detection subtasks, the groups used primarily multi-label
classification systems, with image retrieval systems consistently performing worse for teams
who experimented with them. One team successfully used an image retrieval system as a
fallback when the multi-label classification system did not predict any concepts [ 21]. As in 2022,
the AUEB-NLP-Group once again achieved the top scores by increasing their ensemble from
two to three models [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>The overall F1 scores increased compared to last year which is not surprising considering a
reduced number of concepts for this year’s edition of the challenge.</p>
        <p>While one team experimented with a novel autoregressive multi-label classification system
which tries to model relationships between concepts and another team tried training separate
models for the diferent modalities, these experiments did not yield better results compared to
the winning approach.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Results for the Caption Prediction subtask</title>
        <p>
          In this seventh edition, the caption prediction subtask attracted 13 teams which submitted 69
graded runs. Tables 3 and 4 present the results of the submissions.
CSIRO The CSIRO team achieved first place in the caption prediction subtask with a BERTScore
of 0.6413 [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] and a ROUGE score of 0.2463. The winning approach, which also reached
the highest CLIPScore, consists of an encoder-decoder framework based on the
Convolutional Vision Transformer (CVT) as the encoder and DistilGPT2 as the decoder. This
approach was already used by them in last year’s addition and reached the overall highest
BERTScore then. For this year, they added a reinforcement learning step "to optimize the
model for the primary metric and the means of conditioning the decoder on the visual
features" [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], which further improved the performance and set them apart from the
competition.
closeAI The closeAI team reached the second place spot with a BERTScore of 0.6281 [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and
a ROUGE score of 0.2401. Their approach, which reached top scores in the BLEURT
and CIDEr metrics, consisted of a BLIP-2 framework with a ViT-g image encoder from
EVA-CLIP, a Q-Former and OPT2.7 as the LLM with post-processing to remove duplicate
content from the generated captions.
        </p>
        <p>
          AUEB-NLP-Group The AUEB-NLP-Group reached a BERTScore of 0.6170 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and a ROUGE
score of 0.2130, placing third. Their best approach is a novel captioning pipeline using a
denoising model to rewrite captions produced by a CNN-RNN encoder decoder model
using sequence to sequence models BART and T5.
        </p>
        <p>
          PCLmed The PCLmed team achieved a BERTScore of 0.6152 [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] and a ROUGE score of
0.2528. Much like closeAI, they used a BLIP-2 framework with an EVA-ViT-g encoder,
a Query Transformer, and ChatGLM-6B as the LLM with a final beam search with a
repetition penalty to generate the captions.
        </p>
        <p>VCMI The VCMI team achieved a BERTScore of 0.6147 [21] and a ROUGE score of 0.2175.</p>
        <p>They used an encoder-decoder framework with a Data-eficient image Transformer (DeiT)
as the encoder and DistilGPT-2 as the decoder for their best results.</p>
        <p>
          KDE-Lab_Med The KDE-Lab_Med team reached a BERTScore of 0.6145 [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] and a ROUGE
score of 0.2223. Their best approach is a CNN-RNN system based on Show, Attend,
and Tell with a ResNet152 backbone and LSTM as RNN. They also experimented with a
Caption Transformer system, which did not perform better.
        </p>
        <p>
          SSN_MLRG and SSNSheerinKavitha The team SSN_MLRG achieved a BERTScore of
0.6019 [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and a ROUGE score of 0.2112. They used an encoder-decoder system with
DeiT as the encoder and Distilled-GPT2 as the decoder.
        </p>
        <p>DLNU_CCSE The DLNU_CCSE team reached a BERTScore of 0.6005 and a ROUGE score of
0.2029. They used an encoder-decoder framework with a ResNet-101 encoder and an
LSTM decoder for their best approach. They did not submit working notes.</p>
        <p>
          CS_Morgan The CS_Morgan team reached a BERTScore of 0.5819 [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and a ROUGE score
of 0.1564. They used an encoder-decoder system with a Vision Transformer (ViT) as
the encoder, where the decoder generates keywords which are then transformed into
captions by a T5 generative model fine-tuned for this purpose.
        </p>
        <p>
          Clef-CSE-GAN-Team The Clef-CSE-GAN-Team achieved a BERTScore of 0.5816 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] and
a ROUGE score of 0.2181. They used an encoder-decoder approach with a ResNet101
encoder and an LSTM decoder for their best results.
        </p>
        <p>
          Bluefield The Bluefield team reached a BERTScore of 0.5780 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and a ROUGE score of
0.1534. They first classified the images into six groups roughly corresponding to the
imaging modalities, and then used six diferent CLIP models with a ResNet50 backbone
to generate captions for the images.
        </p>
        <p>
          IUST_NLPLAB Last years winners reached a BERTScore of 0.5669 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] and the overall best
ROUGE score of 0.2898. Like last year, they used a multi-label classification approach
where the top 20 words words are returned as the caption in the order of their probability.
This system once again achieved top scores in the ROUGE, BLEU, and METEOR scores,
but did not perform as well on the remaining metrics, including the primary metric
BERTScore.
        </p>
        <p>To summarize, in the caption prediction subtask most teams experimented with
encoderdecoder frameworks with diferent backbones and LSTM decoders. Unsurprisingly, teams
increasingly used LLMs in the decoding step and to help generate or refine captions. BLIP-2 was
used for the first time and achieved good results (second and fourth place). One novelty was
the use of reinforcement learning to refine and improve upon last year’s best solution in terms
of BERTScore, which ended up winning this year’s competition after the change of primary
scores from BLEU to BERTScore.</p>
        <p>The aforementioned change of evaluation metrics had a big efect on the outcome of the
challenge, with last year’s winner placing second to last according to the BERTScore evaluation
while still winning in terms of the ROUGE, BLEU and METEOR scores with a similar approach
as last year. We will continue to evaluate and explore diferent possible metrics or combination
of metrics, but the evaluation of generated captions remains dificult.</p>
        <p>BERTScore and ROUGE scores were used to predict captions. Unlike the previous edition,
BERTScore replaced BLEU as the primary score for a more refined evaluation of the caption task.
The adoption of BERTScore reflects the intent to prioritize semantic alignment and information
preservation in the generated captions rather than focusing on the frequency of n-gram matches,
which is the basis of BLEU.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This year’s caption task of ImageCLEFmedical once again ran with both subtasks, concept
detection and caption prediction. It once again used a ROCO-based dataset with additional
manual annotations for X-ray directionality. It attracted 13 teams who submitted 116 graded
runs using a cloud file drop instead of the AICrowd platform, which was not available to be used
this year. For the concept detection task, the F1-score and a secondary F1-score, considering
only the manually curated concepts, were used. After adding a number of additional metrics to
the caption prediction task last year, the primary metric was changed from BLEU to BERTScore
for this year, hoping to reward semantic similarity instead of just n-gram overlap. The caption
prediction subtask was more popular than the concept detection subtask this year, with all 9
teams participating in both subtasks, and four teams participating only in the caption prediction
subtask. As before, the teams generally approached the tasks completely separately, not really
making use of generated concepts for the predicted captions. Like last year, teams generally
used multi-label classification systems for the concept detection subtask, last year’s winning
team simply scaling up their approach to use three instead of two ensembles to once again
reach top scores. Retrieval-based systems were still used by some teams, but were consistently
outperformed by multi-label classification systems. For the caption prediction subtask,
encoderdecoder frameworks were used by most teams, with LLMs being used to generate or refine
the captions by some teams. BLIP-2 was used for the first time and achieved good results.
Reinforcement learning helped last year’s top scoring team in terms of BERTScore further
increase last year’s score and take the top spot.</p>
      <p>The scores for both tasks have improved compared to last year. For the concept detection
subtask, this is partly due to the decreased number of concepts. For caption prediction, BERTScore
and ROUGE scores have improved, illustrating the beneficial shift to BERTScore as the
primary metric, which emphasizes semantic alignment and information preservation over n-gram
frequency.</p>
      <p>For next year’s ImageCLEFmedical Caption challenge, some possible improvements include an
improved caption prediction evaluation metric which is specific to medical texts, and improving
manually validated concept quality with the help of a medical professional. It will also be
important to make sure that no models are used that were pre-trained on PubMedCentral data,
since these models will already have seen the original captions.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was partially supported by the University of Essex GCRF QR Engagement Fund
provided by Research England (grant number G026). The work of Louise Bloch and Raphael
Brüngel was partially funded by a PhD grant from the University of Applied Sciences and
Arts Dortmund (FH Dortmund), Germany. The work of Ahmad Idrissi-Yaghir and Henning
Schäfer was funded by a PhD grant from the DFG Research Training Group 2535
Knowledgeand data-based personalisation of medicine at the point of care (WisPerMed).
Notes, CEUR Workshop Proceedings, CEUR-WS.org, Thessaloniki, Greece, 2023.
[21] I. Rio-Torto, C. Patrício, H. Montenegro, T. Gonçalves, J. S. Cardoso, Detecting concepts
and generating captions from medical images: Contributions of the VCMI team to
ImageCLEFmedical caption 2023, in: CLEF2023 Working Notes, CEUR Workshop Proceedings,
CEUR-WS.org, Thessaloniki, Greece, 2023.
[22] R. J. Roberts, PubMed Central: The GenBank of the published literature, Proceedings
of the National Academy of Sciences of the United States of America 98 (2001) 381–382.
doi:10.1073/pnas.98.2.381.
[23] Multi-domain clinical natural language processing with medcat: The medical concept
annotation toolkit, Artificial Intelligence in Medicine 117 (2021) 102083. doi: https:
//doi.org/10.1016/j.artmed.2021.102083.
[24] A. E. Johnson, T. J. Pollard, L. Shen, L. wei H. Lehman, M. Feng, M. Ghassemi, B. Moody,
P. Szolovits, L. A. Celi, R. G. Mark, MIMIC-III, a freely accessible critical care database,
Scientific Data 3 (2016). URL: https://doi.org/10.1038/sdata.2016.35. doi: 10.1038/sdata.
2016.35.
[25] T. M. Lehmann, H. Schubert, D. Keysers, M. Kohnen, B. B. Wein, The IRMA code for unique
classification of medical images, in: H. K. Huang, O. M. Ratib (Eds.), Medical Imaging 2003:
PACS and Integrated Medical Information Systems: Design and Evaluation, SPIE, 2003.
doi:10.1117/12.480677.
[26] T. Deserno, B. Ott, 15.363 IRMA Bilder in 193 Kategorien für ImageCLEFmed
2009, 2009. URL: https://publications.rwth-aachen.de/record/667225. doi:10.18154/
RWTH-2016-06143.
[27] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, Y. Artzi, Bertscore: Evaluating text
generation with BERT, in: 8th International Conference on Learning Representations,
ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020, 2020. URL: https://openreview.net/
forum?id=SkeHuCVFDr.
[28] C.-Y. Lin, ROUGE: A Package for Automatic Evaluation of Summaries, in: Text
Summarization Branches Out, Association for Computational Linguistics, 2004, pp. 74–81. URL:
https://aclanthology.org/W04-1013.
[29] M. Denkowski, A. Lavie, Meteor Universal: Language Specific Translation Evaluation
for Any Target Language, in: Proceedings of the Ninth Workshop on Statistical Machine
Translation, Association for Computational Linguistics, 2014, pp. 376–380. URL: http:
//aclweb.org/anthology/W14-3348. doi:10.3115/v1/W14-3348.
[30] R. Vedantam, C. L. Zitnick, D. Parikh, CIDEr: Consensus-based image description
evaluation, in: 2015 IEEE Conference on Computer Vision and Pattern Recognition
(CVPR), IEEE, 2015, pp. 4566–4575. URL: http://ieeexplore.ieee.org/document/7299087/.
doi:10.1109/CVPR.2015.7299087.
[31] K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, BLEU: a method for automatic evaluation of
machine translation, in: Proceedings of the 40th annual meeting of the Association for
Computational Linguistics, 2002, pp. 311–318.
[32] T. Sellam, D. Das, A. Parikh, BLEURT: Learning robust metrics for text generation, in:
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics,
Association for Computational Linguistics, Online, 2020, pp. 7881–7892. URL: https://
aclanthology.org/2020.acl-main.704. doi:10.18653/v1/2020.acl-main.704.
[33] J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, Y. Choi, CLIPScore: A reference-free
evaluation metric for image captioning, in: Proceedings of the 2021 Conference on
Empirical Methods in Natural Language Processing, Association for Computational
Linguistics, Online and Punta Cana, Dominican Republic, 2021, pp. 7514–7528. URL: https:
//aclanthology.org/2021.emnlp-main.595. doi:10.18653/v1/2021.emnlp-main.595.
[34] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell,
P. Mishkin, J. Clark, G. Krueger, I. Sutskever, Learning transferable visual models from
natural language supervision, in: M. Meila, T. Zhang (Eds.), Proceedings of the 38th
International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event,
volume 139 of Proceedings of Machine Learning Research, PMLR, 2021, pp. 8748–8763. URL:
http://proceedings.mlr.press/v139/radford21a.html.
[35] F. Charalampakos, G. Zachariadis, J. Pavlopoulos, V. Karatzas, C. Trakas, I.
Androutsopoulos, AUEB NLP group at ImageCLEFmed caption 2022, in: CLEF2022 Working Notes,
CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy, 2022.</p>
    </sec>
    <sec id="sec-8">
      <title>A. Full results</title>
      <p>BLEU
METEOR
CIDEr
CLIPScore</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schaer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bromuri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2016 medical task</article-title>
          ,
          <source>in: Working Notes of CLEF 2016 (Cross Language Evaluation Forum)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>219</fpage>
          -
          <lpage>232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Eickhof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Schwall</given-names>
            ,
            <surname>A. G. S. de Herrera</surname>
          </string-name>
          , H. Müller, Overview of ImageCLEFcaption 2017 -
          <article-title>Image Caption Prediction and Concept Detection for Biomedical Images</article-title>
          , in: Working Notes of CLEF 2017 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Dublin, Ireland,
          <source>September 11-14</source>
          ,
          <year>2017</year>
          .,
          <year>2017</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-1866/invited_paper_7.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eickhof</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Andrearczyk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2018 Caption Prediction Tasks</article-title>
          , in: Working Notes of CLEF 2018 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Avignon, France,
          <source>September 10-14</source>
          ,
          <year>2018</year>
          .,
          <year>2018</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2125</volume>
          /invited_paper_4.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
          </string-name>
          , H. Müller,
          <article-title>Overview of the ImageCLEFmed 2019 Concept Detection Task</article-title>
          , in: L.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>D. E.</given-names>
          </string-name>
          <string-name>
            <surname>Losada</surname>
          </string-name>
          , H. Müller (Eds.),
          <source>Working Notes of CLEF 2019 - Conference and Labs of the Evaluation Forum, Lugano, Switzerland, September</source>
          <volume>9</volume>
          -
          <issue>12</issue>
          ,
          <year>2019</year>
          , volume
          <volume>2380</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEURWS.org,
          <year>2019</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2380</volume>
          /paper_245.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , H. Müller,
          <article-title>Overview of the ImageCLEFmed 2020 concept prediction task: Medical image understanding</article-title>
          ,
          <source>in: CLEF2020 Working Notes</source>
          , volume
          <volume>1166</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Thessaloniki, Greece,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>O.</given-names>
            <surname>Bodenreider</surname>
          </string-name>
          ,
          <article-title>The Unified Medical Language System (UMLS): integrating biomedical terminology</article-title>
          ,
          <source>Nucleic Acids Research</source>
          <volume>32</volume>
          (
          <year>2004</year>
          )
          <fpage>267</fpage>
          -
          <lpage>270</lpage>
          . doi:
          <volume>10</volume>
          .1093/nar/gkh061.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jacutprakart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEFmed 2021 concept &amp; caption prediction task</article-title>
          ,
          <source>in: CLEF2021 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Bucharest, Romania,
          <year>2021</year>
          , pp.
          <fpage>1101</fpage>
          -
          <lpage>1112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , L. Bloch,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brüngel</surname>
          </string-name>
          , A. IdrissiYaghir, H. Schäfer,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          , Overview of ImageCLEFmedical 2022 -
          <article-title>Caption Prediction and Concept Detection</article-title>
          , in: CLEF2022 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Bologna, Italy,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Koitka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nensa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <article-title>Radiology Objects in COntext (ROCO): A Multimodal Image Dataset</article-title>
          , in: Intravascular Imaging and Computer Assisted Stenting - and
          <string-name>
            <surname>-</surname>
          </string-name>
          Large-Scale
          <source>Annotation of Biomedical Data and Expert Label Synthesis - 7th Joint International Workshop</source>
          , CVII-STENT 2018 and Third International Workshop, LABELS 2018,
          <article-title>Held in Conjunction with MICCAI 2018, Granada</article-title>
          , Spain,
          <year>September 16</year>
          ,
          <year>2018</year>
          , Proceedings,
          <year>2018</year>
          , pp.
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -01364-6\_
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Drăgulinescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Snider</surname>
          </string-name>
          , G. Adams,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yetisgen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bloch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brüngel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Idrissi-Yaghir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hicks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Riegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Thambawita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Storås</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Papachrysos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schöler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Andrei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radzhabov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Coman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stan</surname>
          </string-name>
          , G. Ioannidis,
          <string-name>
            <given-names>H.</given-names>
            <surname>Manguinhas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ştefan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dogariu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Deshayes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Popescu</surname>
          </string-name>
          , Overview of ImageCLEF 2023:
          <article-title>Multimedia retrieval in medical, socialmedia and recommender systems applications</article-title>
          , in: Experimental IR Meets Multilinguality, Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 14th International Conference of the CLEF Association (CLEF</source>
          <year>2023</year>
          ), Springer Lecture Notes in Computer Science LNCS, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kaliosis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Moschovis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Charalambakos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pavlopoulos</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Androutsopoulos</surname>
          </string-name>
          , AUEB NLP group at ImageCLEFmedical caption
          <year>2023</year>
          , in: CLEF2023 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Aono</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shinoda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Asakawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shimizu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Togawa</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <article-title>Komoda, Multi-stage medical image captioning using classification and CLIP</article-title>
          , in: CLEF2023 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>V.</given-names>
            <surname>Yeshwanth</surname>
          </string-name>
          , P. P, L. Kalinathan,
          <article-title>Concept detection and image caption generation in medical imaging</article-title>
          ,
          <source>in: CLEF2023 Working Notes, CEUR Workshop Proceedings</source>
          , CEURWS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Transferring pre-trained large language-image model for medical image captioning</article-title>
          ,
          <source>in: CLEF2023 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Layode</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <article-title>Concept detection and caption prediction in ImageCLEFmedical caption 2023 with convolutional neural networks, vision and text-to-text transfer transformers</article-title>
          ,
          <source>in: CLEF2023 Working Notes, CEUR Workshop Proceedings</source>
          , CEURWS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nicolson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dowling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Koopman</surname>
          </string-name>
          ,
          <article-title>A concise model for medical image captioning</article-title>
          ,
          <source>in: CLEF2023 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lotfollahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nobakhtian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hajihosseini</surname>
          </string-name>
          , S. Eetemadi, IUST_NLPLAB at ImageCLEFmedical caption tasks
          <year>2023</year>
          , in: CLEF2023 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>H.</given-names>
            <surname>Shinoda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aono</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Asakawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shimizu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Komoda</surname>
          </string-name>
          , T. Togawa, KDE lab at ImageCLEFmedical caption
          <year>2023</year>
          , in: CLEF2023 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>B.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Raza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zou</surname>
          </string-name>
          , T. Zhang,
          <article-title>Customizing general-purpose foundation models for medical report generation</article-title>
          , in: CLEF2023 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S. S. N.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Srinivasan</surname>
          </string-name>
          , SSN MLRG at caption
          <year>2023</year>
          :
          <article-title>Automatic concept detection and caption prediction using ConceptNet and vision transformer</article-title>
          , in: CLEF2023 Working
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>