<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UIT-2Q2T at ImageCLEFmedical 2024 Caption: Multimodal medical image captioning using Bootstrapping Language-Image Pre-training</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thien V. Phan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Trinh K. Nguyen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Quang A.D.D. Hoang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Quan T. Phan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thien B. Nguyen-Tat</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Information Technology</institution>
          ,
          <addr-line>Ho Chi Minh City</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vietnam National University</institution>
          ,
          <addr-line>Ho Chi Minh City</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Introduction: Medical image captioning is an important AI task in healthcare, automating the generation of text descriptions to support the management and interpretation of medical images. Our team, UIT-2Q2T, participated in the second task of the ImageCLEFmedical 2024 Caption challenge using the ROCOv2 dataset with the Bootstrapping Language-Image Pre-training (BLIP) approach. Methods: Our approach leveraged the BLIP architecture for multimodal medical image captioning. This architecture employs a Vision Transformer (ViT) as the image encoder and a Bidirectional Encoder Representations from Transformers (BERT) as the text model. Results: We ranked 5th according to BERTscore and placed 3rd with ROUGEScore, BLEURTScore, and RefCLIPScore. Additionally, we achieved 2nd place for BLEU-1, METEOR, and CIDEr scores. Notably, we obtained the top position with a CLIPScore of 0.827074, demonstrating the efectiveness of our approach in medical image captioning. Conclusion: Our participation in the ImageCLEFmedical 2024 Caption challenge demonstrated the efectiveness of the BLIP architecture for medical image captioning, achieving a high CLIPScore of 0.82707. This result demonstrates the model's potential to generate accurate and informative textual descriptions from medical images, thereby aiding diagnosis and assisting non-experts in understanding medical images.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;CLEF 2024</kwd>
        <kwd>Medical image processing</kwd>
        <kwd>Image captioning</kwd>
        <kwd>BERT</kwd>
        <kwd>Pre-trained models</kwd>
        <kwd>BLIP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Image captioning, a well-established field in artificial intelligence (AI), finds applications across diverse
domains. In healthcare, the increasing availability of medical imaging equipment and the eficiency of
diagnosis based on visual data have fueled the popularity of image-based patient diagnosis. Medical
image captioning models address this need by automatically analyzing and describing medical images.
These models generate textual descriptions that assist doctors in diagnosing diseases, understanding
physiological processes, and enabling non-experts to interpret medical imagery.</p>
      <p>
        This field integrates computer vision and natural language processing, demanding an understanding
of image components and their relationships [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Various models, such as the Show-Attend-Tell, GPT-3,
and BioLinkBERT-Large, have been utilized to generate comprehensive and descriptive captions for
medical images, including radiological scans and histopathological specimens [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Transformer-based
approaches, like the Global-Local Visual Extractor (GLVE) and Cross Encoder-Decoder Transformer
(CEDT), have shown promise in capturing both global and local features of images, enhancing the
accuracy of generated captions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These advancements in medical image captioning not only facilitate
clinical workflows and decision-making but also contribute significantly to medical education by
providing quantitative indicators and assessments for learning outcomes [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        To successfully deploy image captioning in healthcare, it is essential to integrate efective algorithms
and use a suficiently large and diverse training dataset. Our team participated in ImageCLEF 2024
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for the ImageCLEFmedical 2024 Caption [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] task which consists of 2 subtask: Concept Detection,
Caption Prediction. We mainly focus on the latter. Here, participants are required to automatically
generate captions for given medical images, which could be of various modalities, such as ultrasound,
X-Ray, Computer Tomography (CT), Magnetic Resonance Imaging (MRI), etc.
      </p>
      <p>
        Our approach for the caption prediction subtask is based on Bootstrapping Language-Image
Pretraining (BLIP) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] architecture with a Vision Transformer (ViT) image Encoder [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In this paper,
Section 2 outlines the task and dataset descriptions. Section 3 describes our proposed methodology.
Section 4 details the implementation and results of the experiments. Finally, in Section 5, we conclude
by summarizing the results, discussing the weaknesses, and outlining potential improvements for the
future.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Task and Dataset Descriptions</title>
      <p>At ImageCLEFmedical 2024, we participated in the image captioning task. This is the 8th edition of the
ImageCLEFmedical caption task. In this section, we will introduce the task in the ImageCLEFmedical
2024 Caption and the dataset used for this challenge.</p>
      <sec id="sec-2-1">
        <title>2.1. Task Description</title>
        <p>
          ImageCLEFmedical 2024 Caption [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is one of ImageCLEFmedical’s tasks to create descriptive captions
for visual content. The tasks in ImageCLEFmedical Caption include two sub-tasks:
1. Concept detection: Based on the visual image content, this subtask provides the foundation for
the scene understanding step by identifying the individual elements from which the annotation
is generated.
2. Captions prediction: The core task is to create descriptive captions for given images. Leveraging
identified concepts and contextual understanding, the models are tasked with generating concise
and informative textual descriptions that accurately reflect the visual content depicted in the
image.
        </p>
        <p>
          In this study, we focus on the second sub-task based on the provided dataset ROCOv2 [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Dataset Descriptions</title>
        <p>
          The dataset for this task is ROCOv2 [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] - an extended version of ROCO [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. It is a multimodal dataset
consisting of radiological images and associated medical concepts and captions extracted from the
PubMed Open Access subset. All images in the dataset were accompanied by a caption, which form
the labels for the caption prediction task. Each caption was pre-processed by removing links from the
captions. The splits for the dataset are as follows:
• Training Set: Consists of 70,108 radiology images.
• Validation Set: Consists of 9,972 radiology images.
        </p>
        <p>• Test Set: Consists of 17,237 radiology images.</p>
        <p>As shown in Figure 1, the majority of captions in the dataset range from 50 to 150 words in length.
Similarly, Figure 2 illustrates that among the six imaging modalities represented in the dataset, CT scans
and X-rays are predominant, accounting for 24,227 and 19,363 samples in the training set, respectively.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methods</title>
      <p>In this study, the BLIP model was employed to tackle the image captioning task. This approach involves
ifnetuning the BLIP model on the competition dataset, which consists of diverse and challenging
imagecaption pairs. The pipeline of our method is illustrated in Figure 3, showcasing the steps involved in
adapting the BLIP model for our specific image captioning task.</p>
      <sec id="sec-3-1">
        <title>3.1. Models</title>
        <p>
          Bootstrapping Language-Image Pre-training (BLIP) [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] is a Vision-Language Pre-training (VLP)
framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP
efectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates
synthetic captions and a filter removes the noisy ones.
        </p>
        <p>
          The model uses Vision Transformer (ViT) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] which divides the input image into patches and encodes
them as a sequence of embedding with the addition of [CLS] token to represent the globe image feature.
As the authors mentioned ViT uses less computation cost and is a straightforward method, and is being
adopted by recent methods.
        </p>
        <p>N X
Image
Encoder</p>
        <p>Feed Forward
Self Attention</p>
        <p>CC BY-NC
[Bajracharya et al. (2021)]
TEenxctoder</p>
        <p>N X</p>
        <p>Cross Attention
Feed Forward</p>
        <p>Bi Self-Att
"[CLS] +
Feed Forward
Cross Attention</p>
        <p>Bi Self-Att
Feed Forward
Cross Attention</p>
        <p>Causal Self-Att
Image-grounded
Text encoder "[Encode] +</p>
        <p>Image-grounded</p>
        <p>Text decoder "[Decode] +
Axial image of contrast-enhanced computed tomography (CECT)
of lower abdomen shows a mixed solid and cystic mass....</p>
        <p>To be able to train or pre-train the model for understanding and generation tasks, a multimodal
mixture of an encoder and decoder is used, integrating three functionalities and three objectives, as
illustrated in Figure 3. The functionalities include:
• Unimodal Encoder: Encodes either the image or the text separately without considering the
other modality. This helps in understanding individual representations.
• Image-grounded Text Encoder: Encodes text while being conditioned on the image, allowing
the model to capture relationships between visual and textual information.
• Image-grounded Text Decoder: Generates text based on the given image, useful for tasks like
image captioning where the output is text describing the input image.</p>
        <p>The objectives are:
• Image-Text Contrastive Loss (ITC): Ensures that paired image and text representations are
closer together in the embedding space compared to unpaired ones. This helps the model learn
strong associations between images and their corresponding texts.
• Image-Text Matching Loss (ITM): Assesses whether a given image and text pair match or not,
promoting accurate image-text alignment in the embedding space.
• Language Modeling Loss (LM): Focuses on generating coherent and contextually accurate text
based on given inputs, improving the model’s language generation capabilities.</p>
        <p>These functionalities and objectives together enable the BLIP model to perform both vision-language
understanding and generation tasks efectively.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Evaluation Metrics</title>
        <p>
          Following the guidelines provided by the competition organizers, we employed two main metrics:
BERTScore [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and ROUGEScore [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. To calculate BERTScore, we use the
"microsoft/deberta-xlargemnli" model, which can be found on the Hugging Face Model Hub1. Additionally, other metrics
1https://huggingface.co/microsoft/deberta-xlarge-mnli (Last accessed: May 17, 2024)
such as BLEU-1 [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], BLEURT [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], METEOR [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], CIDEr [17], CLIPScore [18], RefCLIPScore [19],
ClinicalBLEURTScore [20], and MedBERTScore [20] were also applied for evaluation.
        </p>
        <p>As the organizers’ instructions, the captions underwent preprocessing through three steps:
conversion to lowercase, replacement of numbers with a special token, and removal of punctuation. This
preprocessing aimed to standardize the text inputs and enhance the quality of evaluation result.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>In this section, we present our experimental setup and results for evaluating the BLIP model in the
ImageCLEFmedical 2024 Caption challenge. The experiments were designed to test the model’s
performance across diferent configurations and metrics, aiming to generate accurate and informative
captions for medical images. We describe the setup in detail and discuss the results obtained from
various test scenarios.</p>
      <sec id="sec-4-1">
        <title>4.1. Experimental Setup</title>
        <p>In our experiments, we employed BLIP model (base/large) from pre-trained checkpoints. For BLIP
base, weights were utilized from the checkpoint "Salesforce/blip-image-captioning-base"2. Training was
conducted over 15 epochs with an initial learning rate of 1e-5. A StepLR scheduler was used to decrease
the learning rate by a factor of 10 every 3 epochs. For the BLIP large model, weights were utilized
from the checkpoint "Salesforce/blip-image-captioning-large"3. Training was conducted over 5 epochs
with an initial learning rate of 1e-5 and was stopped when the loss ceased to decrease. Throughout
all experiments, the AdamW optimizer [21] was used. Input images were resized to 224x224, and the
maximum length of text input was set to 200 tokens. To facilitate model training, a single GPU A100
PCIE 40GB was used.</p>
        <p>For each model, experiments were conducted with four diferent generation settings using
no_repeat_ngram_size = 3 to prevent the model from repeating any n-gram of size 3 within the
generated text. The generation settings are as follows:
(1) Greedy Search: This setting selects the token with the highest probability at each step, ensuring
a straightforward and fast generation process but potentially missing out on more diverse or
optimal sequences.
(2) Beam Search with beam_size = 3: Beam search keeps the top 3 most probable sequences at each
generation step, allowing for more exploration of potential sequences compared to greedy search.
(3) Beam Search with beam_size = 4: Similar to the previous setting, but with a beam size of 4, which
balances between exploration and computational eficiency.
(4) Beam Search with beam_size = 5: This setting further increases the beam size to 5, allowing for
more comprehensive exploration while balancing computational eficiency.
(5) Beam Search with beam_size = 10: With a beam size of 10, this setting aims for a broader
exploration of possible sequences, potentially improving the quality of text generation at the cost
of higher computational resources.</p>
        <p>The source code for our experiments is available on GitHub4.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Experimental Results</title>
        <p>To evaluate the efectiveness of our approach using the BLIP model on the image captioning task,
we conducted a series of experiments on the competition dataset. This section presents the results,
highlighting the model’s performance in generating captions. We compared the experimental results of
the model using diferent configurations and also benchmark our model’s performance against other
teams’.
2https://huggingface.co/Salesforce/blip-image-captioning-base (Last accessed: May 17, 2024)
3https://huggingface.co/Salesforce/blip-image-captioning-large (Last accessed: May 17, 2024)
4https://github.com/QuangHoang059/DS312</p>
        <p>CC BY-NC
[Trowbridge et all. (2022)]</p>
        <sec id="sec-4-2-1">
          <title>4.2.1. Results on Validation Set</title>
          <p>As shown in Table 1, the BLIP large model outperforms the BLIP base model. Both models show
generative capabilities, with beam search outperforming greedy search across ROUGEScore, as well
as BERTScore and BLEUScore. Specifically, the BLIP base model achieves its highest BERTScore and
ROUGE score with a beam size of 5, and its best BLEU score with a beam size of 3. The BLIP large
model attains optimal results across all three metrics with a beam size of 4. Additionally, as illustrated
by the two examples in Figure 4, the model accurately identifies objects and colors (white arrow), as
well as diferent imaging modalities (CT and sagittal T2-weighted MRI).</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>4.2.2. Results on Test Set</title>
          <p>According to the private test results announced by the organizing committee and partially presented
in Table 2, our team ranked 5th based on the BERTScore metric. We achieved 3rd place with ROUGE,
BLEURT, and RefCLIPScore metrics. For BLEU-1, METEOR, and CIDEr scores, we achieved 2nd place.
Notably, we attained 1st place with a CLIPScore of 0.827074. These results demonstrate the model’s
expected performance.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future work</title>
      <p>In this paper, we implemented and experimented with the BLIP model for the task of medical image
captioning in the ImageCLEFmedical 2024 Caption challenge. The experimental results across various
configurations showed promising outcomes, with the model achieving a CLIPScore of 0.82707 on the
test set of the ROCOv2 dataset. Despite these achievements, there is still room for improvement in our
research. The primary weakness of the model is its pre-training on a dataset significantly diferent from
the medical domain, resulting in considerable bias.</p>
      <p>Moving forward, we aim to enhance the model’s accuracy by utilizing pre-trained models with
datasets that are more closely aligned with medical and diagnostic domains. Additionally, we plan
to apply preprocessing methods tailored to diferent types of medical images to further improve the
performance. Exploring domain-specific augmentation techniques and integrating more diverse medical
datasets could also provide substantial gains. By addressing these areas, we hope to develop a more
robust and accurate medical image captioning model, which can be a valuable tool in clinical settings
for aiding diagnosis and assisting non-experts in understanding medical imagery.</p>
      <p>Furthermore, future research will involve a detailed analysis of the model’s errors to understand the
underlying reasons for its mispredictions. This analysis will guide the development of more efective
strategies for fine-tuning and enhancing the model’s capabilities. Our ultimate goal is to contribute to
the advancement of AI in healthcare by providing reliable and interpretable models that can support
medical professionals and improve patient outcomes.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgment References</title>
      <p>This research is funded by University of Information Technology-Vietnam National University
HoChiMinh City under grant number D4-2024-01.</p>
      <p>URL: https://aclanthology.org/W05-0909.
[17] R. Vedantam, C. L. Zitnick, D. Parikh, Cider: Consensus-based image description evaluation, in:
2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 4566–4575.
doi:10.1109/CVPR.2015.7299087.
[18] J. Hessel, A. Holtzman, M. Forbes, R. Bras, C. Yejin, Clipscore: A reference-free evaluation metric
for image captioning, 2021, pp. 7514–7528. doi:10.18653/v1/2021.emnlp-main.595.
[19] L. Jin, G. Luo, Y. Zhou, X. Sun, G. Jiang, A. Shu, R. Ji, Refclip: A universal teacher for weakly
supervised referring expression comprehension, 2023, pp. 01–10. doi:10.1109/CVPR52729.
2023.00263.
[20] A. Ben Abacha, W.-w. Yim, G. Michalopoulos, T. Lin, An investigation of evaluation methods in
automatic medical note generation, in: A. Rogers, J. Boyd-Graber, N. Okazaki (Eds.), Findings of the
Association for Computational Linguistics: ACL 2023, Association for Computational Linguistics,
Toronto, Canada, 2023, pp. 2575–2588. URL: https://aclanthology.org/2023.findings-acl.161. doi: 10.
18653/v1/2023.findings-acl.161.
[21] I. Loshchilov, F. Hutter, Decoupled weight decay regularization, in: 7th International Conference
on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net,
2019. URL: https://openreview.net/forum?id=Bkg6RiCqY7.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <article-title>Skin medical image captioning using multi-label classification and siamese network</article-title>
          ,
          <source>IEEE Access 11</source>
          (
          <year>2023</year>
          )
          <fpage>23447</fpage>
          -
          <lpage>23454</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2023</year>
          .
          <volume>3249462</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Elbedwehy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Medhat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hamza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alrahmawy</surname>
          </string-name>
          ,
          <article-title>Enhanced descriptive captioning model for histopathological patches</article-title>
          ,
          <source>Multimedia Tools and Applications</source>
          <volume>83</volume>
          (
          <year>2023</year>
          )
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s11042-023-15884-y.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Selivanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Rogov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chesakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Fedulova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dylov</surname>
          </string-name>
          ,
          <article-title>Medical image captioning via generative pretrained transformers</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>13</volume>
          (
          <year>2023</year>
          ).
          <source>doi: 10.1038/ s41598-023-31223-5.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Cross encoder-decoder transformer with global-local visual extractor for medical image captioning</article-title>
          ,
          <source>Sensors</source>
          <volume>22</volume>
          (
          <year>2022</year>
          )
          <article-title>1429</article-title>
          . doi:
          <volume>10</volume>
          .3390/s22041429.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.-R.</given-names>
            <surname>Beddiar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Oussalah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Seppänen</surname>
          </string-name>
          ,
          <article-title>Automatic captioning for medical imaging (mic): a rapid review of literature, Artif</article-title>
          .
          <source>Intell. Rev</source>
          .
          <volume>56</volume>
          (
          <year>2022</year>
          )
          <fpage>4019</fpage>
          -
          <lpage>4076</lpage>
          . URL: https://doi.org/10.1007/ s10462-022-10270-w. doi:
          <volume>10</volume>
          .1007/s10462-022-10270-w.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Drăgulinescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , L. Bloch,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brüngel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Idrissi-Yaghir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Pakull</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Damm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bracke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Andrei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Prokopchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Karpenka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radzhabov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macaire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schwab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lecouteux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Esperança-Rodier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yetisgen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hicks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Riegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Thambawita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Storås</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Heinrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kiesel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , Overview of ImageCLEF 2024:
          <article-title>Multimedia retrieval in medical applications, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 15th International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ), Springer Lecture Notes in Computer Science LNCS, Grenoble, France,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Seco de Herrera</surname>
          </string-name>
          , L. Bloch,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brüngel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Idrissi-Yaghir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bracke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Damm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Pakull</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          , Overview of ImageCLEFmedical 2024 -
          <article-title>Caption Prediction and Concept Detection</article-title>
          , in: CLEF2024 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Grenoble, France,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xiong</surname>
          </string-name>
          , S. Hoi, BLIP:
          <article-title>Bootstrapping language-image pre-training for unified visionlanguage understanding and generation</article-title>
          , in: K. Chaudhuri,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jegelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Szepesvari</surname>
          </string-name>
          , G. Niu, S. Sabato (Eds.),
          <source>Proceedings of the 39th International Conference on Machine Learning</source>
          , volume
          <volume>162</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>12888</fpage>
          -
          <lpage>12900</lpage>
          . URL: https://proceedings.mlr.press/v162/li22n.html.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weissenborn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Unterthiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Minderer</surname>
          </string-name>
          , G. Heigold,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Houlsby</surname>
          </string-name>
          ,
          <article-title>An image is worth 16x16 words: Transformers for image recognition at scale</article-title>
          , in: International Conference on Learning Representations,
          <year>2021</year>
          . URL: https://openreview.net/forum?id=YicbFdNTTy.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bloch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brüngel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Idrissi-Yaghir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Koitka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            , H. Müller,
            <given-names>P. A.</given-names>
          </string-name>
          <string-name>
            <surname>Horn</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Nensa</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>M. Friedrich, ROCOv2: Radiology Objects in COntext version 2, an updated multimodal image dataset, Scientific Data (</article-title>
          <year>2024</year>
          ). URL: https://arxiv.org/abs/2405.10004v1.
          <source>doi:10.1038/s41597-024-03496-6.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Koitka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nensa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <article-title>Radiology Objects in COntext (ROCO): A Multimodal Image Dataset:</article-title>
          7th Joint International Workshop, CVII-STENT 2018 and Third International Workshop, LABELS 2018,
          <article-title>Held in Conjunction with MICCAI 2018, Granada</article-title>
          , Spain,
          <year>September 16</year>
          ,
          <year>2018</year>
          , Proceedings,
          <year>2018</year>
          , pp.
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -01364-6_
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang*</surname>
          </string-name>
          , V. Kishore*, F. Wu*,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Artzi</surname>
          </string-name>
          , Bertscore:
          <article-title>Evaluating text generation with bert</article-title>
          ,
          <source>in: International Conference on Learning Representations</source>
          ,
          <year>2020</year>
          . URL: https://openreview. net/forum?id=SkeHuCVFDr.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>C.-Y. Lin</surname>
            ,
            <given-names>ROUGE:</given-names>
          </string-name>
          <article-title>A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics</article-title>
          , Barcelona, Spain,
          <year>2004</year>
          , pp.
          <fpage>74</fpage>
          -
          <lpage>81</lpage>
          . URL: https://www.aclweb.org/anthology/W04-1013.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>K.</given-names>
            <surname>Papineni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roukos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ward</surname>
          </string-name>
          , W.-J. Zhu,
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          , in: P.
          <string-name>
            <surname>Isabelle</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Charniak</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Lin</surname>
          </string-name>
          (Eds.),
          <article-title>Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Philadelphia, Pennsylvania, USA,
          <year>2002</year>
          , pp.
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          . URL: https://aclanthology.org/P02-1040. doi:
          <volume>10</volume>
          .3115/1073083.1073135.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Sellam</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Parikh</surname>
          </string-name>
          ,
          <article-title>BLEURT: Learning robust metrics for text generation</article-title>
          , in: D.
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Chai</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Schluter</surname>
          </string-name>
          , J. Tetreault (Eds.),
          <article-title>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics</article-title>
          , Online,
          <year>2020</year>
          , pp.
          <fpage>7881</fpage>
          -
          <lpage>7892</lpage>
          . URL: https://aclanthology.org/
          <year>2020</year>
          .acl-main.
          <volume>704</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .acl-main.
          <volume>704</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lavie</surname>
          </string-name>
          ,
          <string-name>
            <surname>METEOR:</surname>
          </string-name>
          <article-title>An automatic metric for MT evaluation with improved correlation with human judgments</article-title>
          , in: J.
          <string-name>
            <surname>Goldstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lavie</surname>
            ,
            <given-names>C.-Y.</given-names>
          </string-name>
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          Voss (Eds.),
          <source>Proceedings of the ACL Workshop</source>
          on Intrinsic and
          <article-title>Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Association for Computational Linguistics</article-title>
          , Ann Arbor, Michigan,
          <year>2005</year>
          , pp.
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>