<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the MEDIQA-MAGIC Task at ImageCLEF 2024: Multimodal And Generative TelemedICine in Dermatology⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Wen-wai Yim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asma Ben Abacha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yujuan Fu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhaoyi Sun</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Meliha Yetisgen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fei Xia</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Microsoft Health AI</institution>
          ,
          <addr-line>Redmond, 98052</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Washington</institution>
          ,
          <addr-line>Seattle, 98109</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Multimodal processing and language generation require models to internally represent both language and vision, and then generate contextually appropriate responses. To do so with arbitrary images and textual inputs in the medical field, requires additional high performance and fidelity. This paper presents the overview of the MEDIQA-MAGIC shared task at ImageCLEF 2024. In this dermatological visual question-answering (VQA) task, participants receive the input of an image and a textual consumer health query, and are expected to output a textual medical answer. A total of twenty two runs were submitted with a variety of general language-vision models and fine-tuned models, with the best team achieving 8.969 BLEU points. We hope that the findings and insights explored here will inspire future research directions to support improved patient care.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Visual Question Answering</kwd>
        <kwd>Response Generation</kwd>
        <kwd>Dermatology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Partially in efect after the adoption of meaningful use requirements in the United States, the ability to
message doctors and receive care remotely on patient portals have skyrocketed since the 2019 COVID
pandemic [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Furthermore, the establishment of online tele-health companies outside of traditional
hospital settings, e.g. teledoc [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], icliniq [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and amazon clinic [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], harkens consumer health needs for
on-demand medical care access. Asynchronous online dermatology consultation is one application area
for this new care delivery method. In this scenario, patients may provide their dermatology images
and questions through electronic messaging; medical doctors may then likewise provide electronic
responses related to treatment and medication. However, whether as an extended branch of a hospital
institution or as a standalone online care alternative, these services require a medical doctor in the loop
to deliver safe, reliable care – putting additional demands on provider workload.
      </p>
      <p>
        Automated models have the opportunity to provide response suggestions, which in turn may help
optimize care quality and eficiency, deburdening healthcare workers. While large multi-modal language
models have made significant strides, such as with the results of Gemini [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and GPT-4o [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], there
remain questions about the applicability of such models to real-world unconstrained tasks. The area
of multi-modal consumer health question-answering is a challenging task. It requires processing of
uncontrolled user-generated images, featuring variable angles, lighting, and resolution, as well as
arbitrary textual content and queries. As a medical task, reasonable accuracy and reliability are critical.
      </p>
      <p>
        To benchmark the state-of-the-art large multi-modal language models’ performance for this problem,
we have conducted the MEDIQA-MAGIC shared task as part of ImageCLEF 2024 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Specifically,
participants are required to compose automatic responses to patient dermatological queries with a
textual question and image context. Previous related shared tasks have featured radiology-related
visual question answering (VQA) [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ], as well as text-only consumer health answer generation
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Other medical VQA problems include images from the areas of pathology and GI-tract [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ].
Meanwhile, single modality dermatological image classification tasks include automatic categorizations
to predefined diseases or symptom characteristics [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. This year’s edition tackles answer generation
for multi-modal consumer health question task. A similar task was part of the related NAACL 2024
ClinicalNLP challenge MEDIQA-M3G 2024 [16] featuring a diferent dataset.
      </p>
      <p>In the following sections, we introduce the tasks, describe the evaluation, present the participating
teams’ results, as well as provide some insight into future directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task Description and Dataset</title>
      <p>In this shared-task, an input instance is composed of a single textual query and an image, representing a
consumer health question. The expected output is a free-text response, representing a possible doctor’s
answer to the query. Figure 1 shows an example instance.</p>
      <p>The dataset was sourced from real consumer health queries found on Reddit for posts related to
dermatological problems (subreddit r/DermatologyQuestions) 1. Encounters were filtered out if they
met at least one of the following exclusion criteria: (a) images that included identifying features (e.g.
full faces), (b) queries that were not seeking information (e.g. “look at my tatoo”), (c) images including
genitalia, and (d) images that contained annotations (e.g. drawn arrows). Gold standard responses were
generated by 3 certified practicing dermatologists. The train and validation sets were single annotated;
the test set was double-annotated. To comply with Reddit data usage guidelines, only post IDs and our
response labels were shared with participants. Participants who registered through Reddit could receive
API credentials to access Reddit’s data. Afterwards, the participants could use the supplied download
script2 to retrieve the original input data.</p>
      <p>Table 1 provides summary statistics of the dataset. A single query may involve multiple anatomic
locations. Because users may delete content, the final set of test set IDs was determined by the subset of
test IDs retrieval shortly after the submission deadline. As a consequence, the final number of test set
encounters may include fewer instances than the original labeled test set. The data here used a subset
of the DermaVQA dataset, for which the full corpus creation description can be found in [17].</p>
    </sec>
    <sec id="sec-3">
      <title>3. Evaluation Methodology</title>
      <p>We evaluated the system responses by comparing them with the double-annotated gold standard
responses per query. We used relevant multi-reference metrics/variants including:</p>
      <sec id="sec-3-1">
        <title>1https://www.reddit.com/r/DermatologyQuestions/ 2https://github.com/wyim/MEDIQA-MAGIC-2024</title>
        <p>deltaBLEU. deltaBLEU is a variant of SacreBLEU developed for response generation, in which many
diverse gold standard responses are possible [18]. The metric incorporates human-annotated quality
rating and assigns higher weights to n-grams from responses rated to be of higher quality. The authors
have shown this method produces higher correlation with human rankings compared to previous BLEU
metrics. In this task, we weigh both annotator responses as equal, defaulting to a normal BLEU score
behavior. This metric was used for the shared task ranking.</p>
        <p>BERTScore. BERTScore3 [19] averages the maximum word embedding similarity scores between two
texts based on BERT embeddings. This metric has been shown to work well on a variety of tasks,
including image captioning and machine translation. The maximum was taken over multiple over
pairwise scores when multiple references were available.</p>
        <p>MEDCON. In this task, we propose MEDCON a medical information-extraction-based metric. The
metric uses QuickUMLS4 to identify medical concepts in conjunction with an in-house Llama-based
assertion classifier [ 20]. Concepts identified by QuickUMLS are normalized according to a curated
concept map. Precision, recall, and F1 were calculated based on combined concept and assertion statuses.
The maximum was taken over multiple pairwise scores when multiple references were available. A
variant of this metric, excluding the assertion status, was used in our previous work for measuring
clinical note summarization from medical dialogue[21]. The evaluation code can be found in our GitHub
repo5.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>
        Out of 30 initial registrations, 3 teams submitted results in a total of 22 runs. The final results are
shown in Table 3. The submitted systems represented a variety of solutions, including leveraging
out-of-the-box Gemini [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] models (YuanAI), applying small visual language models (VisionQAries), and
utilizing visual-language encoders with cosine similarities (IRLab@IIT_BHU). The ranges of scores
performed at the lower spectrum for all three metrics (100 total for BLEU, and 1.0 for BERTScore and
      </p>
      <sec id="sec-4-1">
        <title>3github.com/Tiiiger/bert_score 4github.com/Georgetown-IR-Lab/QuickUMLS 5https://github.com/wyim/MEDIQA-M3G-2024</title>
        <p>YuanAI
IRLab@IIT_BHU
IRLab@IIT_BHU
IRLab@IIT_BHU
YuanAI
YuanAI
VisionQAries
VisionQAries
IRLab@IIT_BHU
IRLab@IIT_BHU
IRLab@IIT_BHU
VisionQAries
IRLab@IIT_BHU
VisionQAries
VisionQAries
VisionQAries
VisionQAries
IRLab@IIT_BHU
IRLab@IIT_BHU</p>
        <p>MODELS_EXACT
moondream2
Clip, Cosine similarity
CLIP, Cosine similarity,
data augmentation
llama3,gemini-pro
Clip, triplet loss, textgenie
CLIP, Cosine similarity,
textgenie
CLIP, BIsltm
llama
llama3,gemini-pro
moondream2
moondream2
CLIP, GPT2-xl
BERT, Clip, BIlstm
CLIP, GPT2-xl
moondream2
CLIP, GPT2-xl
tiny-llava-v1-hf
moondream2
tiny-llava-v1-hf
moondream2
CLIP, GPT-2
Clip, Cosine similarity
MEDCON), indicating the dificulty of the task. Each teams’ system descriptions are described in the
following paragraphs:
IRLab@IIT_BHU This team’s approach involved multiple steps, using both pre-trained language
and vision-language models in conjunction with their own neural network architecture. The team
ifrst manually labeled instances into 160 hand-crafted, non-mutually-exclusive categories to be used
as targets. For each query instance, image and text were passed through a CLIP [25] vision and text
encoder respectively. Text-encoded data was then sent through a Bi-LSTM, and the vision-encoded
data, passed through an MLP layer. The results of both were averaged to produce a label vector.
During training, the vector is compared with positive and negative label embeddings using a weighted
cosine similarity loss. During inference, the combined embedding was compared with the closest label
embedding and assigned the corresponding class. The team also experimented with data augmentation
by adding paraphrased versions of the original data using TextGenie [26], and using GPT2 [27] as the
ifnal classifier.</p>
        <p>VisionQAries This team’s solutions focused on applying small multi-modal models. Mainly, they tested
two diferent approaches: (a) direct prompting on pre-trained moondream2 [ 28] and TinyLLaVA [29]
models; and (b) fine-tuning moondream2 model. During testing, they compared two diferent prompts.
Based on their results, they found that fine-tuning models gave better results than direct prompting in
the context of BLEU scores.</p>
        <p>YuanAI This team used a two-step approach. In the first step, they utilized the Gemini as their
image2-text model to create a descriptive text. In a second step, using the textual query and input from the
previous step, they employed a LoRA fine-tuned Llama3 [ 30] to use the output from the previous step
and the query as input to an LLM model to generate the final response.</p>
        <p>As evidenced in the spread of scores of each team across the rankings, the same setups with small
variations may produce widely diferent performances. For example, Team VisionQAries’ moondream2
runs had rankings at 1, 10, and 20 at BLEU scores of 8.969, 3.310, and 1.250 – utilizing diferent
finetuning and prompting variations. Although the best system scored by BLEU was much higher compared
to the next best system, the other metrics did not reflect this diference.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion and Conclusions</title>
      <p>Multi-modal question answering for unconstrained answers is a challenging problem. From the similarity
of scores in a wide variation of systems, it is clear that no one architecture is particularly superior at
this task.</p>
      <p>This year’s related 2024 NAACL ClinicalNLP MEDIQA-M3G [16] task posed the same problem of
dermatological multi-modal answer generation as this shared task, however the data characteristics
of these two challenges were completely diferent. The NAACL ClinicalNLP MEDIQA-M3G shared
task utilized the iiyi subset of the DermaVQA dataset [17]. Particularly, the data was sourced form a
Chinese medical platform where most query posts contained 1-3+ images and the naturally-occurring
responses tended to have shorter replies. In contrast, in this challenge, the MEDIQA-MAGIC shared
task, data originated from Reddit posts, with one image and brief shorter queries; where responses were
generated by medical doctors hired for the dataset creation - often providing very complete well-formed
answers. As a consequence, in this task, the data contained shorter queries and longer answers with an
average of 30 and 95 words, respectively, compared to 80 and 12 words in the MEDIQA-M3G shared
task English version data subset. Although the magnitude of BLEU and BERTscores were similar in
both challenges, the MEDCON scores were much lower (highest 0.1 F1 in this task, compared to a high
of 0.29 in the M3G task). This can be attributed to the long answers expected in this dataset which may
include many more concepts. That said, the overall modest scores across both shared tasks highlight
the need for improved answer generation methods for dermatological VQA. More details on the dataset
can be found in our dataset paper[17] and released dataset(https://osf.io/72rp3/).</p>
      <p>Despite diferences, similar task-related issues arose in both of these shared tasks. Firstly, true
dermatology gold standards for benign maladies are rare as typically suspected malignant lesions are
prioritized for pathologically testing. This may lead to diferences in dermatological expert opinions
which cannot be resolved. Even with access to large private health records, methods to tackle the
determination of the best gold label will require additional investigation. Secondly, in contrast to
previous VQA tasks, this task expected long-form natural language outputs. Although prior VQA
datasets had textual outputs, in reality, the number of question types are limited, with answers on
average at 1-2 words long. In fact, all previous VQA tasks report accuracy as a metric. Natural language
generation evaluation with respect to VQA is an area needing much more future research, particularly
with respect to fairly evaluating instances with multiple diverse possible answers and evaluating
instances of long free-text responses.</p>
      <p>This shared task revealed a multitude of opportunities for future modeling, corpus creation, and
evaluation. Future directions of study includes testing additional vision-language models, incorporating
intermediate image segmentation or image extraction steps, and re-ranking answers. In the future,
these can be tested on the larger combined DermaVQA dataset[17]. In future editions of this task,
we will experiment with other evaluation methods, e.g. ranking or weighting based on normalized
medical concepts. We hope that the benchmarks, insights, and datasets presented here will inspire
future research directions.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We would like to thank Thomas Lin from Microsoft Health AI and the ImageCLEF organizers for their
feedback and support for the MEDIQA-MAGIC 2024 shared tasks. We also thank our annotation team
for preparing the data in time for the challenge and all the participating teams who contributed to
the success of these shared tasks through their interesting approaches and experiments and strong
engagement.
[16] A. Ben Abacha, W. Yim, Y. Fu, Z. Sun, F. Xia, M. Yetisgen, M. Krallinger, Overview of the
mediqam3g 2024 shared tasks on multilingual multimodal medical answer generation, in:
NAACLClinicalNLP 2024, 2024.
[17] W. Yim, Y. Fu, Z. Sun, A. Ben Abacha, M. Yetisgen, F. Xia, Dermavqa: A multilingual visual
question answering dataset for dermatology, in: Medical Image Computing and Computer Assisted
Intervention – MICCAI 2024, 2024.
[18] M. Galley, C. Brockett, A. Sordoni, Y. Ji, M. Auli, C. Quirk, M. Mitchell, J. Gao, B. Dolan, deltaBLEU:
A discriminative metric for generation tasks with intrinsically diverse targets, in: Proceedings of the
53rd Annual Meeting of the Association for Computational Linguistics and the 7th International
Joint Conference on Natural Language Processing (Volume 2: Short Papers), Association for
Computational Linguistics, Beijing, China, 2015, pp. 445–450.
[19] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, Y. Artzi, Bertscore: Evaluating text generation with
bert, ArXiv abs/1904.09675 (2019).
[20] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P.
Bhargava, S. Bhosale, et al., Llama 2: Open foundation and fine-tuned chat models, arXiv preprint
arXiv:2307.09288 (2023).
[21] W. Yim, Y. Fu, A. B. Abacha, N. Snider, T. Lin, M. Yetisgen, Aci-bench: a novel ambient clinical
intelligence dataset for benchmarking automatic visit note generation, Scientific Data 10 (2023).</p>
      <p>URL: https://api.semanticscholar.org/CorpusID:259075199.
[22] A. Agrawal, S. Pal, Irlab@iit_bhu at mediqa-magic 2024: Medical question answering using
classification model, in: Experimental IR Meets Multilinguality, Multimodality, and Interaction,
Proceedings of the 15th International Conference of the CLEF Association (CLEF 2024), Springer
Lecture Notes in Computer Science LNCS, Grenoble, France, 2024.
[23] P. Cieplicka, J. Kłos, M. Morawski, Visionqaries at mediqa-magic 2024: Small vision language
models for dermatological diagnosis, in: Experimental IR Meets Multilinguality, Multimodality,
and Interaction, Proceedings of the 15th International Conference of the CLEF Association (CLEF
2024), Springer Lecture Notes in Computer Science LNCS, Grenoble, France, 2024.
[24] H. Fu, H. Huang, Yuanai at mediqa-magic 2024: Improving medical vqa performance through
parameter-eficient fine-tuning, in: Experimental IR Meets Multilinguality, Multimodality, and
Interaction, Proceedings of the 15th International Conference of the CLEF Association (CLEF
2024), Springer Lecture Notes in Computer Science LNCS, Grenoble, France, 2024.
[25] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell,
P. Mishkin, J. Clark, G. Krueger, I. Sutskever, Learning transferable visual models from
natural language supervision, in: International Conference on Machine Learning, 2021. URL:
https://api.semanticscholar.org/CorpusID:231591445.
[26] Text genie, https://github.com/hetpandya/textgenie, 2024. Accessed: 2024-05-31.
[27] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language models are unsupervised
multitask learners, 2019. URL: https://api.semanticscholar.org/CorpusID:160025533.
[28] moondream.ai, https://moondream.ai/, 2024. Accessed: 2024-05-31.
[29] B. Zhou, Y. Hu, X. Weng, J. Jia, J. Luo, X. Liu, J. Wu, L. Huang, Tinyllava: A framework of small-scale
large multimodal models, 2024. arXiv:2402.14289.
[30] llama3 model, https://llama.meta.com/llama3/, 2024. Accessed: 2024-05-31.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Bishop</surname>
          </string-name>
          , M. J. Press,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Mendelsohn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Casalino</surname>
          </string-name>
          ,
          <article-title>Electronic communication improves access, but barriers to its widespread adoption remain, Health afairs</article-title>
          (
          <issue>Project Hope</issue>
          )
          <volume>32</volume>
          (
          <year>2024</year>
          )
          <volume>10</volume>
          .1377/hlthaf.
          <year>2012</year>
          .
          <volume>1151</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Sinsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. D.</given-names>
            <surname>Shanafelt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Ripp</surname>
          </string-name>
          ,
          <article-title>The electronic health record inbox: Recommendations for relief 37 (</article-title>
          <year>2024</year>
          )
          <fpage>4002</fpage>
          -
          <lpage>4003</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] teledoc, https://www.teladochealth.com/,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -05-31.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4] icliniq, https://www.icliniq.com/,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -05-31.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Amazon</surname>
            <given-names>clinic</given-names>
          </string-name>
          , https://clinic.amazon.com/,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -05-31.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Gemini</surname>
            <given-names>models</given-names>
          </string-name>
          , https://ai.google.dev/gemini-api/docs/models/gemini,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -04-24.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] Gpt-4o, https://openai.com/index/hello-gpt-4o/,
          <year>2024</year>
          . Accessed:
          <fpage>2024</fpage>
          -05-31.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Drăgulinescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Garcıa Seco de Herrera</surname>
          </string-name>
          , L. Bloch,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brüngel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Idrissi-Yaghir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schäfer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Pakull</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Damm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bracke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Andrei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Prokopchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Karpenka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Radzhabov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Macaire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schwab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lecouteux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Esperança-Rodier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yetisgen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hicks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Riegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Thambawita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Storås</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Heinrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kiesel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , Overview of ImageCLEF 2024:
          <article-title>Multimedia retrieval in medical applications, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 15th International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ), Springer Lecture Notes in Computer Science LNCS, Grenoble, France,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Lau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gayen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <article-title>A dataset of clinically generated visual questions and answers about radiology images</article-title>
          ,
          <source>Scientific data 5</source>
          (
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ben Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. V.</given-names>
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Vqa-med: Overview of the medical visual question answering task at imageclef 2019</article-title>
          ,
          <source>in: Proceedings of CLEF (Conference and Labs of the Evaluation Forum) 2019 Working Notes</source>
          ,
          <fpage>9</fpage>
          -
          <issue>12</issue>
          <year>September 2019</year>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ben Abacha</surname>
          </string-name>
          , E. Agichtein,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Pinter</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Demner-Fushman, Overview of the medical question answering task at trec 2017 liveqa</article-title>
          , in
          <source>: TREC</source>
          <year>2017</year>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mou</surname>
          </string-name>
          , E. Xing,
          <string-name>
            <given-names>P.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <article-title>Towards visual question answering on pathology images, in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th</article-title>
          <source>International Joint Conference on Natural Language Processing</source>
          (Volume
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>708</fpage>
          -
          <lpage>718</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hicks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Storås</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          , T. de Lange,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. L.</given-names>
            <surname>Thambawita</surname>
          </string-name>
          ,
          <article-title>Overview of imageclefmedical 2023 - medical visual question answering for gastrointestinal tract</article-title>
          ,
          <source>in: Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Daneshjou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Vodrahalli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Novoa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jenkins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Rotemberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Ko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Swetter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. E.</given-names>
            <surname>Bailey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Gevaert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Phung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yekrang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sahasrabudhe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Chiou</surname>
          </string-name>
          ,
          <article-title>Disparities in dermatology ai performance on a diverse, curated clinical image set</article-title>
          ,
          <source>Science Advances</source>
          <volume>8</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Groh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Harris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Soenksen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lau</surname>
          </string-name>
          , R. Han,
          <string-name>
            <given-names>A</given-names>
            .
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Koochek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Badri</surname>
          </string-name>
          ,
          <article-title>Evaluating deep neural networks trained on clinical images in dermatology with the fitzpatrick 17k dataset</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1820</fpage>
          -
          <lpage>1828</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>