<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>G. Chen); liguanbin@mail.sysu.edu.cn (G. Li)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>SYSU-HCP at VQA-Med 2021: A Data-centric Model with Eficient Training Methodology for Medical Visual Question Answering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Haifan Gong (Co-first author)</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ricong Huang (Co-first author)</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guanqi Chen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guanbin Li (Corresponding author)</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science and Engineering, Sun Yat-sen University</institution>
          ,
          <addr-line>Guangzhou</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This paper describes our contribution to the Visual Question Answering Task in the Medical Domain at ImageCLEF 2021. We propose the method with a core idea that the model design and the training should best suit the feature of the data. Specifically, we design a hierarchical feature extraction structure to capture multi-scale features of medical images. To alleviate the issue of data limitation, we apply the mixup strategy for data augmentation during the training process. Based on the observation that there exist hard samples, we introduce the curriculum learning paradigm to resolve this issue. Last but not least, we apply label smoothing and ensemble training to avoid the model bias on the data. The proposed method achieves 1st place in the competition with 0.382 in accuracy and 0.416 in BLEU. Our code and model are available at https://github.com/Rodger-Huang/SYSU-HCP-at-ImageCLEF-VQA-Med-2021.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Medical visual question answering</kwd>
        <kwd>Classification</kwd>
        <kwd>Curriculum learning</kwd>
        <kwd>Mixup</kwd>
        <kwd>Label smoothing</kwd>
        <kwd>Ensemble learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Q: What is the primary</title>
        <p>abnormality in this image?</p>
      </sec>
      <sec id="sec-1-2">
        <title>A: Sickle cell anemia (a)</title>
      </sec>
      <sec id="sec-1-3">
        <title>Q: What is most alarming about this mri?</title>
      </sec>
      <sec id="sec-1-4">
        <title>A: Leptomeningeal sarcoid (b)</title>
      </sec>
      <sec id="sec-1-5">
        <title>Q: What is abnormal in the x-ray?</title>
      </sec>
      <sec id="sec-1-6">
        <title>A: Vacterl syndrome (c)</title>
        <p>design of visual representation to classify the abnormality of the medical images.</p>
        <p>
          To address the issue of data limitation, we expand the dataset with the data in the previous
VQA-Med competition (i.e., VQA-Med 2019 [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and VQA-Med 2020 [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]). To make full use
of the limited data, we apply mixup technology to create more samples. For eficient visual
feature extraction, we apply label smoothing to stabilize the training progress. Furthermore,
we discover the phenomenon that hard samples restrict the performance of the model and use
a curriculum learning-based loss function to resolve this issue. Last but not least, to prevent
the model from the infliction of intrinsic model bias, we propose to ensemble diferent types of
models in exchange for higher model accuracy(e.g., VGG [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], ResNet [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]).
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>A retrospective analysis is conducted in this section. We first literately review the research of
general VQA, then we conclude the methods of VQA in the medical domain.</p>
      <sec id="sec-2-1">
        <title>2.1. General VQA</title>
        <p>
          The prevailing VQA framework in the general domain is mainly composed of four components:
a visual encoder, a linguistic encoder, a cross-modal feature fusion module, and a classifier. The
visual encoders are usually based on deep CNNs such as ResNet [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], VGG [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], Faster-RCNN [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ],
etc. To extract the linguistic feature, researchers apply the Transformer or RNN based models
(e.g., Bert [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], LSTM [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]). The cross-modal feature fusion modules are dominant in current
VQA systems. To capture the relationship between images and languages, Fukui et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and
Kim et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] apply the compact bilinear pooling methods. In the meanwhile, Yang et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ],
Cao et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], and Anderson et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] investigate to design the model that focus on the
question-related region of the image. MLP-liked classifier is usually used to select the final
answer.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Medical VQA</title>
        <p>Diferent from general VQA, medical VQA usually sufers from limited data and the distinction
between common-sense knowledge and medical domain-specific knowledge. Thus, we first
discuss the current methods to alleviate the data limitation in medical VQA. After that, we
summarize the previous methods for the medical VQA.</p>
        <p>
          Data limitation in medical VQA. Data limitation is an unavoidable topic in the field of
medical image analysis, especially in the domain of medical VQA. To address this issue, Nguyen
et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] combine the meta-learning and the denoising auto-encoder to make use of large-scale
unlabeled data. Nevertheless, they neglect the compatibility between the visual concept and the
questions. Gong et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] propose a novel multi-task pre-train framework, the image encoders
of which are mandatory to not only learn the linguistic compatibility feature but also the vision
concept by performing the original task (i.e., classification &amp; segmentation) on the external
dataset. Still, the easiest way to overcome the data limitation is to collect more data as was
done by Chen et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. In this work, we not only collect the data from the previous VQA-Med
competition but also apply the mixup strategy for data augmentation.
        </p>
        <p>
          Previous methods on VQA-Med challenges. In the 2018 VQA-Med challenge [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], the top
three groups applied the analogous pipeline as the VQA in the general domain. Specifically,
they apply CNNs (i.e., ResNet-152 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], VGG [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], Inception-ResNet-v2 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]) for visual feature
extraction, LSTM [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] or Bi-LSTM for language modeling, and attention based feature fusion
modules (i.e., MFH [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ], SAM [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], BAN [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]).
        </p>
        <p>
          In the 2019 VQA-Med challenge [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], the leading three groups [
          <xref ref-type="bibr" rid="ref19 ref20 ref21">19, 20, 21</xref>
          ] used the
InceptionResnet-v2 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] or a tailor-designed VGG [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] to extract image feature. Bert [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is used to capture
the semantic of the questions, and MFH [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] is applied for feature fusion. The tailor-designed
VGG proposed by Yan et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] replaced the conventional global average pooling layer of
VGG with the hierarchical average pooling layer, which could eficiently capture the multi-scale
feature. Besides, it is worth noting that the top 3 teams apply the question classification method
to figure out the category of the questions.
        </p>
        <p>
          In the 2020 VQA-Med challenge [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], two [
          <xref ref-type="bibr" rid="ref22 ref23">22, 23</xref>
          ] of the top three teams abandon the
conventional VQA framework. Only Bumjun et al. [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] applies the conventional VQA framework,
which uses VGG backbone for visual feature extraction, BioBert [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] for question encoding, and
MFH [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] for feature fusion. The other two teams chose to direct classify the image with the deep
neural networks. The winner of VQA-Med-2020 [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] designed a question skeleton-based
approach to take full use of the linguistic feature, and integrate multi-scale and multi-architecture
models to achieve the best result.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Datasets</title>
      <p>
        In VQA-Med task [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] of ImageCLEF 2021 [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], the original dataset includes a training set of
4500 radiology images with 4500 question-answer (QA) pairs, a validation set of 500 radiology
images with 500 QA pairs, and a test set of 500 radiology images with 500 questions. These
questions focus on the abnormalities of medical images. Figure 1 shows three examples in the
dataset.
      </p>
      <p>
        Since the previous dataset [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] in the VQA-Med competition is allowed to use, we leverage the
Data preparation
      </p>
      <p>Network design
Efficient training</p>
      <p>Ensemble</p>
      <p>Collect training data from the</p>
      <p>previous competition
Design a hierarchical feature
extraction architecture to</p>
      <p>representation image
Ensemble the original CNN and
the hierarchical CNN against</p>
      <p>the data bias
Apply the mixup strategy for
data augmentation</p>
      <p>Introduce curriculum learning
to handle the hard samples</p>
      <p>Use label smoothing to avoid
the data over-fitting
abnormality subset from the VQA-Med 2019, the test set of VQA-Med 2020, and the validation set
of VQA-Med 2021 to extend the VQA-Med 2021 training set. The final training set is composed
of 6183 VQA pairs.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>As this competition focus on the questions about abnormalities, we discard the conventional
VQA framework and regard VQA as an image classification task. To make full use of the data,
we design a data-centric model which is shown in Figure 2. This framework mainly consists of
four parts: data preparation, network design, data-centric training methodologies, and model
ensemble. The data preparation is illustrated in Section 3. Other parts are detailed below.</p>
      <sec id="sec-4-1">
        <title>4.1. Network Architecture</title>
        <p>
          Inspired by Yan et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] and Aisha et al. [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] that multi-scale features contain more abundant
information of medical images, we design a hierarchical feature extraction architecture to
capture the multi-scale features of medical images. Diferent from the conventional
highlevel semantic feature representation architecture with fully-connected layers (Fig. 3 (a)), our
Adaptive Global Average Pooling
(a) High-level semantic representation
(b) Hierarchical multi-scale representation
proposed architecture replaces the fully-connected layers with hierarchical adaptive global
average pooling layers (Fig. 3 (b)). Compared with a similarity work [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] that uses global
average pooling to construct feature vector, the proposed method with adaptive global average
pooling is more flexible to receive arbitrary input size of the image. This hierarchically adaptive
global average pooling (HAGAP) structure is applied to ResNet-50 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], ResNeSt-50 [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], VGG-16
and VGG-19 [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] to extract image feature.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Data-centric eficient training</title>
        <p>In this part, we introduce three eficient strategies to better utilize the training data according
to their characteristic.</p>
        <p>
          Mixup. To alleviate the issue of data limitation in medical image representation learning, we
adopt a simple yet efective data augmentation method called Mixup [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. Given two samples
(, ) and ( ,  ), we create a new image ˆ with label ˆ by linear interpolation with the
following operation:
ˆ =   + (1 −  )
ˆ =   + (1 −  )
where  ∈ [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] is a random value drawn from the (,  ) distribution with the
hyperparameter  = 0.2. It is worth noting that we only use the newly created images during the
training process.
        </p>
        <p>
          Curriculum learning. Based on the observation that in the training set of this competition,
one disease could occur in various images modalities (e.g., CT, MRI.). Some modalities are of
(1)
numerous samples while others are of few. Thus, the imaging modality of the diseases that occur
infrequently is hard to learn. Furthermore, the training set is unavoidable to contain noise. To
resolve these issues, we introduce the idea of curriculum learning [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] into the training process.
To simplify this process, we apply the SuperLoss [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ], which automatically down-weights the
hard samples with a larger loss.
        </p>
        <p>
          Label smoothing. The label smoothing methodology is first proposed to train the
InceptionV2 [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ] network. It works by adjusting the probability of the target label by:
 =
︂{
1
        </p>
        <p>− 
/( − 1)
if  = 
otherwise
(2)
where  is a small constant,  is the number of classes, and  denotes the possibility of category
. As label smoothing groups the representations of the examples from the same class into tight
clusters, the model could achieve a better generalization ability.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Model Ensemble</title>
        <p>
          As the model unavoidably contains bias, we apply multi-architecture ensemble to further
improve the model performance. Comparing to the winner of VQA-Med 2020 [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] that takes
more than 30 models to an ensemble, our best submission only contains 8 models by taking the
advantage of the HAGAP structure. Specifically, the 8 models are ResNet-50, ResNeSt-50,
VGG16, VGG-19, ResNet-50-HAGAP, ResNeSt-50-HAGAP, VGG-16-HAGAP and VGG-19-HAGAP.
        </p>
        <p>224. All backbones are initialized with the ImageNet pre-trained weight.</p>
        <p>Since the data is limited, we do not set the validation set during the training process and select
256 while the input size of VGG based
networks is 224 ×
The input size of ResNet-based networks is 256 ×
the model of a fixed epoch for evaluation.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <sec id="sec-5-1">
        <title>5.1. Implementation details</title>
        <p>As for training data, we leverage the data in Section 3. The models for our best submission are
trained with the combination of mixup loss, SuperLoss, and Label smoothing loss. We used the
SGD optimizer with momentum set to 0.9. The initial learning rate is set to 1e-3, and the weight
decay is 5e-4. All models are trained for 60 epochs, and we select the model for inference on a
ifxed epoch (e.g., 50).</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Evaluation</title>
        <p>
          The VQA-Med competition applies accuracy and BLEU[
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] as the evaluation metrics. Accuracy
is calculated as the number of correct predicted answers among all answers. BLEU measures
the similarity between the predicted answers and ground truth answers. As shown in Fig.4, we
achieved an accuracy of 0.382 and a BLEU score of 0.416 in the VQA-Med-2021 test set, which
won the first place in this competition.
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Ablation study</title>
        <p>To demonstrate the efectiveness of our proposed model, we conduct ablation study with the
VGG-16 network, which is shown in Table 1. Specifically, the input image size is 224 × 224.
The training set contains 5664 images while the validation set contains 500 images. The original
VGG-16 network is set as the baseline, which achieves an accuracy of 66.6%. We utilize the
mixup strategy for data augmentation, which surpasses the baseline by 1.8%. Furthermore, we
adopt label smoothing to avoid the over-fitting of the model, which improves the accuracy to
68.8%. After that, we apply hierarchical architecture to represent the feature of the medical
image, and introduce the curriculum learning paradigm into the framework, which brings an
accuracy gain of 0.4%. With the eforts mentioned above, we achieve 69.2% accuracy on the</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>
        The VQA-Med is challenging task due to the limited data. As this work mainly focused on
distinguish the abnormality between the medical images, we focus on designing the training
scheduler and the feature extract module to make better use of the limited data. It is worth noting
that as the training set, the validation set, and test set may not obey the same distribution, the
Table 1 is of limited value. In other words, Label smoothing, HAGAP, and curriculum learning
may be efective in the test set, but it not brings significant improvement on the validation set.
For the same distribution inconsistent issue, we directly classify the images rather than use the
long-tailed based methods [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Though we achieves 1 st place at this competition, our score is
not high and there is still a long way to go to achieve applicable medical VQA.
      </p>
      <p>For the future works in the medical VQA, we may digger deeper into better feature
representation of image or words with the help of large amount unlabeled data. Besides, generating the
answer word by word rather than regard the answer as a label is more valuable research topic.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>In this paper, we describe our participation at the ImageCLEF 2021 VQA-Med challenge.
Considering most of the questions are about abnormality, we abandon the conventional complex
cross-modal fusion methodologies. With the firm brief that the characteristics of the data
should be fully considered in the construction of the model, we design a data-centric
model with eficient training strategies. Our proposed method achieves the best results among
all participating groups with an accuracy of 0.382 and a BLEU score of 0.416.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Farri</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Lungren</surname>
          </string-name>
          ,
          <article-title>Overview of imageclef 2018 medical domain visual question answering task</article-title>
          ,
          <source>in: Working Notes of CLEF 2018 - Conference and Labs of the Evaluation Forum</source>
          , Avignon, France,
          <source>September 10-14</source>
          ,
          <year>2018</year>
          , volume
          <volume>2125</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ben Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. V.</given-names>
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Vqa-med: Overview of the medical visual question answering task at imageclef 2019</article-title>
          , in: Working Notes of CLEF 2019 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Lugano, Switzerland, September 9-
          <issue>12</issue>
          ,
          <year>2019</year>
          , volume
          <volume>2380</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ben Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. V.</given-names>
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain</article-title>
          ,
          <source>in: Working Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum</source>
          , Thessaloniki, Greece,
          <source>September 22-25</source>
          ,
          <year>2020</year>
          , volume
          <volume>2696</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          , in: Y. Bengio, Y. LeCun (Eds.),
          <source>3rd International Conference on Learning Representations, ICLR</source>
          <year>2015</year>
          , San Diego, CA, USA, May 7-
          <issue>9</issue>
          ,
          <year>2015</year>
          , Conference Track Proceedings,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <surname>Faster</surname>
          </string-name>
          r-cnn:
          <article-title>Towards real-time object detection with region proposal networks</article-title>
          ,
          <source>in: Advances in neural information processing systems</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>91</fpage>
          -
          <lpage>99</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <article-title>Long short-term memory</article-title>
          ,
          <source>Neural computation 9</source>
          (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fukui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. H.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Darrell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          ,
          <article-title>Multimodal compact bilinear pooling for visual question answering and visual grounding</article-title>
          ,
          <source>arXiv preprint arXiv:1606</source>
          .
          <year>01847</year>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>J.-H. Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Jun</surname>
          </string-name>
          , B.-T. Zhang,
          <article-title>Bilinear attention networks</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1564</fpage>
          -
          <lpage>1574</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smola</surname>
          </string-name>
          ,
          <article-title>Stacked attention networks for image question answering</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Visual question reasoning on general dependency tree</article-title>
          ,
          <source>in: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Buehler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Teney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , S. Gould,
          <string-name>
            <surname>L. Zhang,</surname>
          </string-name>
          <article-title>Bottom-up and top-down attention for image captioning and visual question answering</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>6077</fpage>
          -
          <lpage>6086</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , T.-T. Do,
          <string-name>
            <given-names>B. X.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Do</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tjiputra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. D.</given-names>
            <surname>Tran</surname>
          </string-name>
          ,
          <article-title>Overcoming data limitation in medical visual question answering</article-title>
          , in: International Conference on Medical Image Computing and
          <string-name>
            <surname>Computer-Assisted</surname>
            <given-names>Intervention</given-names>
          </string-name>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>522</fpage>
          -
          <lpage>530</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>H.</given-names>
            <surname>Gong</surname>
          </string-name>
          , G. Chen, S. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Cross-modal self-attention with multi-task pretraining for medical visual question answering</article-title>
          ,
          <source>in: ACM International Conference on Multimedia Retrieval (ICMR)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Li, HCP-MIC at vqa-med 2020: Efective visual representation for medical visual question answering</article-title>
          ,
          <source>in: Working Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum</source>
          , Thessaloniki, Greece,
          <source>September 22-25</source>
          ,
          <year>2020</year>
          , volume
          <volume>2696</volume>
          <source>of CEUR Workshop Proceedings</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>C.</given-names>
            <surname>Szegedy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Iofe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vanhoucke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Alemi</surname>
          </string-name>
          ,
          <article-title>Inception-v4, inception-resnet and the impact of residual connections on learning</article-title>
          , in: S.
          <string-name>
            <given-names>P.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          Markovitch (Eds.),
          <source>Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9</source>
          ,
          <year>2017</year>
          , San Francisco, California, USA, AAAI Press,
          <year>2017</year>
          , pp.
          <fpage>4278</fpage>
          -
          <lpage>4284</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Tao, Multi-modal factorized bilinear pooling with co-attention learning for visual question answering</article-title>
          ,
          <source>in: IEEE International Conference on Computer Vision</source>
          , ICCV 2017, Venice, Italy,
          <source>October 22-29</source>
          ,
          <year>2017</year>
          , IEEE Computer Society,
          <year>2017</year>
          , pp.
          <fpage>1839</fpage>
          -
          <lpage>1848</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          , L. Gu, Zhejiang university at imageclef 2019 visual
          <article-title>question answering in the medical domain</article-title>
          ,
          <source>in: Working Notes of CLEF 2019 - Conference and Labs of the Evaluation Forum, Lugano, Switzerland, September</source>
          <volume>9</volume>
          -
          <issue>12</issue>
          ,
          <year>2019</year>
          , volume
          <volume>2380</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <article-title>TUA1 at imageclef 2019 vqa-med: a classification and generation model based on transfer learning</article-title>
          ,
          <source>in: Working Notes of CLEF 2019 - Conference and Labs of the Evaluation Forum, Lugano, Switzerland, September</source>
          <volume>9</volume>
          -
          <issue>12</issue>
          ,
          <year>2019</year>
          , volume
          <volume>2380</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Vu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sznitman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Nyholm</surname>
          </string-name>
          , T. Löfstedt,
          <article-title>Ensemble of streamlined bilinear visual question answering models for the imageclef 2019 challenge in the medical domain</article-title>
          ,
          <source>in: Working Notes of CLEF 2019 - Conference and Labs of the Evaluation Forum, Lugano, Switzerland, September</source>
          <volume>9</volume>
          -
          <issue>12</issue>
          ,
          <year>2019</year>
          , volume
          <volume>2380</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEURWS.org,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. van den Hengel</surname>
          </string-name>
          , J. Verjans, AIML at vqa-med
          <year>2020</year>
          :
          <article-title>Knowledge inference via a skeleton-based sentence mapping approach for medical domain visual question answering</article-title>
          ,
          <source>in: Working Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum</source>
          , Thessaloniki, Greece,
          <source>September 22-25</source>
          ,
          <year>2020</year>
          , volume
          <volume>2696</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Al-Sadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Al-Theiabat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Al-Ayyoub</surname>
          </string-name>
          ,
          <article-title>The inception team at vqa-med 2020: Pretrained VGG with data augmentation for medical VQA and VQG</article-title>
          , in: Working Notes of CLEF 2020 -
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          , Thessaloniki, Greece,
          <source>September 22-25</source>
          ,
          <year>2020</year>
          , volume
          <volume>2696</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>B.</given-names>
            <surname>Jung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gu</surname>
          </string-name>
          , T. Harada, bumjun_jung at vqa-med
          <year>2020</year>
          :
          <article-title>VQA model based on feature extraction and multi-modal feature fusion</article-title>
          ,
          <source>in: Working Notes of CLEF 2020 - Conference and Labs of the Evaluation Forum</source>
          , Thessaloniki, Greece,
          <source>September 22-25</source>
          ,
          <year>2020</year>
          , volume
          <volume>2696</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. H.</given-names>
            <surname>So</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <article-title>Biobert: a pre-trained biomedical language representation model for biomedical text mining</article-title>
          ,
          <source>Bioinform</source>
          .
          <volume>36</volume>
          (
          <year>2020</year>
          )
          <fpage>1234</fpage>
          -
          <lpage>1240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ben Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sarrouti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Overview of the vqa-med task at imageclef 2021: Visual question answering and generation in the medical domain</article-title>
          ,
          <source>in: CLEF 2021 Working Notes, CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Bucharest, Romania,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Péteri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sarrouti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kozlovski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Liauchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dicente</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G. S.</given-names>
            <surname>Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jacutprakart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Berari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tauteanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fichou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Brie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dogariu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. D.</given-names>
            <surname>Ştefan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chamberlain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Campello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. A.</given-names>
            <surname>Oliver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Moustahfid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Popescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Deshayes-Chossart</surname>
          </string-name>
          ,
          <article-title>Overview of the ImageCLEF 2021: Multimedia retrieval in medical, nature, internet and social media applications, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>Proceedings of the 12th International Conference of the CLEF Association, LNCS Lecture Notes in Computer Science</source>
          , Springer, Bucharest, Romania,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Manmatha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Smola</surname>
          </string-name>
          , Resnest:
          <article-title>Split-attention networks</article-title>
          , CoRR abs/
          <year>2004</year>
          .08955 (
          <year>2020</year>
          ). arXiv:
          <year>2004</year>
          .08955.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cisse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. N.</given-names>
            <surname>Dauphin</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Lopez-Paz, mixup: Beyond empirical risk minimization</article-title>
          ,
          <source>arXiv preprint arXiv:1710.09412</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Louradour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Collobert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Weston</surname>
          </string-name>
          ,
          <article-title>Curriculum learning</article-title>
          ,
          <source>in: Proceedings of the 26th Annual International Conference on MachineLearning, ICML</source>
          <year>2009</year>
          , Montreal, Quebec, Canada, June 14-18,
          <year>2009</year>
          , volume
          <volume>382</volume>
          of ACM International Conference Proceeding Series, ACM,
          <year>2009</year>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>T.</given-names>
            <surname>Castells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Weinzaepfel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Revaud</surname>
          </string-name>
          ,
          <article-title>Superloss: A generic loss for robust curriculum learning</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems</source>
          <year>2020</year>
          ,
          <article-title>NeurIPS 2020</article-title>
          , December 6-
          <issue>12</issue>
          ,
          <year>2020</year>
          , virtual,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>C.</given-names>
            <surname>Szegedy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vanhoucke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Iofe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shlens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wojna</surname>
          </string-name>
          ,
          <article-title>Rethinking the inception architecture for computer vision</article-title>
          , in:
          <source>2016 IEEE Conference on Computer Vision</source>
          and Pattern Recognition,
          <string-name>
            <surname>CVPR</surname>
          </string-name>
          <year>2016</year>
          ,
          <string-name>
            <surname>Las</surname>
            <given-names>Vegas</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NV</surname>
          </string-name>
          , USA, June 27-30,
          <year>2016</year>
          , IEEE Computer Society,
          <year>2016</year>
          , pp.
          <fpage>2818</fpage>
          -
          <lpage>2826</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>K.</given-names>
            <surname>Papineni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Roukos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ward</surname>
          </string-name>
          , W. Zhu,
          <article-title>Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics</article-title>
          ,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          ,
          <year>2002</year>
          , pp.
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>