<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>bumjun jung at VQA-Med 2020: VQA model based on feature extraction and multi-modal feature fusion</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bumjun Jung</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lin Gu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tatsuya Harada</string-name>
          <email>haradag@mi.t.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>RIKEN AIP</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>The University of Tokyo</institution>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the submission of University of Tokyo for Medical Domain Visual Question Answering (VQA-Med) task [3] at ImageCLEF 2020 [11]. The data set for the task mostly consists of Medical Images and Question Answer pair considering the abnormality appeared in the images. We extract visual features by VGG16 network [16] with Global Average Pooling (GAP) [14]. Compared to the model [18] that ranked rst in last year's competition that used BERT [6] model to encode semantic features of questions, we used bioBERT model [13], which is a BERT model pre-trained by biomedical textual data. We also apply multi-modal Factorized High-order (MFH) Pooling [20] with coattention which shows higher performance than Multi-modal Factorized Bilinear (MFB) Pooling [19] used in [18], to fuse two feature modalities. The fused features are then fed to a decoder to predict the answer in a manner of classi cation. The score of our model is 0:466 in accuracy, 0:502 in BLEU score, and ranked 3rd among all the participating teams in the VQA-Med task [3] at ImageCLEF 2020 [11].</p>
      </abstract>
      <kwd-group>
        <kwd>Visual Question Answering</kwd>
        <kwd>Medical Imagery</kwd>
        <kwd>Global Av- erage Pooling</kwd>
        <kwd>bioBERT</kwd>
        <kwd>Multi-modal Factorized High-order Pooling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>With many achievements and rapid progress in the eld of Arti cial Intelligence
(AI) related to Computer Vision (CV) and Natural Language Processing (NLP),
recently the AI technology is applied in the medical domain to analyze the
pathological images and medical reports. To be speci c, it is used to detect
abnormalities or symptoms shown in the images or to generate explanations
regarding the medical images.</p>
      <p>Visual Question Answering (VQA) task involves both CV and NLP
techniques to process the data. VQA data set is comprised of both Images and</p>
      <p>Question Answer (QA) pairs about the medical images. The images and
questions become the inputs to VQA system, whose goal is to predict the answers
for the given questions.</p>
      <p>
        Large-scale data sets of VQA for general domain [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] exist and there
are many advanced models and techniques that e ectively solve the task. With
increasing interest in applying AI technology in the medical eld, VQA in medical
domain is drawing attention due to the importance of supporting the doctors'
clinical decision and enhancing the patients' understanding of their conditions
from the medical images especially in patient-centered medical care.
      </p>
      <p>
        VQA for medical domain is a challenging task compared to that of general
domain. First, since the cost of collecting valid data is high, valid medical data
for training are limited compared to those in general domain such as [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
where hundreds of thousands of images and QA pairs are available. Second,
the vocabulary used in QA pairs or medical reports is quite distinct from the
language used in daily life.
      </p>
      <p>
        VQA-Med data set provided by ImageCLEF 2020 consists of 4,000 training
set with radiology images and QA pairs, 500 validation set, and 500 test set only
with questions without answers. As illustrated in Fig.1, VQA-Med 2020 data set
generally asks questions related to the abnormalities shown in the images.
The proposed framework in this paper is shown in Fig.2 and can be described
as following steps with Fig2: 1. VGG16 network [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] with GAP [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] (Green) is
used to extract image features from input image 2. bioBERT model [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] (Blue) is
used to capture the semantic of questions and encode it into textual features. 3.
Visual features and textual features are fused by fusion mechanism called MFH
Pooling [20] (Purple) 4. Co-attention mechanism (Purple) is applied to both
visual and textual features to focus on particular image regions based on the
question features and vice versa. 5. Finally, the features, fused by MFH Pooling
[20] with co-attention, are fed to the decoder to predict the answer in a manner
of classi cation.
      </p>
      <p>*+,-.$'/./</p>
      <p>!"#$%&amp;'()
&gt;?5:#)4#5%"-#)</p>
      <p>*@%-5+%'.ABBC#%DBE83
01(234+3(#"53)
4#5%"-#)4"$&amp;'(
.647)8''9&amp;(:);&amp;%&lt;)</p>
      <p>+'=5%%#(%&amp;'(3
!"#$%&amp;'()*(+',#./&amp;'0*123
6"3)#*-.4+3(#"53)
!"#$"#%&amp;'())*+*,(#*-./</p>
      <p>
        The improvements and contributions we made compared to the method [18]
comprises of three points: First, for the visual feature extraction, the dimension of
extracted feature is reduced to 1472 from 1984 to avoid over- tting problem while
maintaining the quantity of information in extracted features. Second, bioBERT
model [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] is used to extract textual features instead of BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] model used in
[18]. While bioBERT [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] model has the same network structure as BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] ,
the data used in pre-training of bioBERT [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] model were biomedical texts which
are di erent from the text data of general domain used in pre-training of BERT
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] . Third, MFH Pooling [20] is used to fuse the visual and textual features
which is the advanced version of MFB Pooling [19] used in [18]. MFH Pooling
[20] method achieves higher performance than MFB Pooling [19] method.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>
        There has been much of developments in methods and models used in open
domain VQA [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. For these tasks, deep learning models for image
processing based on deep Convolution Neural Networks (CNNs) such as VGGNet [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ],
ResNet [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] are frequently used to extract image features after pre-trained by
large-scale data set in the general domain such as Image net data set [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Regarding the question information processing, models for NLP, which are based
on Recurrent Neural Networks (RNNs) such as long short-term memory (LSTM)
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and grated recurrent units (GRU) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], are frequently used not only to
encode the textual features but also to generate answers as output. Similarlly, NLP
models such as BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] pre-trained by large-scale data are applied to extract
semantic features from text data.
      </p>
      <p>
        Attention mechanism and multi-modal feature fusion methods are the
important factors of VQA system since VQA is a multidisciplinary task that involves
both CV and NLP approaches. Attention mechanisms have been successfully
employed in image captioning [17] and NLP models such as BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] also have
adopted self-attention transformers in the network structure. Multi-modal
feature fusion is essential for VQA task since it combines the information from both
modalities to predict the right answer. Fusion techniques have evolved
starting from hierarchical co-attention model (Hie+CoAtt) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] which employs
coattention mechanism using element-wise summations, concatenation, and fully
connected layers. Multimodal Compact Bilinear (MCB) pooling [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] computes
the outer product between two features to represent every information from
the features and this also reduces the computational cost compared to simple
outer product calculation. Also, Multi-modal Low-rank Bilinear (MLB) pooling
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] generate output features with lower dimensions and models with fewer
parameters compared to MCB Pooling. MFB Pooling [19] method xed the slow
convergence rate of MLB Pooling as well as the part that it is sensitive to the
hyper-parameters. MFH Pooling [20] extends the MFB Pooling to a generalized
high-order setting to fuse the multi-modal features more e ectively.
      </p>
      <p>
        [18] describes the rst rank method of VQA-Med challenge at ImageCLEF
2019 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] which uses VGG16 network with GAP to extract visual features from
input images and BERT model to extract textual features from questions. Those
extracted features are fused by MFB Pooling with co-attention method and the
fused features are used to predict the answer in a manner of classi cation.
      </p>
      <p>
        In addition, bioBERT [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] model is a pre-trained language representation
model for the biomedical domain which shares the same network structure with
BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. bioBERT is pre-trained on biomedical domain corpora and it can
capture semantic features of biomedical texts such as medical reports more
effectively than BERT.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>
        This section describes the whole pipeline of our VQA model submitted for
ImageCLEF 2020 VQA-Med task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. As shown in Fig.2, rst, from the input image
and question the image features and question features are extracted by Image
feature extractor and Question encoder. The extracted features are then fused
with feature fusion method with co-attention to a classi cation network for the
answer selecting.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Image feature extractor</title>
        <p>
          In our VQA framework, VGG16 network pre-trained by ImageNet data set [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
is used to extract image features. GAP [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] strategy is applied with VGG16
network to prevent over- tting problem. The GAP method take the average of last
convolution outputs of each layers that have di erent number of channels. If the
input image shape is 224x224x3 as in our model, the output shapes of VGG16
network layers are as follows: 224x224x64, 112x112x128, 56x56x256, 28x28x512,
14x14x512, 7x7x512. The last number of each output shape is the channel size of
the convolution layer outputs. After taking the average of the outputs by channel,
the extracted features' dimension become the channel size of the layers. Those
features are concatenated to form a 1472-dimensional (64+128+256+512+512=1472)3
vector and it is used as image features and fed to the next network.
3.2
        </p>
        <p>
          Question encoder
bioBERT [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] is used to extract the semantic features of the given questions.
bioBERT is pre-trained with biomedical text and has the same network
structure as BERT [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. bioBERT largely outperforms BERT and previous
state-ofthe-art models in a variety of biomedical text mining tasks when pre-trained
on biomedical corpora. To extract the textual features that can represent the
question sentences, we average the last layer of bioBERT-base model to obtain
a 768-dimensional question feature vector.
3.3
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Feature fusion with co-attention</title>
        <p>Fusing multi-modal features is essential and important technique to improve the
performance of VQA model. As mentioned in section 2, Multi-modal
Factorized High-order (MFH) Pooling [20] method can fuse multi-modal features with
less computational cost and improved performance. Co-attention mechanism can
help the model to learn the importance of each part in both visual and textual
features. It can use the relative information from both modalities to learn which
parts of the features are important and to ignore the irrelevant information. We
therefore employ the MFH Pooling with co-attention to fuse visual and textual
features.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Training</title>
      <p>Our model is trained for 990 epochs on one Quadro GV100 for about 4 hours.
This section describes the detailed process and parameters used in the actual
training.
4.1</p>
      <sec id="sec-4-1">
        <title>Train data extension</title>
        <p>
          Besides the data set provided in the ImageCLEF 2020 VQA-Med task [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], we
also took advantage of VQA-Med data set of ImageCLEF 2019 [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. From the
3 The last convolution layer output that shaped 7x7x512 is precluded when extracting
features because it is only the output of Max pooling layer that represents the same
information as the former layer output.
data set in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], only the data comprised of QA pair existing in VQA-Med 2020
data set is used to train the model. 978 pairs in training set and 143 pairs in
validation set from [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] are used to extend the VQA-Med 2020 data set.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Hyper parameters</title>
        <p>Hyper parameters are set according to the performance on the validation data
set. We used Binary cross-entropy loss as loss function, ADAM optimizer with
initial learning rate of 3e-5 and the L1 regularization with co-e cient of 5e-11.
MFH Pooling [20] is used with default parameters explained in [20] except for
the dropout co-e cient which is set to 0.85 to prevent the over- tting problem.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>Two evaluation methods were adopted to VQA-Med 2020 competition, accuracy
(strict) and BLEU score. The accuracy measures the ratio of correct prediction
and the BLEU score measures the similarity between the real answer and
predicted answer. The max validation accuracy of our model was 0.612 and the
accuracy transition by training epoch is shown in Fig.3. For the actual training
of submitted model, the validation accuracy became 1.0 since the validation data
set was also included during the training.</p>
      <p>
        Among the 5 valid submissions, the model described in this paper that
comprises of VGG16 (with GAP) + bioBERT + MFH Pooling (with co-attention)
achieved the accuracy score of 0.466 and BLEU score of 0.502 for the test data
set. Our submission took 3rd place in the competition. Fig.4 shows the
leaderboard page of VQA-Med competition.
This paper describes the model submitted in ImageCLEF 2020 VQA-Med
challenge. Our model ranked 3rd place and achieved accuracy of 0.466 and BLEU
score of 0.502 on test data set. We applied bioBERT [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] model to extract textual
features which has stronger performance on encoding biomedical texts compared
BERT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Also MFH Pooling [20] is used to fuse the multi-modal features that
extends the MFB Pooling [19] to a generalized high-order setting to perform
better. For the future work, we will continue to improve the current network
and apply it to other data set or tasks.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgement</title>
      <p>This work was supported by JSPS KAKENHI Grant Number JP20H05556,
JST AIP Acceleration Research Grant Number JPMJCR20U3 and JST ACT-X
Grant Number JPMJAX190D. We would like to thank Kohei Uehara, Ryohei
Shimizu, Dr. Hiroaki Yamane, and Dr. Yusuke Kurose for helpful discussion.
17. Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R.,
Bengio, Y.: Show, attend and tell: Neural image caption generation with visual
attention. In: International conference on machine learning. pp. 2048{2057 (2015)
18. Yan, X., Li, L., Xie, C., Xiao, J., Gu, L.: Zhejiang university at imageclef 2019 visual
question answering in the medical domain. In: CLEF (Working Notes) (2019)
19. Yu, Z., Yu, J., Fan, J., Tao, D.: Multi-modal factorized bilinear pooling with
coattention learning for visual question answering. In: Proceedings of the IEEE
international conference on computer vision. pp. 1821{1830 (2017)
20. Yu, Z., Yu, J., Xiang, C., Fan, J., Tao, D.: Beyond bilinear: Generalized multimodal
factorized high-order pooling for visual question answering. IEEE transactions on
neural networks and learning systems 29(12), 5947{5959 (2018)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abacha</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Datla</surname>
            ,
            <given-names>V.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demner-Fushman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Muller, H.:
          <article-title>Vqa-med: Overview of the medical visual question answering task at imageclef 2019</article-title>
          . In: CLEF (Working Notes) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Antol</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            , J., Mitchell,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Lawrence Zitnick,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          : Vqa:
          <article-title>Visual question answering</article-title>
          .
          <source>In: Proceedings of the IEEE international conference on computer vision</source>
          . pp.
          <volume>2425</volume>
          {
          <issue>2433</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.V.</given-names>
            ,
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.A.</given-names>
            ,
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          , Muller, H.:
          <article-title>Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain</article-title>
          .
          <source>In: CLEF 2020 Working Notes. CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Thessaloniki,
          <source>Greece (September</source>
          <volume>22</volume>
          -25
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chung</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gulcehre</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Empirical evaluation of gated recurrent neural networks on sequence modeling</article-title>
          .
          <source>arXiv preprint arXiv:1412.3555</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fei-Fei</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Imagenet: A largescale hierarchical image database</article-title>
          .
          <source>In: 2009 IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>248</volume>
          {
          <fpage>255</fpage>
          .
          <string-name>
            <surname>Ieee</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fukui</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohrbach</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darrell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohrbach</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Multimodal compact bilinear pooling for visual question answering and visual grounding</article-title>
          .
          <source>arXiv preprint arXiv:1606</source>
          .
          <year>01847</year>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khot</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Summers-Stay</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parikh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Making the v in vqa matter: Elevating the role of image understanding in visual question answering</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <volume>6904</volume>
          {
          <issue>6913</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>He</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , J.:
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>770</volume>
          {
          <issue>778</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9(8)</source>
          ,
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Peteri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.A.</given-names>
            ,
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Kozlovski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Liauchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Cid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.D.</given-names>
            ,
            <surname>Kovalev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.M.</given-names>
            ,
            <surname>de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.G.S.</given-names>
            ,
            <surname>Ninh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.T.</given-names>
            ,
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.K.</given-names>
            ,
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Piras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Riegler</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , l Halvorsen,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.T.</given-names>
            ,
            <surname>Lux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gurrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Dang-Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.T.</given-names>
            ,
            <surname>Chamberlain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Campello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Fichou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Berari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Brie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Dogariu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Stefan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.D.</given-names>
            ,
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.G.</surname>
          </string-name>
          :
          <article-title>Overview of the ImageCLEF 2020: Multimedia retrieval in medical, lifelogging, nature, and internet applications</article-title>
          .
          <source>In: Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the 11th International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ), vol.
          <volume>12260</volume>
          .
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Thessaloniki,
          <source>Greece (September</source>
          <volume>22</volume>
          - 25
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>S.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kwak</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heo</surname>
            ,
            <given-names>M.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ha</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          , Zhang, B.T.:
          <article-title>Multimodal residual learning for visual qa</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>361</volume>
          {
          <issue>369</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoon</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>So</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kang</surname>
          </string-name>
          , J.:
          <article-title>Biobert: a pre-trained biomedical language representation model for biomedical text mining</article-title>
          .
          <source>Bioinformatics</source>
          <volume>36</volume>
          (
          <issue>4</issue>
          ),
          <volume>1234</volume>
          {
          <fpage>1240</fpage>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Network in network</article-title>
          .
          <source>arXiv preprint arXiv:1312.4400</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parikh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Hierarchical question-image co-attention for visual question answering</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>289</volume>
          {
          <issue>297</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Simonyan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          .
          <source>arXiv preprint arXiv:1409.1556</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>