<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>kdevqa at VQA-Med 2020: focusing on GLU-based classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hideo Umada</string-name>
          <email>umada@kde.cs.tut.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Masaki Aono</string-name>
          <email>aono@tut.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering, Toyohashi University of Technology</institution>
          ,
          <addr-line>Aichi</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Interpretation of medical images is a challenging research problem with increasing interest in medical applications of artificial intelligence. In particular, the ImageCLEF2020 visual question answering (VQA) task is expected to have applications such as a second opinion. The purpose of this research is to find an effective VQA-Med system method. We propose neural networks using the Gated Linear Unit for effective fusion of image and question features. Before training, we perform pre-processes and conduct pre-training. We apply so called “inpainting” to remove a logo or text embedded in images so that we attempt to extract image features with less noise. And we use the VQA-Med2019 dataset to train some of the weights of the proposed model. We consider the VQA task as a 332-dimensional classification task. The score of our proposed model turns out to be 0.314 in Accuracy and 0.350 in Bleu in VQA-Med2020 task.</p>
      </abstract>
      <kwd-group>
        <kwd>VQA-Med</kwd>
        <kwd>Visual Question Answering</kwd>
        <kwd>Classification</kwd>
        <kwd>Inpainting</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>With increasing interest in artificial intelligence to support clinical
decisionmaking and to improve patient engagement, the application to automated
medical image interpretation is currently getting much popularity. In particular, it is
expected that the second opinion provided by the automated system will enhance
the judgment of clinicians.</p>
      <p>Visual Question-Answering (VQA) is the task to generate a plausible answer
presented with an image-question pairs such as left of Fig. 1. The task requires
expertise in both natural language processing (NLP) and computer vision (CV)
so that researchers have been attempting to solve the problem from various
standpoints with Deep Neural Networks (DNN).</p>
      <p>
        In this paper, we describe our approach to ImageCLEF2020 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] visual
question answering (VQA) task [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] in medical domain at VQA such as right of Fig. 1.
The nature of medical images are quite different from general images such as
Imagenet [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] in many aspects. The knowledge on medical vocabulary seems to the
must to better understand both the questions and answers written in medical
terminologies.
      </p>
      <p>
        In the following, we first describe related work on VQA task and VQA-Med
task in Section 2, followed by the description of the dataset provided for
VQAMed2020 dataset in Section 3. In Section 4, we describe details of the method
we propose, and then of our experiments we have conducted in Section 5. We
finally conclude this paper in Section 6.
Convolution Neural Networks (CNNs) for image recognition, such as VGG and
ResNet, has been used extensively. Similarly, multiple Transformers for sentence
comprehension, such as BERT, has been getting popular recently. Accordingly,
feature extraction from pretrained neural network models, transfer learning and
fine tuning with the pretrained models have been actively investigated. Visual
Question Answering, or VQA, stands between image recognition and sentence
comprehension, and is regarded as a bridge application between them. Research
on VQA is actively carried out through the VQA Challenge using VQA v2.0 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
For example, P.Anderson et al. proposed DNN using Bottom-up Attention [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
obtained by using pretrained Faster R-CNN [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] which is one of CNN used for
object detection. In addition, as a VQA-Med task, there are competitions at
ImageCLEF2018 and 2019. Yan et al. [7] proposed dividing the dataset into
subcategories and attempted to solve the tasks by transforming the opriginal
problem into a classification problem with categories in VQA-Med2019.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Dataset of VQA-Med2020</title>
      <p>
        The VQA-Med2020 dataset consists of 5,000 pairs of medical image and
questionanswering. Specifically, the dataset consists of 4,000 training, 500 validation, and
500 test data. Most of the images in the VQA-Med2020 dataset are non-colored,
and they potentially include non-essential logos and texts. The question pattern
can be classified into 39 different types for training and validation data. In our
analysis, the top 10 patterns cover more than 94% of the total data. On the
other hand, there are 332 different answer patterns, and the top 10 patterns
cover approximately 12% of the total data. Table 1 summarizes top 5 frequent
questions and answers.
This section presents our methods in VQA-Med2020. The overview of our system
is illustrated in Fig 2 with the yellow layers having trainable weights. We deal
with VQA as a classification task of 332-dimension. All the images make some
pre-processes shown in subsection 4.1 and are later characterized by VGG. We
use VGG16 with batch normalization model [8] pretrained at Imagenet [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to
extract image features. However, since there is a large difference in distribution
between medical images and general images, fine-tuning is performed using
VQAMed2019 data [9]. We extract question features from pretrained BERT-Base,
Cased [10]. All the questions are then embedded by the WordPiece which is
used by BERT. On the other hand, all the answers are embedded by one-hot
encoding. Proposed model consists of DNN, and detailed architecture of DNN
is mentioned in subsection 4.3.
We process image normalization, standardization and inpainting [11]. We show
the flow of image pre-processing in Fig. 3. Firstly, all the images are grayscaling
and resizing at 255 × 255 shape. Secondly, we make masks for inpainting in
following four steps.
      </p>
      <p>– Casting laplacian filter on resized images.
– Binarizing images with a threshold 50.
– Closing images with kernel size 5.
– Opening images with kernel size 3.</p>
      <p>Thirdly, we cast inpainting images using the masks. We illustrate Fig. 4 where
you can compare the raw images with the inpainting images. Finally, we make
center crop images at 224 × 224 and normalize images as described in [8].
e
g
a
m
I
e
lrsyaa zseeR
c i
G</p>
      <p>Mask</p>
      <p>n
licapLaan iiitrzanao lisongC iepnngO</p>
      <p>B
g
n
it
n
i
a
p
n
I
trvceonBRG trropeneCC ilrzeaoNm
s
s
rcoe age
rep Im
P
Our networks, illustrated at left of the Fig. 5, classify the images of
VQAMed2019 into each attribute as a pre-training task. VQA-Med2019 dataset
consists of question-answering classified into 4 categories of Modality, Plane, Organ
and Abnormality per image. Similar to VQA-Med2020, question pattern is
typical, and the answer can be predicted from the image alone, almost regardless
of the question. We regard the answer of each category except Abnomality
attached to the image as the attributes of the image, and perform the task of
classifying each attribute of the image. However, for Modality, the classification
is too subdivided, so use the rough classification given in the VQA-Med2019
dataset paper [9]. The pre-training model consists of VGG16 and two FC layers.
We input the pre-processed image in Section 4.1 into VGG16 and obtain 4096-d
features and multiply the matrix W1 ∈ R4096×1000 by the 4096-d features, and
perform batchnorm [12], ReLU [13] and dropout [14] ratio= 0.5 at FC1, and
obtain 1000-d features. Then we multiply the matrix W2 ∈ R1000×27 by the 1000-d
features and obtain each attribute probability of softmax function. W1, W2 and
VGG16 have trainable parameters.
Proposed model, illustrated right of the Fig. 5, generates an answer as a
classification problem. Proposed model has VGG16 and FC1, 2, 3, 4, the weights of
FC1, 2, VGG16 is trained by a pre-training task. VGG16 weights are frozen, and
FC1, 2 are fine-tuned. Only FC3, 4 are trained from the beginning. FC3, 4 are
based on the Gated Linear Unit (GLU) [15], and FC3 consists of W3 ∈ R1000×332,
bias b ∈ R332. FC4 consists of matrix W4 ∈ R795×332, batchnorm and sigmoid.
Then outputs of FC3, 4 are fused by element-wise multiplication and obtained
using the softmax function to get the probability of answers.</p>
      <p>GLU masks each dimension of the features obtained by FC3 with a real
number from 0 to 1. The introduction of GLU is based on the consideration that
it is possible to narrow down the answers to some extent only from the attribute
information of question texts and images.</p>
      <sec id="sec-2-1">
        <title>Preprocess VGG16</title>
      </sec>
      <sec id="sec-2-2">
        <title>FC1+relu</title>
      </sec>
      <sec id="sec-2-3">
        <title>FC2+linear</title>
        <p>split
softmax</p>
      </sec>
      <sec id="sec-2-4">
        <title>Plane's Probability Train</title>
      </sec>
      <sec id="sec-2-5">
        <title>Transfer</title>
        <p>softmax
Modality's
Probability
softmax</p>
      </sec>
      <sec id="sec-2-6">
        <title>Organ's Probability 2020 Image Preprocess</title>
        <p>Experiments are also performed on the baseline model, proposed model and
proposed model without pre-training to show the usefulness of the proposed
model. We first describe the baseline model in subsection 5.1, followed by the
description of experimental conditions and evaluations in subsection 5.2. We
finally describe experimental results and computational scores in subsection 5.3.
The overview of the baseline model is shown in Fig. 6, feature fusion is used by
concatenation instead of GLU. Baseline model has VGG16, FC1, 2 and FC3.
FC3 has weights
W ∈ R1895×332. Other layers are in accordance with the proposed model in
subsection 4.3.
We train our models using the train set and verify training model by the
validation set. We determined the following hyper-parameters; loss function as cross
entropy loss, the number of epoch 300, batch size of 64, optimizer as RMSprop [16]
with a learning rate of 0.001. When training, we shuffle the training set order
for each epoch, and training images are randomly flipped left and right with
probability of 0.5.</p>
        <p>The VQA-Med task adopts two evaluation method, accuracy and BLEU [17].
BLEU score measures the similarity between the predicted and correct answers.
5.3</p>
        <p>Results
We submitted the baseline model and the proposed model and obtained the
evaluation on the test set.</p>
        <p>The results of our models show in Table 2. These results show that the fusion
method using GLU is superior to the concatenation fusion, and according to the
verification results, it can be seen that the accuracy is slightly improved by the
pre-training task. VQA-Med2020 competition result is shown in Table 3, and
our rank is 8th.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>In this research, we describe the models we submitted in ImageCLEF2020
VQAMed task. We proposed a model of feature connection by GLU and a pre-training
task by VQA-Med2019 dataset. We also introduced the removal of a logo and
texts using inpainting as image pre-processing. We show that fusion of functions
using GLU is superior to simple concatenation, and slightly improved score using
pre-training task. Proposed model scores 0.314 in accuracy and 0.350 in BLEU
in VQA-Med2020 task, and our rank is 8th.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgment</title>
      <p>A part of this research was carried out with the support of the Grant-in-Aid for
Scientific Research (B) (issue number 17H01746).
7. Xin Yan, Lin Li, Chulin Xie, Jun Xiao, and Lin Gu. Zhejiang university at imageclef
2019 visual question answering in the medical domain. In CLEF, 2019.
8. Torchvision.models. https://pytorch.org/docs/master/torchvision/models.html.
9. Asma Ben Abacha, Sadid A. Hasan, Vivek V. Datla, Joey Liu, Dina
DemnerFushman, and Henning Mu¨ller. VQA-Med: Overview of the medical visual question
answering task at imageclef 2019. In CLEF2019 Working Notes, CEUR
Workshop Proceedings, Lugano, Switzerland, September 09-12 2019. CEUR-WS.org
&lt;http://ceur-ws.org&gt;.
10. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert:
Pretraining of deep bidirectional transformers for language understanding, 2018. cite
arxiv:1810.04805Comment: 13 pages.
11. Alexandru Telea. An image inpainting technique based on the fast marching
method. Journal of Graphics Tools, 9(1):23–34, 2004.
12. Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep
network training by reducing internal covariate shift. In Proceedings of the 32nd
International Conference on International Conference on Machine Learning - Volume
37, ICML’15, page 448–456. JMLR.org, 2015.
13. Abien Fred Agarap. Deep learning using rectified linear units (relu), 2018. cite
arxiv:1803.08375Comment: 7 pages, 11 figures, 9 tables.
14. Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan
Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting.</p>
      <p>J. Mach. Learn. Res., 15(1):1929–1958, January 2014.
15. Yann N. Dauphin, Angela Fan, Michael Auli, and David Grangier. Language
modeling with gated convolutional networks. In Proceedings of the 34th International
Conference on Machine Learning - Volume 70, ICML’17, page 933–941. JMLR.org,
2017.
16. T. Tieleman and G. Hinton. Lecture 6.5—RmsProp: Divide the gradient by a
running average of its recent magnitude. COURSERA: Neural Networks for Machine
Learning, 2012.
17. Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method
for automatic evaluation of machine translation. In Proceedings of the 40th Annual
Meeting of the Association for Computational Linguistics, pages 311–318,
Philadelphia, Pennsylvania, USA, July 2002. Association for Computational Linguistics.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Bogdan</given-names>
            <surname>Ionescu</surname>
          </string-name>
          , Henning Mu¨ller, Renaud P´eteri, Asma Ben Abacha, Vivek Datla, Sadid A.
          <string-name>
            <surname>Hasan</surname>
          </string-name>
          , Dina Demner-Fushman, Serge Kozlovski, Vitali Liauchuk, Yashin Dicente Cid, Vassili Kovalev, Obioma Pelka,
          <string-name>
            <surname>Christoph M. Friedrich</surname>
          </string-name>
          , Alba Garc´ıa Seco de Herrera,
          <string-name>
            <surname>Van-Tu</surname>
            <given-names>Ninh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tu-Khiem</surname>
            <given-names>Le</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liting Zhou</surname>
            , Luca Piras,
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Riegler</surname>
          </string-name>
          , P˚al Halvorsen,
          <string-name>
            <surname>Minh-Triet</surname>
            <given-names>Tran</given-names>
          </string-name>
          , Mathias Lux, Cathal Gurrin,
          <string-name>
            <surname>Duc-Tien</surname>
          </string-name>
          Dang-Nguyen, Jon Chamberlain, Adrian Clark, Antonio Campello, Dimitri Fichou, Raul Berari, Paul Brie, Mihai Dogariu, Liviu Daniel S¸tefan, and Mihai Gabriel Constantin.
          <source>Overview of the ImageCLEF</source>
          <year>2020</year>
          :
          <article-title>Multimedia retrieval in lifelogging, medical, nature, and internet applications</article-title>
          .
          <source>In Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , volume
          <volume>12260</volume>
          <source>of Proceedings of the 11th International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ), Thessaloniki, Greece,
          <source>September 22-25 2020. LNCS Lecture Notes in Computer Science</source>
          , Springer.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Asma</given-names>
            <surname>Ben</surname>
          </string-name>
          <string-name>
            <surname>Abacha</surname>
          </string-name>
          , Vivek V.
          <article-title>Datla, Sadid A. Hasan, Dina Demner-Fushman, and Henning Mu¨ller. Overview of the vqa-med task at imageclef 2020: Visual question answering and generation in the medical domain</article-title>
          .
          <source>In CLEF 2020 Working Notes, CEUR Workshop Proceedings</source>
          , Thessaloniki, Greece,
          <source>September</source>
          <volume>22</volume>
          -25
          <year>2020</year>
          .
          <article-title>CEUR-WS.org</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei. ImageNet: A LargeScale Hierarchical Image</surname>
          </string-name>
          <article-title>Database</article-title>
          .
          <source>In CVPR09</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Yash</given-names>
            <surname>Goyal</surname>
          </string-name>
          , Tejas Khot, Douglas Summers-Stay,
          <article-title>Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering</article-title>
          .
          <source>In Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Peter</given-names>
            <surname>Anderson</surname>
          </string-name>
          , Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang.
          <article-title>Bottom-up and top-down attention for image captioning and visual question answering</article-title>
          .
          <source>In CVPR</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Ross</given-names>
            <surname>Girshick. Fast</surname>
          </string-name>
          r-cnn.
          <source>In Proceedings of the 2015 IEEE International Conference on Computer Vision</source>
          (ICCV),
          <source>ICCV '15</source>
          , pages
          <fpage>1440</fpage>
          -
          <lpage>1448</lpage>
          , Washington, DC, USA,
          <year>2015</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>