<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Leveraging Medical Visual Question Answering with Supporting Facts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tomasz Kornuta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Deepta Rajan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chaitanya Shivade</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexis Asseman</string-name>
          <email>alexis.asseman@ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ahmet S. Ozcan</string-name>
          <email>asozcang@us.ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IBM Research AI, Almaden Research Center</institution>
          ,
          <addr-line>San Jose</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this working notes paper, we describe IBM Research AI (Almaden) team's participation in the ImageCLEF 2019 VQA-Med competition. The challenge consists of four question-answering tasks based on radiology images. The diversity of imaging modalities, organs and disease types combined with a small imbalanced training set made this a highly complex problem. To overcome these di culties, we implemented a modular pipeline architecture that utilized transfer learning and multitask learning. Our ndings led to the development of a novel model called Supporting Facts Network (SFN). The main idea behind SFN is to cross-utilize information from upstream tasks to improve the accuracy on harder downstream ones. This approach signi cantly improved the scores achieved in the validation set (18 point improvement in F1 score). Finally, we submitted four runs to the competition and were ranked seventh.</p>
      </abstract>
      <kwd-group>
        <kwd>ImageCLEF 2019</kwd>
        <kwd>VQA-Med</kwd>
        <kwd>Visual Question Answering</kwd>
        <kwd>Supporting Facts Network</kwd>
        <kwd>Multi-Task Learning</kwd>
        <kwd>Transfer Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In the era of data deluge and powerful computing systems, deriving meaningful
insights from heterogeneous information has shown to have tremendous value
across industries. In particular, the promise of deep learning-based
computational models [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] in accurately predicting diseases has further stirred great
interest in adopting automated learning systems in healthcare [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A
daunting challenge within the realm of healthcare is to e ciently sieve through vast
amounts of multi-modal information and reason over them to arrive at a di
erential diagnosis. Longitudinal patient records including time-series measurements,
text reports and imaging volumes form the basis for doctors to draw conclusive
insights. In practice, radiologists are tasked with reviewing thousands of
imaging studies each day, with an average of about three seconds to mark them as
anomalous or not, leading to severe eye fatigue [24]. Moreover, clinical
workows have a sequential nature tending to cause delays in triage situations, where
the existence of answers to key questions about a patient's holistic conditions
can potentially expedite treatment. Thus, building e ective question-answering
systems for the medical domain by bringing advancements in machine learning
research will be a game changer towards improving patient care.
      </p>
      <p>
        Visual Question Answering (VQA) [
        <xref ref-type="bibr" rid="ref1">17, 1</xref>
        ] is a new exciting problem domain,
where the system is expected to answer questions expressed in natural language
by taking into account the content of the image. In this paper, we present results
of our research on the VQA-Med 2019 dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], an open challenge associated
with the ImageCLEF 2019 initiative [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The main issue here, in comparison
to the other recent VQA datasets such as TextVQA [23] or GQA [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], is dealing
with scattered, noisy and heavily biased data. Hence, the dataset serves as a
great use-case to study challenges encountered in practical clinical scenarios.
      </p>
      <p>
        In order to address the data issues, we designed a new model called
Supporting Facts Network (SFN) that e ciently shares knowledge between upstream
and downstream tasks through the use of a pre-trained multi-task solver in
combination with task-speci c solvers. Note that posing the VQA-Med challenge as
a multi-task learning problem [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] allowed the model to e ectively leverage and
encode relevant domain knowledge. Our multi-task SFN model outperforms the
single task baseline by better adapting to label distribution shifts.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The VQA-Med dataset</title>
      <p>
        The VQA-Med 2019 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is a Visual Question Answering (VQA) dataset embedded
in the medical domain, with a focus on radiology images. It consists of:
{ a training set of 3,200 images with 12,792 Question-Answer (QA) pairs,
{ a validation set of 500 images with 2,000 QA pairs, and
{ a test set of 500 images with 500 questions (answers were released after the
end of the VQA-Med 2019 challenge).
      </p>
      <p>In all splits the samples were divided into four categories, depending on the main
task to be solved:
{ C1: determine the modality of the image,
{ C2: determine the plane of the image,
{ C3: identify the organ/anatomy of interest in the image, and
{ C4: identify the abnormality in the image.</p>
      <p>Our analysis of the dataset (distribution of questions, answers, word
vocabularies, categories and image sizes) led to the following ndings and system-design
related decisions:
{ merge of the original training and validation sets, shu e and re-sample new
training and validation sets with a proportion of 19:1,
{ use of weighted random sampling during batch preparation,
{ addition of a fth Binary category for samples with Y/N type questions,
{ focus on accuracy-related metrics instead of the BLEU score,
{ avoid label (answer classes) uni cation and cleansing,
{ consider C4 as a downstream task and exclude it from the pre-training of
input fusion modules,
{ utilization of image size as an additional input cue to the system.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Supporting Facts Network</title>
      <p>Typical VQA systems process two types of input, visual (image) and language
(question), that need to undergo various transformations to produce the answer.
Fig. 1 presents a general architecture of such systems, indicating four major
modules: two encoders responsible for encoding raw inputs to more useful
representations, followed by a reasoning module that combines them and nally, an
answer decoder that produces the answer.</p>
      <p>Image
Question</p>
      <p>Image
Encoder
Question
Encoder</p>
      <p>Reasoning
Module</p>
      <p>Answer
Decoder</p>
      <p>
        In the early prototypes of VQA systems, reasoning modules were rather
simple and relied mainly on multi-modal fusion mechanisms. These fusion
techniques varied from concatenation of image and question representations, to more
complex pooling mechanisms such as Multi-modal Compact Bilinear pooling
(MCB) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and Multi-modal Low-rank Bilinear pooling (MLB) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Further,
diverse attention mechanisms such as question-driven attention over image
features [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] were also used. More recently, researchers have focused on complex
multi-step reasoning mechanisms such as Relational Networks [
        <xref ref-type="bibr" rid="ref5">21, 5</xref>
        ] and
Memory, Attention and Composition (MAC) networks [
        <xref ref-type="bibr" rid="ref8">8, 18</xref>
        ]. Despite that, certain
empirical studies indicate early fusion of language and vision signals signi cantly
boosts the overall performance of VQA systems [16]. Therefore, we explored the
nding of an "optimal" module for early fusion of multi-modal inputs.
One of our ndings from analyzing the dataset was to use the image size as
additional input cue to the system. This insight triggered an extensive architecture
search that included, among others, comparison and training of models with:
{ di erent methods for question encoding, from 1-hot encoding with
Bag-ofWords to di erent word embeddings combined with various types of
recurrent neural networks,
{ di erent image encoders, from simple networks containing few convolutional
layers trained from scratch to ne-tuning of selected state-of-the-art models
pre-trained on ImageNet,
{ various data fusion techniques as mentioned in the previous section.
Question
      </p>
      <p>Image</p>
      <p>
        The nal architecture of our model is presented in Fig. 2a. We used GloVe
word embeddings [20] followed by Long Short-Term Memory (LSTM) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The
LSTM outputs along with feature maps extracted from images using
VGG16 [22] were passed to the Fusion I module, implementing question-driven
attention over image features [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Next, the output of that module was
concatenated in the Fusion II module with image size representation created by
passing image width and height through a fully connected (FC) layer.
      </p>
      <p>Note that the green colored modules were initially pre-trained on external
datasets (ImageNet and 6B tokens from Wikipedia 2014 and Gigaword 5 datasets
for VGG-16 and GloVe models respectively) and later ne-tuned during training
on the VQA-Med dataset.
3.2</p>
      <p>Architectures of the Reasoning Modules
During the architecture search of the Input Fusion module we used the model
presented in Fig. 3, with a simple classi er with two FC layers. These models were
trained and validated on C1, C2 and C3 categories separately, while excluding
C4. In fact, to test our hypothesis we trained some early prototypes only on
samples from C4 and the models failed to converge.</p>
      <p>After establishing the Input Fusion module we trained it on samples from
C1, C2 and C3 categories. This served as a starting point for training more
complex reasoning modules. At rst, we worked on a model that exploited
information about 5 categories of questions by employing 5 separate classi ers
which used data produced by the Input Fusion module. Each of these
classi ers essentially specialized in one question category and had its own answer
Question</p>
      <p>Image
Image
size</p>
      <p>Image
label dictionary and associated loss function. The predictions were then fed to
the Answer Fusion module, which selected answers from the right classi er
based on the question category predicted by the Question Categorizer
module, whose architecture is shown in Fig. 2b. Please note that we pre-trained the
module in advance on all samples from all categories and froze its weights during
the training of classi ers.</p>
      <p>Support C1
Support C2
Support C3
Fusion III</p>
      <p>Classifier C1
Classifier C2
Classifier C3
Classifier C4
Classifier Y/N</p>
      <p>Answer
Fusion</p>
      <p>The architecture of our nal model, Supporting Facts Network is
presented in Fig. 4. The main idea here resulted from the analysis of questions about
the presence of abnormalities { to answer which the system required knowledge
on image modality and/or organ type. Therefore, we divided the classi cation
modules into two networks: Support networks (consisting of two FC layers) and
nal classi ers (being single FC layers). We added Plane (C2) as an additional
supporting fact. The supporting facts were then concatenated with output from
Input Fusion module in Fusion III and passed as input to the classi er
specialized on C4 questions. In addition, since Binary Y/N questions were present
in both C1 and C4 categories, we followed a similar approach for that classi er.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Experimental Results</title>
      <p>
        All experiments were conducted using PyTorchPipe [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], a framework that
facilitates development of multi-modal pipelines built on top of PyTorch [19]. Our
models were trained using relatively large batches (256), dropout (0:5) and Adam
optimizer [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] with a small learning rate (1e 4). For each experimental run, we
generated a new training and validation set by combining the original sets,
shufing and sampling them in proportions of 19 : 1, thereby resulting in a validation
set of size 5%.
      </p>
      <sec id="sec-4-1">
        <title>Resampled Valid. Set</title>
      </sec>
      <sec id="sec-4-2">
        <title>Original Train. Set</title>
      </sec>
      <sec id="sec-4-3">
        <title>Original Valid. Set Model IF-1C SFN</title>
        <p>Prec.</p>
        <p>In Tab. 1 we present a comparison of average scores achieved by our baseline
models using single classi er (IF-1C) and the Supporting Facts Networks (SFN).
Our results clearly indicate the advantage of using 'supporting facts' over the
baseline model with a single classi er. The SFN model by our team achieved a
best score of (0:558 Accuracy, 0:582 BLEU score) on the test set as indicated
by the CrowdAI leaderboard. One of the reasons for such a signi cant drop in
performance is due to the presence of new answers classes in the test set that
were not present both in the original training and validation sets.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Summary</title>
      <p>In this work, we introduced a new model called Supporting Facts Network
(SFN), that leverages knowledge learned from combinations of upstream tasks
in order to bene t additional downstream tasks. The model incorporates domain
knowledge that we gathered from a thorough analysis of the dataset, resulting in
specialized input fusion methods and ve separate, category-speci c classi ers.
It comprises of two pre-trained shared modules followed by a reasoning module
jointly trained with ve classi ers using the multi-task learning approach. Our
models were found to train faster and to deal much better with label distribution
shifts under a small imbalanced data regime.</p>
      <p>Among the ve categories of samples present in the VQA-Med dataset, C4
and Binary turned out to be extremely di cult to learn, for several reasons.
First, there were 1483 unique answer classes assigned to 3082 training samples
related to C4. Second, both C4 and Binary required more complex reasoning
and, besides, might be impossible to conclude by looking only at the question
and content of the image. However, our observation that some of the information
from simpler categories might be useful during reasoning on more complex ones,
we re ned the model by adding supporting networks. Given, modality, imaging
plane and organ typically help narrow down the scope of disease conditions
and/or answer whether or not an abnormality is present. Our empirical studies
prove that this approach performs signi cantly better, leading to an 18 point
improvement in F-1 score over the baseline model on the original validation set.
16. Malinowski, M., Doersch, C.: The Visual QA devil in the details: The impact of
early fusion and batch norm on CLEVR. In: ECCV'18 Workshop on Shortcomings
in Vision and Language (2018)
17. Malinowski, M., Fritz, M.: A multi-world approach to question answering about
real-world scenes based on uncertain input. In: Advances in neural information
processing systems. pp. 1682{1690 (2014)
18. Marois, V., Jayram, T., Albouy, V., Kornuta, T., Bouhadjar, Y., Ozcan, A.S.: On
transfer learning using a MAC model variant. In: NeurIPS'18 Visually-Grounded
Interaction and Language (ViGIL) Workshop (2018)
19. Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z.,</p>
      <p>Desmaison, A., Antiga, L., Lerer, A.: Automatic di erentiation in pytorch (2017)
20. Pennington, J., Socher, R., Manning, C.: Glove: Global vectors for word
representation. In: Proceedings of the 2014 conference on empirical methods in natural
language processing (EMNLP). pp. 1532{1543 (2014)
21. Santoro, A., Raposo, D., Barrett, D.G., Malinowski, M., Pascanu, R., Battaglia, P.,
Lillicrap, T.: A simple neural network module for relational reasoning. In: Advances
in Neural Information Processing Systems. pp. 4967{4976 (2017)
22. Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale
image recognition. arXiv preprint arXiv:1409.1556 (2014)
23. Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D.,
Rohrbach, M.: Towards vqa models that can read. arXiv preprint arXiv:1904.08920
(2019)
24. Syeda-Mahmood, T., Walach, E., Beymer, D., Gilboa-Solomon, F., Moradi, M.,
Kisilev, P., Kakrania, D., Compas, C., Wang, H., Negahdar, R., et al.: Medical
sieve: a cognitive assistant for radiologists and cardiologists. In: Medical Imaging
2016: Computer-Aided Diagnosis. vol. 9785, p. 97850A. International Society for
Optics and Photonics (2016)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Antol</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            , J., Mitchell,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Lawrence Zitnick,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          : VQA:
          <article-title>Visual Question Answering</article-title>
          .
          <source>In: Proceedings of the IEEE international conference on computer vision</source>
          . pp.
          <volume>2425</volume>
          {
          <issue>2433</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ardila</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiraly</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bharadwaj</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reicher</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tse</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etemadi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ye</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naidich</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shetty</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography</article-title>
          .
          <source>Nature Medicine</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Ben</given-names>
            <surname>Abacha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.A.</given-names>
            ,
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.V.</given-names>
            ,
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          , Muller, H.:
          <article-title>VQA-Med: Overview of the medical visual question answering task at imageclef 2019</article-title>
          .
          <source>In: CLEF 2019 Working Notes. CEUR Workshop Proceedings</source>
          , CEURWS.org &lt; http : ==ceur ws:org= &gt;, Lugano,
          <source>Switzerland (September</source>
          <volume>09</volume>
          -12
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Caruana</surname>
          </string-name>
          , R.:
          <article-title>Multitask learning</article-title>
          .
          <source>Machine learning 28(1)</source>
          ,
          <volume>41</volume>
          {
          <fpage>75</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Desta</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kornuta</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Object-based reasoning in VQA</article-title>
          .
          <source>In: 2018 IEEE Winter Conference on Applications of Computer Vision (WACV)</source>
          . pp.
          <year>1814</year>
          {
          <year>1823</year>
          . IEEE (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Fukui</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohrbach</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darrell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohrbach</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Multimodal compact bilinear pooling for visual question answering and visual grounding</article-title>
          .
          <source>In: EMNLP</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9(8)</source>
          ,
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hudson</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Compositional attention networks for machine reasoning</article-title>
          .
          <source>In: CVPR</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hudson</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Gqa: a new dataset for compositional question answering over real-world images</article-title>
          . arXiv preprint arXiv:
          <year>1902</year>
          .
          <volume>09506</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Peteri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cid</surname>
            ,
            <given-names>Y.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liauchuk</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klimuk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarasau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ben</surname>
            <given-names>Abacha</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.A.</given-names>
            ,
            <surname>Datla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Dang-Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.T.</given-names>
            ,
            <surname>Piras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Riegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.T.</given-names>
            ,
            <surname>Lux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Gurrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.M.</given-names>
            ,
            <surname>de Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.G.S.</given-names>
            ,
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Kavallieratou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>del Blanco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.R.</given-names>
            , Rodr guez, C.C.,
            <surname>Vasillopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Karampidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Chamberlain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Campello</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>ImageCLEF 2019: Multimedia retrieval in medicine, lifelogging, security and nature</article-title>
          . In:
          <article-title>Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the 10th International Conference of the CLEF Association (CLEF</source>
          <year>2019</year>
          ),
          <source>LNCS Lecture Notes in Computer Science</source>
          , Springer, Lugano,
          <source>Switzerland (September 9-12</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kazemi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elqursh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Show, ask, attend, and answer: A strong baseline for visual question answering</article-title>
          .
          <source>arXiv preprint arXiv:1704.03162</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>On</surname>
            ,
            <given-names>K.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ha</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          , Zhang, B.T.:
          <article-title>Hadamard product for low-rank bilinear pooling</article-title>
          .
          <source>In: ICLR</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
          </string-name>
          , J.:
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>arXiv preprint arXiv:1412.6980</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kornuta</surname>
          </string-name>
          , T.: PyTorchPipe. https://github.com/ibm/pytorchpipe (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>LeCun</surname>
          </string-name>
          , Y.,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.:
          <article-title>Deep learning</article-title>
          .
          <source>nature</source>
          <volume>521</volume>
          (
          <issue>7553</issue>
          ),
          <volume>436</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>