<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>BROWALLIA at Memotion 2.0 2022 : Multimodal Memotion Analysis with Modified OGB Strategies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Baishan Duan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuesheng Zhu</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Emotion analysis with social media is important for many social psychology tasks such as hatespeech detection. Internet memes comprises of various modalities including textual, visual and audio, multimodal information computational processing needs superior methods. Recently, Memotion 2.0 attracted widespread attention. In this paper, we propose a novel multimodal system to analyze textual- visual pair information. We choose LSTM and ResNet50 as backbone, then employee late-fusion to fusion two modalities. In addition, to train network well, we adopt modified ofline-gradient-blending strategy to alleviate overfitting. Our approach achieves good performance in three sub-tasks, ranking 2 in Subtask A, 3 in Subtask B, 2 in Subtask C.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Internet media has taken a large part in human life, and internet media platforms such as
facebook, twitter, and weibo contain a large amount of information, which mostly consists of
images and text. People like to express their emotions on social media, which may be humorous
or contain sarcasm, may be ofensive, harmful to individuals, governments, organizations, and
races, or motivational. Sentiment analysis of social media can be a good solution to these
problems.</p>
      <p>
        Common text sentiment classification has been relatively improved, such as RNN, LSTM[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
BERT[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] et al. Multimodal sentiment classification faces a big challenge. The MEmotion task
collects 10K annotated memes online, including images and text information obtained from
image OCR. This data format will make the sentiment information mining more dificult, the text
in the images will generate noise to the image content, and the OCR results are not guaranteed to
be exactly the same as in image. In addition, the relationship between some picture contents and
textual expressions is not very close, which brings trouble to the relationship mining between
modalities. In addition to the semantic information of images, imbalanced dataset also brings
challenges to the final results.
      </p>
      <p>
        There are three subtasks in Memotion 2.0[
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. Subtask A is overall sentiment (positive,
negative, and neutral) analysis of memes, Subtask B is binary emotion classification from humor,
sarcasm, ofensive, and motivational. Subtask B is extended into Subtask C, which is
Finegrained sentiment classification for four categories. We consider Subtask A as the pretraining
task and Subtask B and C as downsteam tasks.
      </p>
      <p>
        In this paper, we propose a novel multimodal system to analyze textual- visual pair
information. We use ResNet-50[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and LSTM to be backbone and late fusion features extracted
from backbone. We additionally adopt modified ofline-gradient-blending(OGB)[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to alleviate
overfitting. Our contributions are as follows:
• We propose a multimodal information fusion network, which can extract text and image
features well.
• We use modified OGB strategy to train multimodal networks, which alleviate overfitting.
• Experiments show that our model achieves good results in all three subtasks.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>
        Analysis of Memotion 2.0 is a part of sentiment analysis.In the field of text-only sentiment
classification, there are many excellent methods, such as RNN, LSTM, Transformer[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and BERT.
Attention mechanisms are proposed to pay better attention to the contextual information of the
text, and BERT with variants like RoBerta[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and XLNet[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] are the state-of-the-art methods in
text-only tasks . In the field of image classification, ResNet is the most used backbone, which is
pretrained on ImageNet[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. However, there is still much to be mined in multimodal sentiment
classification. TFN[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] uses inner product to fuse the features of three modalities. LMF utilizes
a low-rank matrix decomposition of the weights, where each mode is first individually linearly
transformed before multi-dimensional dot product, which can be viewed as a sum of the results
of multiple low-rank vectors, thus reducing the number of parameters in the model. MulT[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
attends to interactions between multimodal sequences across distinct time steps and latently
adapt streams from one modality to another.
      </p>
      <p>In our work, considering the small dataset, we prefer LSTM to transformer-based models
as textual backbone. ResNet is used to extract image features. Our architecture is light but
efective.</p>
      <p>We proposed a multimodal network, based on pretrained ResNet-50 and LSTM. We regard
TaskA as pretraining task and then finetune it to complete TaskB and TaskC. We also refer to
uni-model architecture and compare the result to multi-model result. As a result, we found that
multi-model can extract features well.</p>
      <p>We experiment with text-only and image-only models for Memotion classification. The result
analysis is justified in subsection 4.2.1
To extract image feature, we opted for the ResNet-50 architecture, which is the most common
visual backbone. The residual block can avoid network degradation and alleviate gradient
vanishing problems. In the experiment, we use pretrained ResNet-50 from timm1.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <sec id="sec-3-1">
        <title>3.1. Overview</title>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Uni-model</title>
        <p>3.2.1. Visual
3.2.2. Textual
To extract textual features, we refer to two basic language modeling network LSTM and BERT.
LSTM is the variant of RNN, it is mainly to solve the gradient disappearance and gradient
explosion during long sequence training. Bert is the variant of Transformer proposed by Google
research. Through training on huge corpora, Bert achieved state of the art result on many
NLP tasks. For LSTM, we realize it by pytorch, for Bert, we adopted the pretrained version by
HuggingFace2.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Multi-model</title>
        <p>In order to better mining relationship between image and text, we combined two single feature
extraction components into a unified network. For Multi-model design, we just add late fusion
layer to fusion features extracted from visual and textual backbone. Note that we remove
classification layer in ResNet-50 to obtain only convolutions used in extracting the features
from the input image . In addition, we adopted modified ofline-gradient-blending strategy to
alleviate overfitting. For TaskB and TaskC we just add four FC layer based on pretrained TaskA
multi-modal network.
3.4. OGB
We add ofline-gradient-blending to alleviate overfitting. There are three branches in our
multimodal network, and overfitting of each branch will happen at diferent time as well as the
1https://github.com/rwightman/pytorch-image-models
2https://github.com/huggingface/transformers
overfiting ratio is also diferent. Overfiting-to-Generalization Ratio is proposed to measure
overfitting as follows
 = |
Δ  ,
Δ  ,
| = |
  ,</p>
        <p>−  
 ∗ −  ∗+
|
Where  ∗ indicates the validation loss of N epoch,  indicates the gap between validation loss
and train loss.</p>
        <p>Then compute the weight of each branch loss:</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.5. Objective Function</title>
        <p>We sum three branch losses with the weight computed by OGB:
{  ∗}=+11 =
1  
   2</p>
        <p>+1
= ∑    
=1
(1)
(2)
(3)</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiment</title>
      <sec id="sec-4-1">
        <title>4.1. Experiment setup</title>
        <sec id="sec-4-1-1">
          <title>4.1.1. Dataset and preprocessing</title>
          <p>Our experiments are conducted on the oficial dataset of Memotion 2.0. The training dataset
consists of 7k human annotated internet memes. Each sample is composed of image and text
extracted from image by OCR as shown in Fig 1. Besides, the class distribution is imbalanced
shown in Table 1, thus we adopt focal loss in our experiments.</p>
          <p>For image we use transformers such as random flip and random crop, finally the input
image is resized to 320x320. For text, we padding it to the fixed length 30, and tokenize it by
BertTokenizer additionally in Bert Model.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Implementation</title>
          <p>We conducted our model by pytorch 1.7.1, we use ResNet-50 from timm and Bert-base-uncased
from Huggingface. We set the hyper-parameters manually, the batch-size, epoch, learning rate,
weight decay are 32, 40, 1e-5, 5e-3. The Optimizer is Adam with linear lr decay. We train all
models on a single 2080Ti GPU on Ubuntu system.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Result</title>
        <sec id="sec-4-2-1">
          <title>4.2.1. Uni-model and Multi-model</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>4.2.2. Modified Online Gradient Blending</title>
          <p>Table 3 contains the result of ogb strategy training. At first, we compute the weight by Eq . 2,
the weight for fusion branch, visual branch and textual branch weight are 0.22, 0.17, 0.61. The
result shows worse performance than 1: 1: 1. We visualize the loss as shown in Fig 3.</p>
          <p>According to Eq .2. G represents the degree to which the validation loss decreases. The
greater the G, the greater the weight, but in this data set, the validation loss has been increasing.
Because of the long-tailed problem, the accuracy has been decreasing, so this formula is not
suitable. What we need to do is to suppress the growth of validation loss and decrease the gap
between validation loss and train loss. After the analysis, We tried to modify the formula to
{  ∗}=+11 =
1 1
   2 
(4)
which means that the weight is inversely proportional to the degree of increase in the validation
loss, and the calculated weights are 0.5, 0.3, 0.2, and have relatively good results. Additionally,
for the imbalanced dataset, we use focal loss and get the best F1 score in validation set.</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>4.2.3. Final Result</title>
          <p>Result is shown in Table 4 , ⋆ indicates our model. Our method gets decent performance
compared with baseline and ranking2 in Subtask A, 3 in Subtask B, 2 in Subtask C.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this paper, we proposed a multimodal architecture with modified OGB strategy for Memotion
2.0 tasks. Experiments show that our methods outperform strong baseline in three subtasks. In
addition our score ranks high in the test sets.</p>
      <p>
        For future work, we will try more powerful backbones like ResNext[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], DenseNet[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] in
visual and RoBERTa, XLNet in textual. Strong backbone allows for better extraction of modal
features. For multimodal architecture, we will try joint end-to-end model such as VLBert[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ],
VisualBert[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. At the same time, we find that imbalenced dataset cause low score in subtask
B and subtask C, maybe some strategy to solve long-tailed problem make sence such as loss
function[
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ] and new architecture like BBN[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <article-title>Long short-term memory</article-title>
          ,
          <source>Neural computation 9</source>
          (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramamoorthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gunti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Ahuja</surname>
          </string-name>
          ,
          <article-title>Memotion 2: Dataset on sentiment and emotion analysis of memes</article-title>
          , in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and
          <article-title>Hate Speech Detection</article-title>
          ,
          <string-name>
            <surname>CEUR</surname>
          </string-name>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramamoorthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gunti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Ahuja</surname>
          </string-name>
          ,
          <article-title>Findings of memotion 2: Sentiment and emotion analysis of memes</article-title>
          , in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and
          <article-title>Hate Speech Detection</article-title>
          ,
          <string-name>
            <surname>CEUR</surname>
          </string-name>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Feiszli</surname>
          </string-name>
          ,
          <article-title>What makes training multi-modal classification networks hard?</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>12695</fpage>
          -
          <lpage>12705</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>arXiv</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach (</article-title>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          , J. Carbonell,
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <article-title>Xlnet: Generalized autoregressive pretraining for language understanding</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Imagenet: A large-scale hierarchical image database</article-title>
          ,
          <source>Proc of IEEE Computer Vision Pattern Recognition</source>
          (
          <year>2009</year>
          )
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Poria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Cambria</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Morency</surname>
          </string-name>
          ,
          <article-title>Tensor fusion network for multimodal sentiment analysis</article-title>
          ,
          <source>arXiv</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Y.-H. H. Tsai</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>P. P.</given-names>
          </string-name>
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>J. Z.</given-names>
          </string-name>
          <string-name>
            <surname>Kolter</surname>
            ,
            <given-names>L.-P.</given-names>
          </string-name>
          <string-name>
            <surname>Morency</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <article-title>Multimodal transformer for unaligned multimodal language sequences</article-title>
          ,
          <source>in: Proceedings of the conference. Association for Computational Linguistics. Meeting</source>
          , volume
          <year>2019</year>
          , NIH Public Access,
          <year>2019</year>
          , p.
          <fpage>6558</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollár</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>Aggregated residual transformations for deep neural networks</article-title>
          ,
          <source>CoRR abs/1611</source>
          .05431 (
          <year>2016</year>
          ). URL: http://arxiv.org/abs/1611.05431. arXiv:
          <volume>1611</volume>
          .
          <fpage>05431</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          ,
          <article-title>Densely connected convolutional networks</article-title>
          ,
          <source>CoRR abs/1608</source>
          .06993 (
          <year>2016</year>
          ). URL: http://arxiv.org/abs/1608.06993. arXiv:
          <volume>1608</volume>
          .
          <fpage>06993</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>W.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Dai,</surname>
          </string-name>
          <article-title>VL-BERT: pre-training of generic visuallinguistic representations</article-title>
          , CoRR abs/
          <year>1908</year>
          .08530 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1908</year>
          . 08530. arXiv:
          <year>1908</year>
          .08530.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yatskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-J. Hsieh</surname>
            ,
            <given-names>K.-W.</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
          </string-name>
          ,
          <article-title>Visualbert: A simple and performant baseline for vision and language</article-title>
          , arXiv preprint arXiv:
          <year>1908</year>
          .
          <volume>03557</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jia</surname>
          </string-name>
          , T.-
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Belongie</surname>
          </string-name>
          ,
          <article-title>Class-balanced loss based on efective number of samples</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>9268</fpage>
          -
          <lpage>9277</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gaidon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Arechiga</surname>
          </string-name>
          , T. Ma,
          <article-title>Learning imbalanced datasets with labeldistribution-aware margin loss</article-title>
          , arXiv preprint arXiv:
          <year>1906</year>
          .
          <volume>07413</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.-S.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.-M.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Bbn:
          <article-title>Bilateral-branch network with cumulative learning for long-tailed visual recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>9719</fpage>
          -
          <lpage>9728</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>