<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bert with Dynamic Masked Softmax and Pseudo Labeling for Hierarchical Product Classi cation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Li Yang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shijia E</string-name>
          <email>e.shijia@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shiyao Xu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yang Xiang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Tencent</institution>
          ,
          <addr-line>Shanghai</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Tongji University</institution>
          ,
          <addr-line>Shanghai</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>ierarchical product classi cation (HPC) aims to assign prede ned product categories stored in a hierarchical structure to product instances. Categories at di erent levels of a product tend to have dependencies. However, most previous studies either decompose the original problem into a set of at classi cation sub-problems or deal with all categories simultaneously, lacking a focus on the dependencies among di erent category levels. In this paper, we propose a BERT-based ensemble model to address the HPC challenge in SWC2020MWPD Task 21. We devise a masked matrix for each category level based on the hierarchical category structure, which can dynamically lter out the child categories unrelated to the current parent category and eliminate the negative e ect of category inconsistency. Further through a two-level ensemble strategy and pseudo labeling, our team Rhinobird wins rst place with a weighted-average macro-F1 score of 88.62 on the testing dataset. Our source code is publicly available on Github.2</p>
      </abstract>
      <kwd-group>
        <kwd>Hierarchical Product Classi cation BERT Dynamic Masked Softmax Pseudo Labeling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Recent years have seen signi cant use of semantic annotations in the e-commerce
domain, where online shops (e-shops) are increasingly adopting semantic markup
languages to describe their products to improve their visibility. Hierarchical
product classi cation is a fundamental but challenging task of product semantic
annotation, where product instances are assigned to multiple levels of categories
which are stored hierarchically. Automated product classi cation based on the
product o er information made on the web has become an essential tool for
searching, retrieving, and managing the products.</p>
      <p>Copyright c 2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
1 https://ir-ischool-uos.github.io/mwpd/index.html#task2
2 https://github.com/AlexYangLi/iswc2020_prodcls</p>
      <p>For the hierarchical classi cation problem, prior studies predict only the
categories of the last level by reducing the problem into a at multi-class
problem [2]. Unfortunately, these at-based approaches ignore the hierarchical
category structure information. To this end, some works have considered the
hierarchical structure, which can be categorized into two approaches: 1) Local
approaches [4] generate a unique classi er for each parent node in the category
hierarchy; 2) Global approaches [3,5] deal with all the levels of categories by
using a single classi er. The local approaches su er the inherited disadvantage
that the number of sub-models grows exponentially concerning the number of
category levels. This is especially problematic for neural-based models with a
large number of parameters.</p>
      <p>In this paper, we propose an end-to-end global approach that overcomes the
problem of exploding models and explicitly considers the dependencies among
di erent category levels. We use BERT [1] as the base model to generate a
rich semantic product representation and predict the categories level by level,
conditioned on a dynamic masked matrix obtained based on the hierarchical
category structure. The masked matrix serves as an information lter that
dynamically lters out the child categories unrelated to the current parent
category, which contributes to addressing the category inconsistency problem. To
enhance the generalization ability of our model, we adopt a two-level
ensemble strategy, which combines the results of 17 di erent BERT models to make
the nal decisions. Furthermore, we utilize pseudo labeling, a semi-supervised
method that uses the unlabeled data to enhance the performance. Our solution
achieves a weighted-average macro-F1 score of 88.62 on the testing dataset of
SWC2020MWPD Task2.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Dataset</title>
      <p>For this challenge, the organizers have provided 10012 labeled products for
training, 3008 labeled products for validation, and 3107 unlabeled products for
testing. Each product instance is provided with its id, name, description,
websitespeci c product category, and original web page URL. In this paper, we use
only the name and description text to solve the product classi cation problem.
The left sub- gure of Figure 1 shows the distribution of text length of product
names and descriptions, and we can nd that most of them are less than 250.
The product labels are organized in three levels, of which the rst level has
37 categories, the second level has 76 categories and the third level has 281
categories. We show the distribution of categories labels of three levels in the
right sub- gure in Figure 13. We can nd that the category distribution is fairly
imbalanced, which increases the di culty of the task.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Model Description</title>
      <p>In this section, we describe the details of our proposed approach. The overall
framework and processing pipeline are shown in Figure 2, including the base
model construction, model ensembling, and pseudo labeling for data
augmentation.</p>
      <p>Test Data</p>
      <sec id="sec-3-1">
        <title>BERT1</title>
        <p>C. pseudo labeling
Ensembel Model
Labeled Data
B. model voting ensemble</p>
      </sec>
      <sec id="sec-3-2">
        <title>BERT2 BERT3</title>
      </sec>
      <sec id="sec-3-3">
        <title>BERT17</title>
        <p>A. 5-fold single model averaging ensemble
BERT17-1 BERT17-2 BERT17-3 BERT17-4 BERT17-5</p>
      </sec>
      <sec id="sec-3-4">
        <title>BERT2-1 BERT2-2 BERT2-3 BERT2-4 BERT2-5</title>
      </sec>
      <sec id="sec-3-5">
        <title>BERT1-1 BERT1-2 BERT1-3 BERT1-4 BERT1-5 fold1 fold2 fold3</title>
        <p>Train Data
fold4</p>
        <p>fold5
D. data augment
3 We only show the categories with top-10 frequencies for clarity, the other label
accounts for the rest categories.</p>
        <sec id="sec-3-5-1">
          <title>SLoeftvmela1x SMLoeafvtsmeklea2dx SMLoeafvtsmkeleadx3</title>
          <p>P_O
H
H</p>
          <p>H
BERT</p>
          <p>H</p>
          <p>H
name desc
BERT-PO</p>
        </sec>
        <sec id="sec-3-5-2">
          <title>SLoefvtmela1x SMLoeafvtsmeklea2dx SMLoeafvtsmeklea3dx</title>
          <p>H0 H0 H0</p>
        </sec>
        <sec id="sec-3-5-3">
          <title>SLoefvtmela1x SMLoeafvtsmeklea2dx SMLoeafvtsmeklea3dx</title>
          <p>H0 H0 H0 P_O
H
H</p>
          <p>H
BERT</p>
          <p>H
H
H
H</p>
          <p>H
BERT</p>
          <p>H
H</p>
        </sec>
        <sec id="sec-3-5-4">
          <title>SLoefvtmela1x SMLoeafvtsmeklea2dx SMLoeafvtsmeklead3x</title>
          <p>O AVG MAX P_O
BiLSTM / BiGRU / CNN
C T1 … TN Tsep T’1 … T’M</p>
          <p>H
BERT</p>
          <p>H
name desc</p>
          <p>BERT-seq
name desc
BERT-K-hidden</p>
          <p>name desc</p>
          <p>BERT-K-hidden-PO
jointly conditioning on both left and right context in all layers, which achieved
state-of-the-art performance in several language understanding tasks.</p>
          <p>In practice, we ne-tune the BERT model to generate a contextual product
representation and then add three output layers on top of this representation to
jointly classify the products into the categories of three levels. For the input to
BERT, we use the name and description text of products provided by the data.
Instead of concatenating the name and description into one sentence, we regard
them as one text pair and use a special symbol, "[sep]", to separate them before
feeding into the BERT model. The intuition behind this is that the product
name provides more concise and useful information about the product. We want
to enhance the in uence of the product name and expect that the model can
distinguish them.</p>
          <p>
            To generate rich semantic product representations, We make full use of the
hidden states from the last or more hidden layers of BERT, resulting in 17
di erent BERT base models. Figure 3 presents four di erent ways:
(
            <xref ref-type="bibr" rid="ref1">1</xref>
            ) BERT-PO uses the pooler output of BERT as the product representation,
which is the most common way to adopt BERT for classi cation problems.
(
            <xref ref-type="bibr" rid="ref2">2</xref>
            ) BERT-K-hidden concatenates the rst hidden state from the last K
hidden layers of BERT as the product representation. We range K from 1 to 5,
resulting in 5 di erent models.
(
            <xref ref-type="bibr" rid="ref3">3</xref>
            ) BERT-K-hidden-PO concatenates the rst hidden state from the last K
hidden layers as well as the pooler output of BERT as the product
representation. We range K from 1 to 5, resulting in 5 di erent models.
(
            <xref ref-type="bibr" rid="ref4">4</xref>
            ) BERT-seq uses the hidden states from the last hidden layer of BERT as
the input of another sequence layer, and then concatenates the pooler
output of BERT, with the last hidden output as well as the max-pooling and
mean-pooling over the hidden states of sequence layer, as the nal product
representation. We try 5 di erent sequence layers: BiLSTM, BiGRU, CNN,
BiLSTM+CNN, and BiGRU+CNN.
          </p>
          <p>Dynamic Masked Softmax
Suppose the product representation generated by the BERT base model is
denoted by R. We then feed it into three feed-forward neural layers, to compute
the scores of categories for each level:</p>
          <p>Ol = W lR + bl;
where W l and b1 are the weight matrix and bias term for the category level l.
Hereafter, we can directly adopt a plain softmax layer to normalize the scores
and select the category with the maximum probability. However, this approach
chooses category independently for each level, ignoring the dependencies among
di erent category levels, which can cause the problem of category inconsistency.</p>
          <p>In this paper, we propose a category hierarchy based mask matrix to address
the above problem by dynamically ltering out the child categories which are
unrelated to the current parent category. Speci cally, we take advantage of the
pre-de ned hierarchical category structure information and devise a mask matrix
for each sub-level M l 2 f0; 1gNl 1 Nl , where N l is the total number of categories
of the current level l and N l 1 is the amount of categories of the last level. In
this matrix, each M ul;v = 1 indicates that the v-th category of level l is the child
category of the u-th category of level l 1, otherwise M ul;v = 0. We then compute
the normalized scores of categories of level l using the masked softmax function
with M l as follows:</p>
          <p>l
P yv j s;
=
exp Ovl</p>
          <p>M ul;v + exp( 8)
PN
v0=1 exp Ovl0</p>
          <p>M ul;v0 + exp( 8)
;
where P yvl j s; is the probability of assigning the v-th categories of level l to
the product with sentence s by the model with parameter , and u is the index of
the parent category. Notably, the probability of the categories unrelated to the
current parent category will be fairly small, and therefore, their negative e ect
can be eliminated.</p>
          <p>With this design, we successfully introduce the hierarchical category
structure information to the model. Besides, by this ltering mechanism, we can
reduce the number of categories to be classi ed for each sub-levels and classify
dynamically according to the predicted parent category, which can ensure the
consistency of categories between di erent levels.
3.3</p>
          <p>
            Model Ensemble
To combine di erent single models, we adopt a two-level ensemble strategy. In
the rst level, we rst combine the original training and validation dataset and
split them into ve folds by using the cross-validation technique. For every
single model that we design above, we train it for ve times by sequential choosing
one fold for validation and the rest four folds for training. We then average the
probability outputs from these ve single models with the same model
architecture but trained on a di erent dataset. In the second level, we apply the voting
(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )
(
            <xref ref-type="bibr" rid="ref2">2</xref>
            )
ensemble strategy to the 17 averaged ensemble single models, by choosing the
most voted category as the nal prediction. With the two ensemble strategy, we
can generate a robust model with better generalization ability than any single
model.
3.4
          </p>
          <p>Pseudo Labeling
To further improve the performance, we utilize pseudo labeling, a simple yet
e ective semi-supervised method that allows us to make full use of the unlabeled
data. To be speci c, after training the single models with the labeled data and
ensembling them, we use the ensemble model to predict labels for the unlabeled
testing data. Together with the pseudo-labeled data and training data, we then
retrained the single models again and thus obtain a new ensemble model. Pseudo
labeling can be regarded as an e ective way of data augmentation, which can
alleviate the over- tting problem by increasing the amount of training data.
Since the ensemble model performs better than any of the single models in most
cases, the pseudo-labels predicted by the ensemble model can correct the bias
made by the single models, thus leading to performance improvements.
4
4.1</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <p>Experimental Setup
Evaluation Metric For the challenge, a weighted-average macro-F1 score
(WAF1) will be calculated over all the categories for each classi cation level.
Then the average of the WAF1 of the three levels will be calculated and used as
the evaluation metric to rank the participating systems.</p>
      <p>Parameter Settings Our proposed approach is implemented by Tensor ow.
We use the BERT-Base-Uncased version of Google's pre-trained BERT models4
for ne-tuning, which has 12 transformer layers, 12 attention heads, and 768
hidden sizes, and is trained on lower-cased English text. We use Adam as the
optimizer and set the initial learning rate to be 2e-5. The batch size is 32. All
the hidden states of BiLSTMs and BiGRU and feature maps of CNNs used in
the sequence layer have the same dimension as BERT's. For training every single
model, early-stopping is applied to avoid over- tting: training will be stopped
when no performance improvement is observed on the validation dataset after
three epochs. We then store the single model with the best performance on the
validation dataset.
4.2</p>
      <p>Experimental Results of single models
We rst present the experimental results of the single models in Table 1. Due
to the limited space, we only show the performance of the best single model,
4 https://storage.googleapis.com/bert_models/2020_02_20/uncased_L-12_
H-768_A-12.zip</p>
      <p>Model</p>
      <p>BERT-5-hidden
BERT-5-hiddenname
BERT-5-hiddendesc</p>
      <p>BERT-5-hiddensent
BERT-5-hiddenw=o mask</p>
      <p>BERT-5-hidden, which concatenates the rst hidden state of the last ve
hidden layers as the product representation. As indicated from the table, the best
performance among 17 BERT models can achieve a weighted-average macro-F1
score of 0.8730. To gain a better understanding of the impact of the BERT's
input and dynamic masked softmax, we further perform an ablation study.</p>
      <p>The di erence between the 1st to 4th single models is the input they are
trained with, where BERT-5-hiddenname and BERT-5-hiddendesc are the
variants which only use the product's name or description as input, while
BERT-5hiddensent uses both name and description as input but concatenates them as
one sentence. As shown in Table 1, BERT-5-hiddendesc is inferior to
BERT-5hiddenname. It is because the description of a product contains more noise (e.g.,
misspelled words, useless URLs, etc.), while the name can convey more
concise and accurate information about a product. Compare to BERT-5-hiddenname
and BERT-5-hiddendesc, BERT-5-hiddensent achieves better performance, which
demonstrates using both name and description as input can provide more useful
and complementary semantic information for product classi cation.
BERT-5hidden distinguishes the name and description by regarding them as one text
pair when adopting BERT to model them, thus performs better than
BERT-5hiddensent.</p>
      <p>Furthermore, we observe performance degradation when eliminating the use
of dynamic masked softmax (BERT-5-hiddenw=o mask), highlighting the
importance of incorporating hierarchical category structure information for
multilevels classi cation.
4.3</p>
      <p>Experiment results of the ensemble model
Table 2 shows the experimental results of the ensemble models. Compared to
the single models, our ensemble model brings a signi cant performance gain,
which achieves a weighted-average macro-F1 score of 0.8662 and is ranked the
rst among all the participating systems. This illustrates the utility of our
twolevel ensemble strategy. Removing the use of dynamic masked softmax
(BERTensemblew=o mask) degrades the model's performance, which is consistent with the
observation in the single models. We also develop another variant that does not
utilize the pseudo-labeled data to retrain the single models, which we denote
BERT-ensemblew=o pseudo. We can see from Table 2 that the ensemble model
achieves worse performance without pseudo labeling. This strongly veri es that
pseudo labeling plays a great role in enhancing the generalization ability of the
model.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we propose a BERT-based ensemble model for the HPC
challenge of SWC2020MWPD Task2. In our solution, we construct 17 di erent single
BERT models to generate rich semantic product representations for classi
cation and then adopt a two-level ensemble strategy to combine their predictions.
To consider the hierarchical category structure, we devise a masked matrix for
each category level that can dynamically lter out the child categories unrelated
to the parent category. Besides, we apply pseudo labeling to make full use of the
unlabeled test data for better performance. Our team Rhinobird win rst place
with a weighted-average macro-F1 score of 88.62 on the testing dataset.
Acknowledgments This work was supported by the National Key Research
and Development Project of China (2019YFB1704402), and the 2019 Tencent
Marketing Solution Rhino-Bird Focused Research Program.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          : BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics</source>
          . pp.
          <volume>4171</volume>
          {
          <issue>4186</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Fall</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          , Torcsvari,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Benzineb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Karetka</surname>
          </string-name>
          , G.:
          <article-title>Automated categorization in the international patent classi cation</article-title>
          .
          <source>SIGIR Forum</source>
          <volume>37</volume>
          (
          <issue>1</issue>
          ),
          <volume>1025</volume>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Hierarchical multi-label text classi cation: An attention-based recurrent network approach</article-title>
          .
          <source>In: Proceedings of the 28th ACM CIKM International Conference on Information and Knowledge Management</source>
          . pp.
          <volume>1051</volume>
          {
          <issue>1060</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kowsari</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brown</surname>
          </string-name>
          , D.E.,
          <string-name>
            <surname>Heidarysafa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Jafari</given-names>
            <surname>Meimandi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Gerber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Barnes</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.E.</surname>
          </string-name>
          : Hdltex:
          <article-title>Hierarchical deep learning for text classi cation</article-title>
          .
          <source>In: 2017 16th IEEE International Conference on Machine Learning and Applications (ICMLA)</source>
          . pp.
          <volume>364</volume>
          {
          <issue>371</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Sinha</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cheung</surname>
            ,
            <given-names>J.C.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruths</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A hierarchical neural attentionbased text classi er</article-title>
          .
          <source>In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <volume>817</volume>
          {
          <issue>823</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>