<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>H-Bert: Enhancing Chinese Pretrained Models with Attention to HowNet</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Wei Zhu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DataSelect AI Technology</institution>
          ,
          <addr-line>Shanghai</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>East China Normal University</institution>
          ,
          <addr-line>Shanghai</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Pretrained transformers for Chinese show remarkable performances on various natural language processing tasks. However, these models are purely data-driven and fail to incorporate explicit semantic knowledge, like HowNet. In this paper, we propose H-BERT, which enhances the semantic representations of Chinese BERT by incorporating sememe knowledge from HowNet in the pretraining stage via multi-head attention. Our experiments demonstrate that HBERT can significantly outperform the vanilla BERT on the downstream tasks. Ablation study compares different settings of H-BERT and shows that and case study also shows that knowledge injection is required at both the pretraining and fine-tuning stage.3</p>
      </abstract>
      <kwd-group>
        <kwd>pretrained language models knowledge graph knowledge enhanced pretraining</kwd>
        <kwd>Type of submission</kwd>
        <kwd>Poster</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Since the rise of BERT, pretrained language models (PLMs) have dominated state of
the art (SOTA) for a comprehensive list of natural language tasks [
        <xref ref-type="bibr" rid="ref2 ref3 ref7">2, 7, 3</xref>
        ]. Despite their
powerfulness, PLMs still fall short on a series of tasks that requires entity level and
domain level knowledge [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. As a result, a branch of literature has been dedicated to
injecting structured knowledge into PLMs, both in pretraining and fine-tuning stages.
One approach is to inject structure information of knowledge graph via entity embedding
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Similarly, [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] pretrains BERT jointly with knowledge embedding training.
Another approach is to inject knowledge by adding them into the original sentence.
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] explicitly injects related triples extracted from KG into the sentence to obtain an
extended tree-form input for BERT.
      </p>
      <p>
        However, the literature falls short on three aspects. First, many PLMs that inject
knowledge from pretraining suffer from catastrophic forgetting and thus can only perform
well on entity-related tasks such as named entity recognition and relation classification,
but performs poorly on sentence-level tasks like GLUE [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Second, there are few
studies on enriching Chinese PLMs with Knowledge. K-BERT studied incorporating
HowNet without pretraining by adding the knowledge facts into the input sentence.
However, it requires manually select the most important two sememes for each word,
which is unsuitable for scale-up and tasks of different domains.
3 Copyright c 2021 for this paper by its authors. Use permitted under Creative Commons License
      </p>
      <p>Attribution 4.0 International (CC BY 4.0).</p>
      <p>
        In this work, we adopt the core data of HowNet [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] as our knowledge source.
HowNet was initially designed and constructed in the 1990s. Furthermore, it has kept
frequently updating since it was published in 1999. In HowNet, each word has several
sememes, which by linguistic definition, are the minimum semantic units of language
and can well represent implicit semantic meanings behind words. For example, as shown
in Figure 1, the word 中国(China) has a series of sememes, i.e., 国家(country), 中
国(China), 亚洲(Asia), 政治(politics), 地方(place). The sememe set of HowNet is
determined by extracting, analyzing, merging, and filtering semantics of thousands of
Chinese characters. HowNet is widely applied for knowledge-enhanced word/sentence
representations [
        <xref ref-type="bibr" rid="ref14 ref8">8, 14</xref>
        ] and is shown to be beneficial for a wide range of NLP tasks.
However, the previous work does not combine HowNet with language model pretraining.
      </p>
      <p>This work proposes HowNet BERT (H-BERT), a transformer-based model that
uses a simple multi-head attention module to incorporate How-Net knowledge. Our
H-BERT model is depicted in Figure 1. A sentence is encoded via two modules. First,
it is tokenized and embedded via the token embedding layer. Second, we recognize
the words which are included in How-Net and obtain their sememes. Tokens in the
same word will have the same sememes. We can treat the sememes of a word as
subword features. Sememes are treated as the minimum units, and the sememes of the
tokens will be embedded via the sememe embedding layer. Knowledge from How-Net
is injected via a multi-head attention layer from the token representation to the tokens’
sememe representation (denoted as attn-to-sememes). Here, this attention module can
be implemented right after the token embedding layer or after obtaining the sequential
output of the token encoder.</p>
      <p>We conduct experiments on sentence classification (CLS) and sentence pair
classification (NLI), which are whole-sentence level tasks, and named entity recognition (NER),
which is an entity-level task. Tasks from different domains are selected. Experimental
results show that our model consistently outperforms the vanilla ALBERT and K-BERT
on a series of tasks, indicating that our model can handle both whole sentence-level and
entity-level tasks equally well. Moreover, we find that pretraining with attn-to-sememe
but exclude this module during fine-tuning also improves the performance of the vanilla
PLM.</p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <sec id="sec-2-1">
        <title>Sentence Encoder</title>
        <p>
          Now we discuss how to incorporate the HowNet knowledge. First, in a sentence, we
match all the words (not overlapping) included in the sememe via FlashText [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Then
sememes of these words are obtained. The tokens will have the same sememe if a word
has more than one token after sub-word tokenization. Now we have a sememe sequence
S0 = (s1; s2; :::; sTt ), in which si = [sem1; sem2; :::; semli ] means the word si is in
has li sememes according to HowNet. For tokens in words that are not in HowNet, li = 1
and we given it a special padding sememe, denoted as &lt; s pad &gt;. For example, in
Figure 1, different words have different sememes, and some have no sememes. The
tokens’ sememes are embedded to tensor SE0 = (se1; se2; :::; seTt ). The sememe
embedding layer is randomly initialized and is learned along with pretraining.
Knowledge is injected via multi-head attention. The token embeddings H0 are treated
as query, and the sememe embeddings SE0 are treated as key and value. Knowledge
enriched representation of the sentence HS is obtained by multi-head attention from the
query to the sememes. We call this knowledge injection module as attn-to-sememes.
We include the attn to sememes in the pretraining stage. During pretraining, for masked
language modeling (MLM), the masked tokens will be treated as tokens with no sememes,
i.e., it only has a padding sememe &lt; s pad &gt;. We believe that including sememe
knowledge during pretraining can help to speed up learning semantic meanings of tokens.
        </p>
        <p>For pretrained H-BERT, we can fine-tune it in two approaches: (a) discarding the
attn-to-sememe module and fine-tuning H-BERT in the same way as BERT; (b) taking
advantage of HowNet during fine-tuning, that is, extracting the sememe informaitons
from the sentence, and using the attn-to-sememe module to inject sememe information
to H-BERT’s encoder.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>3.1</p>
      <sec id="sec-3-1">
        <title>Experimental Setup</title>
        <p>
          Pretraining is done on the Chinese Wikipedia corpus. We use the vocabulary of Google
Chinese Bert [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] for tokenization. We pretrain three models totally from scratch: (a)
ALBERT base [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]; (b) ALBERT large [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]; (c) H-BERT v0, which uses a randomly
initialized ALBERT base as the encoder and includes an attn-to-sememe module on
the embedding layer; (d) H-BERT v1, which is in the large model setting; (e) H-BERT
v2, which is a base model and puts the attn-to-sememe module on the last layer of its
Transformer encoder. (f) H-BERT v3, which is H-BERT v0 fine-tuned without
attn-tosememe. For H-BERT v0 and H-BERT v2, the hidden size of transformers is reduced
to 640. The hidden size of H-BERT v1 is set to 980. Moreover, the attn-to-sememe
module reuses the encoder’s parameters. Thus, the number of parameters is comparable
to ALBERT base. We apply our pretrained vanilla Albert base on the open-sourced codes
of K-Bert [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] to obtain the results.
        </p>
        <p>
          During fine-tuning, we mainly follow the hyper-params of [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Each model runs 10
times to ensure reproducibility.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Datasets</title>
        <p>
          We experiment across a diverse set of 6 benchmark NLP tasks and demonstrate the
effectiveness of our model. For text classification, we select ChnSentiCorp (chn)4. For
NLI, we select LCQMC [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] (lcqmc) and XNLI [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] (xnli). For named entity recognition
task, three datasets from different domains are selected. MSRA NER (msra ner) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] is
from the open domain, Finance NER5 (fin ner) is from the financial domain, and CCKS
NER6 (ccks ner) is collected from the medical records.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Results and Analysis</title>
        <p>The experimental results are reported in Table 1. The main takeaways are:
– H-BERT v0 consistently outperforms ALBERT base and K-BERT with comparable
parameters, demonstrating H-BERT’s effectiveness.
– H-BERT v1 outperforms ALBERT large, demonstrating that our method can also
work for large pretrained models.
– H-BERT v3 performs worse than H-BERT v0, but it is better than ALBERT base,
showing that attn-to-sememe helps improve the generalization ability of pretrained
models. In addition, adopting attn-to-sememe during fine-tuning is beneficial for
downstream tasks.
– Comparing H-BERT v2 with H-BERT v0, we can see that it is better to apply
attn-to-sememes at the embedding layer of the encoder model.
4 https://github.com/pengming617/bert classification
5 https://embedding.github.io/evaluation/#extrinsic
6 https://biendata.com/competition/CCKS2017 2/</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 Conclusion</title>
      <p>This article proposes to enhance the Chinese pre-trained language models with simple
attention to sememes module. Experiments on 6 benchmark datasets shows that: From
the architecture point of view, attn-to-sememes should be applied at the embedding layer;
(2) attn-to-sememes are required at both pretraining and fine-tuning stages for better
downstream performances. Our model can beat the vanilla ALBERT significantly across
6 datasets with roughly the same amount of parameters, showing that our model can
effectively inject knowledge into PLMs.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Conneau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rinott</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lample</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwenk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoyanov</surname>
          </string-name>
          , V.:
          <article-title>XNLI: Evaluating cross-lingual sentence representations</article-title>
          .
          <source>In: EMNLP 2018</source>
          . pp.
          <fpage>2475</fpage>
          -
          <lpage>2485</lpage>
          . Association for Computational Linguistics, Brussels, Belgium (Oct-Nov
          <year>2018</year>
          ). https://doi.org/10.18653/v1/
          <fpage>D18</fpage>
          -1269, https://www.aclweb.org/anthology/D18-1269
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goodman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gimpel</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soricut</surname>
            ,
            <given-names>R.: ALBERT:</given-names>
          </string-name>
          <article-title>A Lite BERT for Self-supervised Learning of Language Representations</article-title>
          . arXiv e-prints arXiv:
          <year>1909</year>
          .
          <volume>11942</volume>
          (
          <year>Sep 2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Levow</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          :
          <article-title>The third international Chinese language processing bakeoff: Word segmentation and named entity recognition</article-title>
          .
          <source>In: Proceedings of the Fifth SIGHAN Workshop on Chinese Language Processing</source>
          . pp.
          <fpage>108</fpage>
          -
          <lpage>117</lpage>
          . Association for Computational Linguistics, Sydney,
          <source>Australia (Jul</source>
          <year>2006</year>
          ), https://www.aclweb.org/anthology/W06-0115
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ju</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
          </string-name>
          , P.:
          <string-name>
            <surname>K-BERT: Enabling Language</surname>
          </string-name>
          <article-title>Representation with Knowledge Graph</article-title>
          . arXiv e-prints arXiv:
          <year>1909</year>
          .
          <volume>07606</volume>
          (
          <year>Sep 2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>LCQMC:a large-scale Chinese question matching corpus</article-title>
          .
          <source>In: CL</source>
          . pp.
          <fpage>1952</fpage>
          -
          <lpage>1962</lpage>
          . ACL,
          <string-name>
            <surname>Santa</surname>
            <given-names>Fe</given-names>
          </string-name>
          , New Mexico, USA (Aug
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoyanov</surname>
          </string-name>
          , V.:
          <article-title>RoBERTa: A Robustly Optimized BERT Pretraining Approach</article-title>
          . arXiv e-prints arXiv:
          <year>1907</year>
          .
          <volume>11692</volume>
          (
          <year>Jul 2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Niu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Improved word representation learning with sememes</article-title>
          .
          <source>In: ACL</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Logan</surname>
            , Robert L.,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          :
          <article-title>Knowledge Enhanced Contextual Word Representations</article-title>
          . arXiv e-prints arXiv:
          <year>1909</year>
          .
          <volume>04164</volume>
          (
          <year>Sep 2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Qi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Openhownet: An open sememe-based lexical knowledge base</article-title>
          . arXiv preprint arXiv:
          <year>1901</year>
          .
          <volume>09957</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Replace or Retrieve Keywords In Documents at Scale</article-title>
          . arXiv e-prints
          <source>arXiv:1711.00046 (Oct</source>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michael</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          :
          <article-title>Glue: A multi-task benchmark and analysis platform for natural language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1804</year>
          .
          <volume>07461</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
          </string-name>
          , J.:
          <article-title>Kepler: A unified model for knowledge embedding and pre-trained language representation</article-title>
          . arXiv preprint arXiv:
          <year>1911</year>
          .
          <volume>06136</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Enhancing transformer with sememe knowledge</article-title>
          .
          <source>In: RepL4NLP@ACL</source>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          :
          <article-title>ERNIE: Enhanced Language Representation with Informative Entities</article-title>
          . arXiv e-prints arXiv:
          <year>1905</year>
          .
          <volume>07129</volume>
          (May
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>