<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Understood in Translation: Transformers for Domain Understanding</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dimitrios Christofidellis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Manica</string-name>
          <email>tte@zurich.ibm.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leonidas Georgopoulos</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hans Vandierendonck</string-name>
          <email>h.vandierendonck@qub.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IBM Research Europe</string-name>
          <email>dic@zurich.ibm.com</email>
          <email>leg@zurich.ibm.com</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Queen's University Belfast</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Knowledge acquisition is the essential first step of any Knowledge Graph (KG) application. This knowledge can be extracted from a given corpus (KG generation process) or specified from an existing KG (KG specification process). Focusing on domain specific solutions, knowledge acquisition is a labor intensive task usually orchestrated and supervised by subject matter experts. Specifically, the domain of interest is usually manually defined and then the needed generation or extraction tools are utilized to produce the KG. Herein, we propose a supervised machine learning method, based on Transformers, for domain definition of a corpus. We argue why such automated definition of the domain's structure is beneficial both in terms of construction time and quality of the generated graph. The proposed method is extensively validated on three public datasets (WebNLG, NYT and DocRED) by comparing it with two reference methods based on CNNs and RNNs models. The evaluation shows the efficiency of our model in this task. Focusing on scientific document understanding, we present a new health domain dataset based on publications extracted from PubMed and we successfully utilize our method on this. Lastly, we demonstrate how this work lays the foundation for fully automated and unsupervised KG generation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Knowledge Graphs (KGs) are among the most
popular data management paradigms and their application is
widespread across different fields, e.g., recommendation
systems, question-answering tools and knowledge discovery
applications. This is due to the fact that KGs share
simultaneously several advantages of databases (information
retrieval via structured queries), graphs (representing loosely
or irregularly structured data) and knowledge bases
(representing semantic relationship among the data). KG research
can be divided in two main streams
        <xref ref-type="bibr" rid="ref13">(Ji et al. 2020)</xref>
        :
knowledge representation learning, which investigates the
representation of KG into vector representations (KG
embeddings), and knowledge acquisition, which considers the KG
generation process. The latter being a fundamental aspect
since a malformed graph will not be able to serve reliably
any kind of downstream task.
      </p>
      <p>
        The knowledge acquisition process is either referred to
KG construction, where the KG is built from scratch
using a specific corpus, or KG specification, where a subgraph
of interest is extracted from an existing KG. In both cases,
the acquisition process can follow a bottom-up or top-down
approach
        <xref ref-type="bibr" rid="ref51">(Zhao, Han, and So 2018)</xref>
        . In a bottom-up
approach, all the entities and their connections are extracted
as a first step of the process. Then, the underlying hierarchy
and structure of the domain can be inferred from the
entities and their connections. Conversely, a top-down approach
starts with the definition of the domain’s schema that is then
used to guide the extraction of the needed entities and
connections. For general KG generation, a bottom-up approach
is usually preferred as we typically wish to include all
entities and relations that we can extract from the given
corpus. Contrarily, a top-down approach better suits a
domainspecific KG generation or KG specification, where entities
and relations are strongly linked to the domain of interest.
The structure of typical bottom-up and top-down pipelines,
focusing on the case of KG generation, are presented in
figures 1a and 1b respectively.
      </p>
      <p>
        Herein, we focus on domain-specific, i.e., top-down,
acquisition for two main reasons. Firstly, the acquisition
process can be faster and more accurate in this way. By
specifying the schema of the domain of interest, then we only need
to select the proper and needed tools (i.e. pretrained
models) for the actual entity and relation extraction. Secondly,
such approach minimizes the presence of irrelevant data and
restricts queries and graph operations to a carefully tailored
KG. This generally improves the accuracy of KG
applications
        <xref ref-type="bibr" rid="ref18">(Lalithsena, Kapanipathi, and Sheth 2016)</xref>
        .
Furthermore, the graph’s size is significantly reduced by excluding
irrelevant content. Thus, execution time of queries can be
reduced by more than one order of magnitude
        <xref ref-type="bibr" rid="ref18">(Lalithsena,
Kapanipathi, and Sheth 2016)</xref>
        .
      </p>
      <p>The domain definition is usually performed by subject
matter experts. Yet, knowledge acquisition by expert
curation can be extremely slow as the process is essentially
manual. Moreover, human error may affect the data
quality and lead to malformed KGs. In this work, we propose
to overcome these issues by introducing an automated
machine learning-based approach to understand the domain of
a collection of text snippets. Specifically, given sample
input texts, we infer the schema of the domain to which they
(a) Bottom-up pipeline.
(b) Top-down pipeline.
belong. This task can be incorporated into both
domainspecific KG generation and KG specification process, where
the domain definition is the essential first step. For the KG
generation, the input texts can be samples from the corpus
of interest, while for the KG specification, these text
snippets can express possible questions that need to be answered
from the specified KG. We introduce a seq2seq-based model
relying on transformer architecture to infer the relation types
characterizing the domain of interest. Such model lets us
to define the domain’s schema including all the needed
entity and relation types. The model can be trained using any
available previous schema (i.e., schema of a general KG like
DBpedia) and respective text examples for each possible
relation type. We show that our proposed model outperforms
other baseline approaches, it can be successfully utilized for
scientific documents and it has interesting potential
extension in the field of automated KG generation.</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        At the best of our knowledge, our method is the first attempt
to introduce a supervised machine learning based domain
understanding tool that can be incorporated into
domainspecific KG generation and specification pipelines.
Currently, the main research interest related to KG generation
workflows is associated with attempts to improve the named
entity recognition (NER) and the relation extraction tasks or
provide end-to-end pipelines for general or domain-specific
KG generation
        <xref ref-type="bibr" rid="ref13">(Ji et al. 2020)</xref>
        . The majority of such work
focuses on the actual generation step and rely solely on
manual identification of the domain definition
        <xref ref-type="bibr" rid="ref24 ref27 ref43">(Luan et al. 2018;
Manica et al. 2019; Wang et al. 2020)</xref>
        .
      </p>
      <p>
        As it concerns the KG specification field, the subgraph
extraction is usually based on graph traversals or more
sophisticated heuristics techniques and some providing initial
entities or entity types
        <xref ref-type="bibr" rid="ref18">(Lalithsena, Kapanipathi, and Sheth
2016)</xref>
        . Such approaches are effective, yet a significant
engineering effort is required to tune the heuristics for each
different case. Let alone, the crucial task of proper
selection of the initial entities or entity types is mainly performed
manually.
      </p>
      <p>
        The relation extraction task is also related to our work.
It aims at the extraction of triplets of the form of (subject,
relation, object) from the texts. The neural network based
methods, such as Nguyen and Grishman; Zhou et al.; Zhang
et al., dominate the field. These methods are CNN
        <xref ref-type="bibr" rid="ref31 ref47">(Zeng et
al. 2014; Nguyen and Grishman 2015)</xref>
        or LSTM
        <xref ref-type="bibr" rid="ref50">(Zhou et al.
2016; Zhang et al. 2017)</xref>
        models, which attempt to identify
relations in a text given its content and information about
the position of entities in it. The positional information of
the entities is typically extracted in a previous step of KG
generation using NER methods
        <xref ref-type="bibr" rid="ref30">(Nadeau and Sekine 2007)</xref>
        .
Lately, there is a high interest of methods that can combine
the NER and relation extraction tasks into a single model
        <xref ref-type="bibr" rid="ref41 ref49 ref7">(Zheng et al. 2017; Zeng et al. 2018; Fu, Li, and Ma 2019)</xref>
        .
      </p>
      <p>While our work is linked to relation extraction it has two
major differences. Firstly, we focus on the relation type and
the entity types that compose a relation rather than the actual
triplet. Secondly, the training process differs and requires
coarser annotations. We solely provide texts and the
respective existing sequence of relation types. Contrarily in a
typical relation extraction training process, information about
the position of the entities in the text is also needed. Here,
we propose to improve knowledge acquisition by
performing a data-driven domain definition providing an approach
that is currently unexplored in KG research.</p>
    </sec>
    <sec id="sec-3">
      <title>Seq2seq-based model for domain understanding</title>
      <p>The domain understanding task attempts to uncover the
structured knowledge underlying a dataset. In order to
depict this structure we can leverage the so called domain’s
metagraph. A domain’s metagraph is a graph that has as
vertices all the entity types and as edges all their
connections/relations in the context of this domain. The generation
of such a metagraph entails obtaining all the entity types and
their relations. Assuming that each of entity types that are
presented in the domain has at least one interaction with
another entity type, the metagraph of this domain can be
produced by inferring all the possible relation types as all the
entity types are included in at least one of them. Thus, our
approach aims to build an accurate model to detect a
domain’s relation types, and leverages this model to extract
those relations from a given corpus. Aggregating all
extracted relations yields the domain’s metagraph.</p>
      <sec id="sec-3-1">
        <title>Seq2seq model for domain’s relation types extraction</title>
        <p>
          Sequence to sequence models (seq2seq)
          <xref ref-type="bibr" rid="ref1 ref1 ref15 ref17 ref17 ref3 ref3 ref32 ref32 ref39 ref46 ref46">(Bahdanau, Cho,
and Bengio 2014; Cho et al. 2014; Sutskever, Vinyals, and
Le 2014; Jozefowicz et al. 2016)</xref>
          attempt to learn the
mapping from an input X to its corresponding target Y where
both of them are represented by sequences. To achieve this
they follow an encoder-decoder based approach. Encoders
and decoders can be recurrent neural networks
          <xref ref-type="bibr" rid="ref1 ref3">(Cho et al.
2014)</xref>
          or convolutional based neural networks
          <xref ref-type="bibr" rid="ref9">(Gehring et
al. 2017)</xref>
          . In addition, an attention mechanism can also be
incorporated into the encoder
          <xref ref-type="bibr" rid="ref1 ref17 ref26 ref3 ref31 ref32 ref46">(Bahdanau, Cho, and Bengio
2014; Luong, Pham, and Manning 2015)</xref>
          for further boosting
of the model’s performance. Lately, Transformer
architectures
          <xref ref-type="bibr" rid="ref22 ref35 ref4 ref40">(Vaswani et al. 2017; Devlin et al. 2018; Liu et al. 2019;
Radford et al. 2018)</xref>
          , a family of models whose components
are entirely made up of attention layers, linear layers and
batch normalization layers, have established themselves as
the state of the art for sequence modeling, outperforming
the typically recurrent based components. Seq2seq models
have been successfully utilized for various tasks such as
neural machine translation
          <xref ref-type="bibr" rid="ref1 ref17 ref3 ref32 ref46">(Bahdanau, Cho, and Bengio 2014)</xref>
          and natural language generation (Pust et al. 2015). Recently,
their scope has also been extended beyond language
processing in fields such as chemical reaction prediction
          <xref ref-type="bibr" rid="ref37">(Schwaller
et al. 2019)</xref>
          .
        </p>
        <p>We consider the domain’s relation type extraction task as
a specific version of machine translation from the language
of the corpus to the “relation” language that includes all the
different relations between the entity types of the domain.
A relation type R which connects the entity type i to j is
represented as “i.R.j” in the “relation” language. In the case
of undirected connections, “i.R.j” is the same as “j.R.i” and
for simplicity we can discard one of them.</p>
        <p>
          Seq2seq models have been designed to address tasks
where both the input and the output sequences are ordered.
In our case the target “relation” language does not have any
defined ordering as per definition the edges of a graph do
not have any ordering. In theory the order does not
matter, yet in practice unordered sequences will lead to slower
convergence of the model and requirements of more
training data to achieve our goal
          <xref ref-type="bibr" rid="ref19 ref31 ref42">(Vinyals, Bengio, and Kudlur
2015)</xref>
          . To overcome this issue, we propose a specific
ordering of the “relation” language influenced from the semantic
context that the majority of the text snippets hold.
        </p>
        <p>
          According to
          <xref ref-type="bibr" rid="ref49">(Zeng et al. 2018)</xref>
          , in the context
of relation extraction, text snippets can be divided
into three types: Normal, EntityPairOverlap and
SingleEntityOverlap. A text snippet is categorized
as Normal if none of its triplets have overlapping entities.
If some of its triplets express a relation on the same pair
of entities, then it belongs to the EntityPairOverlap
category and if some of its triplets have one entity in
common but no overlapped pairs, then it belongs to
SingleEntityOverlap class. These three categories
are also relevant in the metagraph case, even if we are
working with entity types and relation types rather than the actual
entities and their relations.
        </p>
        <p>Based on the given training set, we consider that the
model is aware of a general domain anatomy, i.e., the sets
of possible entity types and relation types are known, and
we would like to identify which of them are depicted in a
given corpus. In both cases of EntityPairOverlap and
SingleEntityOverlap type text snippets, there is one
main entity type from which all the other entity types can be
found by performing only one hop traversal in the general
domain’s metagraph. The class of Normal text snippets is a
broader case in which one can identify heterogeneous
connectivity patterns among the entity types represented. Yet,
a sentence typically describes facts that are expected to be
connected somehow, thus the entity types included in such
texts usually are not more than 1 or 2 hops away from each
other in the general metagraph. In the light of the
considerations above, we propose to sort the relations in a
breadthfirst-search (BFS) order starting from a specific node (entity
type) in the general metagraph. In this way, we confine the
output in a much lower dimensional space by adhering to a
semantically meaningful order.</p>
        <p>Inspired by state-of-the-art approaches in the field of
neural machine translation, our model architecture is a
multi-layer bidirectional Transformer. We follow the lead
of Vaswani et al. in implementing the architecture, with the
only difference that we adopt a learned positional
encoding instead of a static one (see Appendix for further details
on the positional encoding). As the overall architecture of
the encoder and the decoder are otherwise the same as in
Vaswani et al., we omit an in-depth description of the
Transfomer model and refer readers to original paper.</p>
        <p>To boost the model’s performance, we also propose an
ensemble approach exploiting different Transformers and
aggregating their results to construct the domain’s metagraph.
Each of the Transformers differs in the selected ordering of
the “relation” vocabulary. The selection of different starting
entity type for the breadth-first-search will lead to different
orderings. We expect that multiple orderings could facilitate
the prediction of different connection patterns that can not
be easily detected using a single ordering. The sequence of
steps for an ensemble domain understanding is the
following: Firstly, train k Transformers using different orderings.
Secondly, given a set of text snippets, predict sequences of
relations using all the Transformers. Finally, use late fusion
to aggregate the results and form the final predictions.</p>
        <p>It is worth mentioning that in the last step, we omit the
underlying ordering that we follow in each model and we
perform a relation-based aggregation. We examine each relation</p>
        <sec id="sec-3-1-1">
          <title>Dataset</title>
          <p>WebNLG</p>
          <p>NYT</p>
          <p>
            DocRED
PubMed-DU
separately in order to include it or not in the final metagraph.
For the aggregation step, we use the standard Wisdom of
Crowds (WOC)
            <xref ref-type="bibr" rid="ref29">(Marbach et al. 2012)</xref>
            consensus technique,
yet other consensus methods can also be leveraged for the
task. The overall structure of our approach is summarized in
Figure 2.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <p>
        We evaluate our Transformer-based approach against three
baselines on a selection of datasets representing different
domains. As baselines, we use CNN and RNN based methods
influenced by
        <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
        and (Zhou et al.
2016) respectively. For the CNN based method, we slightly
modified the architecture to exclude the component which
provides information about the position of the entities in the
text snippet, as we do not have such information available in
our task. Additionally, we also include a Transformer-based
model without applying any ordering in the target sequences
as an extra baseline.
      </p>
      <p>
        To our knowledge, there is no standard dataset available
for the relation type extraction task in the literature.
However there is a plethora of published datasets for the
standard task of relation extraction that can be utilized for our
case with limited effort. For our task, the leveraged datasets
should contain tuples of texts and their respective sets of
relation types. We use WebNLG
        <xref ref-type="bibr" rid="ref8">(Gardent et al. 2017)</xref>
        , NYT
        <xref ref-type="bibr" rid="ref36">(Riedel, Yao, and McCallum 2010)</xref>
        and DocRED
        <xref ref-type="bibr" rid="ref45">(Yao et al.
2019)</xref>
        , three of the most popular datasets for relation
extraction. Both NYT and DocRED datasets provide the needed
information such as entity types and relation type for the
triplets of each instance. Thus their transformation for our
task can be conducted by just converting these triplets to the
relation type format, for instance the triplet (x,y,z) will be
transformed as type(x).type(y).type(z). On the other hand,
WebNLG doesn’t share such information for the entity types
and thus manual curation is needed. Therefore, all the
possible entities are examined and replaced with the proper entity
type. For the WebNLG dataset, we avoid including rare
entity and relation types which are occurred less than 10 times
in the dataset. We either omit them or replace them with
similar or more general types that exists in it.
      </p>
      <p>
        To emphasize the application of such model in the
scientific document understanding, we produce a new
taskspecific dataset called PubMed-DU related to the general
health domain. We download paper abstracts from PubMed
focusing on work related to 4 specific health subdomains:
Covid-19, mental health, breast cancer and coronary heart
disease. We split the abstracts into sentences. The entities
and their types for each text have been extracted using
PubTator
        <xref ref-type="bibr" rid="ref44">(Wei et al. 2019)</xref>
        . The available entity types are Gene,
Mutation, Chemical, Disease and Species. The respective
relation types are in form x.to.y where x and y are two of the
possible entity types. We assume that the relations are
symmetric. For text annotation, the following rule was used: a
text has the relation x.to.y if two entities with types x and
y co-occurred in the text and the syntax path between them
contains at least one keyword of this relation type. These
keywords have been manually identified based on the
provided instances and are words, mainly verbs, related to the
relation. Table 1 depicts the statistics of all four utilised
datasets.
      </p>
      <p>
        For all datasets, we use the same model parameters.
Specifically, we use Adam
        <xref ref-type="bibr" rid="ref17 ref32 ref46">(Kingma and Ba 2014)</xref>
        optimizer
with a learning rate of 0.0005. The gradients norm is clipped
to 1.0 and dropout
        <xref ref-type="bibr" rid="ref19 ref31 ref42">(LeCun, Bengio, and Hinton 2015)</xref>
        is set
to 0.1. Both encoder and decoder consist of 2 layers with
10 attention heads each, the positional feed-forward hidden
dimension is 512. Lastly, we utilize the token embedding
layers using GloVe pretrained word embeddings
        <xref ref-type="bibr" rid="ref17 ref32 ref33 ref46">(Pennington, Socher, and Manning 2014)</xref>
        which have
dimensionality of m=300. Our code and the datasets are available at
https://github.com/christofid/DomainUnderstanding.
      </p>
      <p>
        The evaluation of the models is performed both at instance
and graph level. In the instance level, we examine the ability
of the model to predict the relation types that exist in a given
text. To investigate this, we use F1-score and accuracy.
F1score is the harmonic mean of model’s precision and recall.
Accuracy is computed at an instance level and it measures
for how many of the testing texts, the model manage to infer
correctly the whole set of their relation types. For the
metagraph level evaluation, we use our model to predict the
metagraph of a domain and we examine how close to the actual
metagraph is. For this comparison, we utilize F1-score for
both edges and nodes of the metagraph as well as the
similarity of the distribution of the degree and eigenvector
centrality
        <xref ref-type="bibr" rid="ref17 ref32 ref46">(Zaki and Meira 2014)</xref>
        of the two metagraphs. For the
comparison of the centralities distribution, we construct the
histogram of the centralities for each graph using 10 fixed
size bins and we utilize Jensen-Shannon Divergence (JSD)
metric
        <xref ref-type="bibr" rid="ref6">(Endres and Schindelin 2003)</xref>
        to examine the
similarity of the two distributions (see Appendix for the definition
of JSD). We have selected degree and eigenvector
centralities as the former gives as localized structure information as
measure the importance of a node based on the direct
connections of it and the latter gives as a broader structure
information as measure the importance of a node based on infinite
walks.
      </p>
      <sec id="sec-4-1">
        <title>WebNLG NYT</title>
      </sec>
      <sec id="sec-4-2">
        <title>DocRED</title>
      </sec>
      <sec id="sec-4-3">
        <title>PubMed-DU</title>
        <p>
          CNN
          <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
          *
RNN (Zhou et al. 2016)*
        </p>
        <p>Transformer - unordered
Transformer - BFSrecord label</p>
        <p>
          Transformer - WOC k=20
CNN
          <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
          *
RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSperson
        </p>
        <p>
          Transformer - WOC k=8
CNN
          <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
          *
RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSORG
        </p>
        <p>
          Transformer - WOC k=6
CNN
          <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
          *
RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSSpecies
Transformer - WOC k=5
0.8156
0.8517
0.8798
0.9000
0.9235
        </p>
        <sec id="sec-4-3-1">
          <title>Instance level evaluation of the models</title>
          <p>To study the performance of our model, we perform 10
independent runs each with different random splitting of the
datasets into training, validation and testing set. Table 2
depicts the median value and the standard error of the baselines
and our method for the two metrics. Our method is better in
terms of accuracy for all the four datasets and in terms of
F1score for the WebNLG and DocRED datasets. For the NYT
and PubMed-DU datasets, the F1-score of CNN and RNN
models outperform our approach. We observed that the
baseline models profit from the fact that, in these datasets, the
majority of the instances depict only one relation and many
of the relations appear in a limited number of instances. In
general, there is lack of sequences of relations that hinders
the Transformer’s ability to learn the underlying distribution
in these two cases(see Appendix). Lastly, the decreased
performances of all the models in the DocRED dataset is due
to the long tail characteristic that this dataset shows as 66%
of the relations appeared in no more than 50 instances (see
Appendix).</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>Metagraph level evaluation of the models</title>
          <p>The above comparisons focus only on the ability of the
model to predict the relation types given a text snippet. Since
our ultimate goal is to infer the domain’s metagraph from a
given corpus, we divide the testing sets of the datasets into
small corpora and we attempt to define their domain using
our model. For WebNLG, NYT and DocRED dataset, 10
artificial corpora and their respective domains have been
created by selecting randomly 10 instances from each of the
testing sets. We set two constraints into this selection to
assure that the produced metagraphs are meaningful. Firstly,
each subdomain should have a connected metagraph and
secondly each existing relation type is appeared at least two
times in the provided instances. For the PubMed-DU dataset,
we already know the existence of 4 subdomains in it, so we
focus on the inference of them. For each subdomain, we
select randomly 100 instances from the testing set that belong
to this subdomain and we attempt to produce the domain
based on them. We infer the relation types for each instance
and then we generate the domain’s metagraph by including
all the relation types that were found in the instances. Then,
we compare how close the actual domain’s metagraph and
the predicted metagraph are.</p>
          <p>Table 3 presents the results of the evaluation of the
predicted versus the actual domain’s metagraph for 10
subdomains extracted from the testing set of the WebNLG, NYT
and DocRED datasets. All the presented values for these
datasets are the mean over all the 10 subdomains. For the
PubMed-DU dataset, we include only the Covid-19
subdomain case. Results for the remaining subdomains of this
dataset can be found in the Appendix. Our approach using
Transformer + BFS based ordering outperforms or is close
to the baselines for all cases in terms of edges and nodes
F1score. Furthermore, the degree and eigenvector centralities
distribution of the generated metagraphs using our method
are closer to the groundtruth in comparison to other
methods in all cases. This indicates that the graphs produced with
our method are both element-wise and structurally closer to
the actual ones. More detailed comparisons of the different
methods at both instance and metagraph level have been
included in the Appendix.</p>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>CNN (Nguyen and Grishman 2015)* RNN (Zhou et al. 2016)* Transformer - unordered</title>
        <p>Transformer - BFSrecord label</p>
        <p>
          Transformer - WOC k=5
CNN
          <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
          *
        </p>
        <p>RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSperson</p>
        <p>
          Transformer - WOC k=8
CNN
          <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
          *
        </p>
        <p>RNN (Zhou et al. 2016)*
Transformer - unordered</p>
        <p>Transformer - BFSPER</p>
        <p>
          Transformer - WOC k=6
CNN
          <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
          *
        </p>
        <p>RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSChemical
Transformer - WOC k=5
0.9747
0.9639
0.9598
0.9806
0.9808
0.9059
0.9205
0.8184
0.8806
0.8672
0.4819
0.6823
0.7530
0.7830
0.8045
0.9140
0.9736
0.9631
0.9583
0.9789</p>
        <p>The ensemble variant of our approach, based on the WOC
consensus strategy, outperforms the simple Transformer +
BFS ordering in all cases. Based on the evaluation at both
instance and metagraph level, our ensemble variant seems
to be the most reliable approach for the task of domain’s
relation type extraction as it achieves some of the best scores
for any dataset and metric.</p>
        <sec id="sec-4-4-1">
          <title>Towards automated KG generation</title>
          <p>
            The proposed domain understanding method enables the
inference of the domain of interest and its components. This
enables a partial automation and a speed up of the KG
generation process as, without manual intervention, we are able
to identify the metagraph, and inherently the needed
models for the entity and relation extraction in the context of the
domain of interest. To achieve this, we adopt a
Transformerbased approach that relies heavily on attention mechanisms.
Recent efforts are focusing on the analysis of such
attention mechanisms to explain and interpret the predictions
and the quality of the models
            <xref ref-type="bibr" rid="ref12 ref41 ref41">(Vig and Belinkov 2019;
Hoover, Strobelt, and Gehrmann 2019)</xref>
            . Interestingly, it has
been shown how the analysis of the attention pattern can
elucidate complex relations between the entities fed as input to
the Transformer, e.g., mapping atoms in chemical reactions
with no supervision
            <xref ref-type="bibr" rid="ref38">(Schwaller et al. 2020)</xref>
            . Even if it is out
of the scope of our current work, we observe that a similar
analysis of the attention patterns in our model can identify
not only parts of text in which relations exist but directly
the entities of the respective triplets. To illustrate this and
emphasize its application in the domain understanding field,
we extract 24 text instances from the PubMed-DU dataset
related to the COVID-19 domain. After generating the
domain’s metagraph, we analyze the attention to triples to build
a KG. We rely on the syntax dependencies to propagate the
attention weights throughout the connected tokens and we
examine the noun chunks to extract the entities of interest
based on their accumulated attention weight (see Appendix
for further details). We select the head which achieves the
best accuracy in order to generate the KG. Figure 3 depicts
the generated metagraph and the KG. Using the
aforementioned attention analysis, we manage to achieve 82% and
64% accuracy in the entity extraction and the relation
extraction respectively. These values might not be able to compete
the state of the art respective models and the investigation
is limited in only few instances. Yet it indicates that a
completely unsupervised generation based on attention analysis
is possible and deserves further investigation.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>Herein, we proposed a method to speed up the knowledge
acquisition process of any domain specific KG application
by defining the domain of interest in an automated manner.
This is achieved by using a Transformer-based approach to
estimate the metagraph representing the schema of the
domain. Such schema can indicate the proper and needed tools
for the actual entity and relation extraction. Thus our method
can be considering as the stepping stone in any KG
generation pipeline. The evaluation and the comparison over
different datasets against state-of-the-art methods indicates
that our approach produces accurately the metagraph.
Especially, in datasets where text instances contain multiple
relation types our model outperforms the baselines. This is an
important observation as text describing multiple relations is
the most common scenario. Based on that and relying on the
capability of the transformers to catch longer dependencies,
future investigation of how our model performs in larger
pieces of texts, like full paragraphs, could be interesting and
indicate a clearer advantage of our work. The needed
definition of a general domain for the training phase might be a
limitation of this method. However, schema and data from
already existing KGs can be utilized for training purposes.
Unsupervised or semi-supervised extension of this work can
also be explored in the future to mitigate the issue.</p>
      <p>
        Our work paves the way towards an automated
knowledge acquisition, as our model minimizes the need of
human intervention in the process. So in the near future the
currently needed manual curation can be avoided and lead
to faster and more accurate knowledge acquisition.
Interestingly, using the PubMed-DU dataset, we underline that our
method can be utilized for scientific documents. The
inference of their domain can assist both in their general
understanding but also lead to more robust knowledge acquisition
from them. As a side effect, it is also important to notice that,
such attention-based model can be directly applied to triplet
extraction from the text without retraining and without
supervision. Triplet extraction in an unsupervised way
represents a breakthrough, especially if combined with most
recent advances in zero-shot learning for NER
        <xref ref-type="bibr" rid="ref17 ref21 ref32 ref43 ref46">(Li et al. 2020;
Pasupat and Liang 2014; Guerini et al. 2018)</xref>
        . Further
analysis of our Transformer-based approach could give a better
insight into these capabilities.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Appendix</title>
      <sec id="sec-6-1">
        <title>Learned positional encoding</title>
        <p>In our model, we adopt a learned positional encoding instead
of a static one. Specifically, the tokens are passed through a
standard embedding layer as a first step in the encoder. The
model has no recurrent layers and therefore it has no idea
about the order of the tokens within the sequence. To
overcome this, we utilize a second embedding layer called a
positional embedding layer. This is a standard embedding layer
where the input is not the token itself but the position of the
token within the sequence, starting with the first token, the
&lt;sos&gt; (start of sequence) token, in position 0. The position
embedding has a ”vocabulary” size equal to the maximum
length of the input sequence. The token embedding and
positional embedding are element-wise summed together to get
the final token embedding which contains information about
both the token and its position within the sequence. This
final token embedding is then provided as input in the stack
of attention layers of the encoder.</p>
      </sec>
      <sec id="sec-6-2">
        <title>Dataset characteristics</title>
        <p>For a better understanding of the datasets, we analyzed the
distribution of occurrences for all relation types. These
distributions are depicted in Figure 4. A percentage of relation
types with number of appearances close or less to 10 is
observed for all datasets. The lack of many examples can pose
problems in the learning process for these specific relation
types. This is highlighted especially in the DocRED case, as
we attributed the decreased performances of all the models
in this dataset in its long tail characteristic that it holds.
Especially for the DocRED, almost the 50% of the relations
appeared in no more than 10 instances and the 66% of the
relations appeared in no more than 50 instances (4).</p>
      </sec>
      <sec id="sec-6-3">
        <title>Jensen-Shannon distance</title>
        <p>The Jensen-Shannon divergence metric between two
probability vectors p and q is defined as:
r D(p k m) + D(q k m)</p>
        <p>2
where m is the pointwise mean of p and q and D is the
Kullback-Leibler divergence.</p>
        <p>The Kullback-Leibler divergence for two probability
vectors p and q of length n is defined as:</p>
        <p>D(p k q) =
n
X pilog2( pqi )</p>
        <p>i
i=1</p>
        <p>The Jensen–Shannon metric is bounded by 1, given that
we use the base 2 logarithm.</p>
      </sec>
      <sec id="sec-6-4">
        <title>Triplets extraction based on attention analysis</title>
        <p>In this section, the procedure of automated triplets
extraction based on the predicted relation types and the respective
attention weights is described . We generate an undirected
graph that connects the tokens of the sentence based on their
syntax dependencies for each instance. Then for each
different predicted relation type, we define the final attention
weights of a token based on the attention weights of
itself and its neighbors in the syntax dependencies graph. Let
ar be the attention vector of a predefined model’s attention
head, which contains all the attention weights related to the
relation type r. The final attention weight w of the token i
for the relation r is defined as:
wir = 2 air +</p>
        <p>X ajr
j2neigi
where neigi is the set containing all the neighbors of i in
the syntax dependencies graph. Then for each noun chunk k
(nk) of the text we compute its total attention weight for the
relation type r as:
nrk =</p>
        <p>X f (wjr)
j2nck
where nck is the set of tokens which belong to the nk and
f is a function defined as
f (wjr) =
wr</p>
        <p>j
2
wr
j
if j is stop-word
if j is not stop-word</p>
        <p>Finally, we extract as entities which are connected via the
relation type r the two noun chunks with the highest weight
nr. As this work is a proof of concept rather than an
actual method, the selection of the attention head is based on
whichever gives as the best outcome. Yet, in actual scenarios
it is recommended the use of a training set, based on which
the optimal head will be identified. For the creation of the
syntax dependencies graph and the extraction of the noun
chunks of the text we use spacy1 and its en core web lg
pretrained model. Table 4 includes all the texts that have
used in the proof of concept that is presented in the main
paper and the respective predicted triplets for each of them.</p>
      </sec>
      <sec id="sec-6-5">
        <title>Models’s comparison</title>
        <p>Table 5 presents a detailed evaluation of our approach and
the baselines models. We have included the 3 best BFS
ordering variants (in terms of accuracy) and 3 consensus
variants. To cover the range of all the available values of k, [1,
number of entity types], we select a case with just a few
Transformers, one with a value close to half of the total
number of entity types and one close to the total number
of entity types. For each different k value, we utilize the top
k best orderings based on their accuracy. In addition to the
per instance accuracy and the per relation F1-score, the table
also includes the per relation precision and the recall of each
model. Our proposed method, especially its ensemble
variant, produces the best outcome in all datasets apart from the
NYT case where the CNN and RNN models manage to be
more precise. This is attributed to the characteristics of NYT
dataset, where there are many one-only relation instances.</p>
        <p>Similarly, in tables 6 and 7 we perform an in-depth
metagraph level evaluation of the models. For all three datasets,
our proposed method and its ensemble extension produce
the best or one of the top-3 best outcomes.</p>
        <sec id="sec-6-5-1">
          <title>1https://spacy.io/</title>
          <p>(a) WebNLG
(b) NYT
(c) DocRED (d) PubMed-DU</p>
          <p>Figure 4: Appearances distribution for the relation types of all utilized datasets.
Text
acute deep vein thrombosis in covid-19 hospitalized patients.
(ace-2) receptor for its attachment similar to sars-cov-1, which
is followed by priming of spike protein by transmembrane protease serine 2
(tmprss2) which can be targeted by a proven inhibitor of tmprss2, camostat.
temporal trends in decompensated heart failure and outcomes during covid-19:
a multisite report from heart failure referral centres in london.
prevalence, risk factors and clinical correlates of depression in quarantined
population during the covid-19 outbreak.
tocilizumab plus glucocorticoids in severe and critically covid-19 patients.
effects of progressive muscle relaxation on anxiety and sleep quality in
patients with covid-19.
special section: covid-19 among people living with hiv.
ace2 and tmprss2 variants and expression as candidates to sex and country
differences in covid-19 severity in italy.
did the covid-19 pandemic cause a delay in the diagnosis of acute appendicitis?
practice observed in managing gynaecological problems in post-menopausal
women during the covid-19 pandemic.
hypogammaglobulinemia causing pneumonia: an overlooked curable entity
in the chaotic covid-19 pandemic.
chemokine receptor gene polymorphisms and covid-19: could knowledge
gained from hiv/aids be important?
serotonin syndrome in two covid-19 patients treated with lopinavir/ritonavir.
mortality rate and predictors of mortality in hospitalized covid-19
patients with diabetes.
in conclusion, self-reported depression occurred at an early stage in convalescent
covid-19 patients, and changes in immune function were apparent during
short-term follow-up of these patients after discharge.
gb-2 inhibits ace2 and tmprss2 expression: in vivo and in vitro studies.
preemptive interleukin-6 blockade in patients with covid-19.
response to: ’clinical course of covid-19 in patients with systemic lupus
erythematosus under long-term treatment with hydroxychloroquine’ by carbillon et al.
obsessive-compulsive disorder during the covid-19 pandemic
repeated monitoring of ferritin, interleukin-6, c-reactive protein, lactic acid dehydrogenase,
and erythrocyte sedimentation rate during covid-19 treatment may assist the prediction of
disease severity and evaluation of treatment effects.
respiratory and pulmonary complications in head and neck cancer patients: evidence-based
review for the covid-19 era.
targeting the immune system for pulmonary inflammation and cardiovascular complications
in covid-19 patients.
risk of peripheral arterial thrombosis in covid-19.
preadmission diabetes-specific risk factors for mortality in hospitalized patients with
diabetes and coronavirus disease 2019.</p>
          <p>Predicted triplets
(thrombosis, Disease.to.Disease, COVID 19)
(spike, Gene.to.Gene, Transmembrane protease serine 2)
(heart failure, Disease.to.Disease, COVID-19)
(depression, Disease.to.Disease, COVID-19)
(Tocilizumab, Chemical.to.Species, patients)
(anxiety, Disease.to.Disease, COVID-19)
(people, Disease.to.Species, HIV)
(ACE2, Gene.to.Gene, TMPRSS2)
(COVID-19, Disease.to.Gene, TMPRSS2)
(COVID-19, Disease.to.Disease, Acute Appendicitis)
(COVID-19, Disease.to.Species, women)
(hypogammaglobulinemia, Disease.to.Disease, COVID-19)
(AIDS, Disease.to.Gene, Chemokine receptor)
(lopinavir/ritonavir, Chemical.to.Species, patients)
(COVID-19, Disease.to.Disease, Diabetes)
(depression, Disease.to.Disease, COVID-19)
(ACE2, Gene.to.Gene, TMPRSS2)
(COVID-19, Disease.to.Gene, interleukin-6)
(hydroxychloroquine, Chemical.to.Species, patients)
(Obsessive-compulsive disorder, Disease.to.Disease, COVID-19)
(COVID-19, Disease.to.Gene, interleukin-6)
(head and neck cancer, Disease.to.Disease, COVID-19)
(Cardiovascular Complications, Disease.to.Disease, COVID-19)
(Thrombosis, Disease.to.Disease, COVID-19)
(Diabetes, Disease.to.Disease, COVID-19)</p>
          <p>
            NYT
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            * 0.8156 0.0071
          </p>
          <p>RNN (Zhou et al. 2016)* 0.8517 0.0058
Transformer - unordered 0.8798 0.0053
Transformer - BFSoccupation 0.8987 0.0068
Transformer - BFSmusic genre 0.8983 0.0053
Transformer - BFSrecord label 0.9000 0.0046</p>
          <p>Transformer - WOC k=5 0.9210 0.0017
Transformer - WOC k=20 0.9235 0.0014</p>
          <p>
            Transformer - WOC k=45 0.9235 0.0002
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            * 0.7341 0.0035
          </p>
          <p>RNN (Zhou et al. 2016)* 0.7520 0.0027
Transformer - unordered 0.7426 0.0061
Transformer - BFSlocation 0.7461 0.0053
Transformer - BFSperson 0.7491 0.0048
Transformer - BFScompany 0.7461 0.0081
Transformer - WOC k=4 0.7547 0.0038
Transformer - WOC k=8 0.7669 0.0011</p>
          <p>
            Transformer - WOC k=12 0.7698 0.0007
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            * 0.1096 0.0073
          </p>
          <p>RNN (Zhou et al. 2016)* 0.2178 0.0088
Transformer - unordered 0.4869 0.0069
Transformer - BFSLOC 0.5235 0.0049
Transformer - BFSPER 0.5234 0.0077
Transformer - BFSORG 0.5252 0.0048
Transfomer - WOC k=4 0.5678 0.0037
Transformer - WOC k=5 0.5697 0.0016</p>
          <p>
            Transformer - WOC k=6 0.5722 0.0001
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            * 0.5573 0.0030
          </p>
          <p>RNN (Zhou et al. 2016)* 0.5772 0.0041
Transformer - unordered 0.5693 0.0075
Transformer - BFSSpecies 0.5707 0.0060
Transformer - BFSChemical 0.5703 0.0048</p>
          <p>Transformer - BFSGene 0.5718 0.0068
Transformer - WOC k=3 0.5909 0.0070
Transformer - WOC k=4 0.5685 0.0069
Transformer - WOC k=5 0.5960 0.0001</p>
          <p>Precision
WebNLG</p>
          <p>
            NYT
DocRED
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            *
          </p>
          <p>RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSoccupation
Transformer - BFSmusic genre
Transformer - BFSrecord label</p>
          <p>Transformer - WOC k=5
Transformer - WOC k=20</p>
          <p>
            Transformer - WOC k=45
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            *
          </p>
          <p>RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSlocation
Transformer - BFSperson
Transformer - BFScompany
Transformer - WOC k=4
Transformer - WOC k=8</p>
          <p>
            Transformer - WOC k=12
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            *
          </p>
          <p>RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSLOC
Transformer - BFSPER
Transformer - BFSORG
Transformer - WOC k=4
Transformer - WOC k=5
Transformer - WOC k=6
0.9879
0.9735
0.9775
0.9805
0.9746
0.9772
0.9772
0.9840
0.9916
0.9800</p>
          <p>1
0.9800
0.9800
0.9666
0.9657
1
1
1
0.9019
0.9714
1
1
0.9777
0.9777
1
1
1</p>
        </sec>
        <sec id="sec-6-5-2">
          <title>Covid-19</title>
        </sec>
        <sec id="sec-6-5-3">
          <title>Breast cancer</title>
        </sec>
        <sec id="sec-6-5-4">
          <title>Coronary heart diseases</title>
        </sec>
        <sec id="sec-6-5-5">
          <title>Mental health Model</title>
        </sec>
        <sec id="sec-6-5-6">
          <title>CNN (Nguyen and Grishman 2015)* RNN (Zhou et al. 2016)* Transformer - unordered</title>
          <p>Transformer - BFSSpecies
Transformer - BFSChemical</p>
          <p>Transformer - BFSGene
Transformer - WOC k=3
Transformer - WOC k=4</p>
          <p>
            Transformer - WOC k=5
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            *
RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSSpecies
Transformer - BFSChemical
          </p>
          <p>Transformer - BFSGene
Transformer - WOC k=3
Transformer - WOC k=4</p>
          <p>
            Transformer - WOC k=5
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            *
RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSSpecies
Transformer - BFSChemical
          </p>
          <p>Transformer - BFSGene
Transformer - WOC k=3
Transformer - WOC k=4</p>
          <p>
            Transformer - WOC k=5
CNN
            <xref ref-type="bibr" rid="ref31">(Nguyen and Grishman 2015)</xref>
            *
RNN (Zhou et al. 2016)*
Transformer - unordered
Transformer - BFSSpecies
Transformer - BFSChemical
          </p>
          <p>Transformer - BFSGene
Transformer - WOC k=3
Transformer - WOC k=4
Transformer - WOC k=5
0.9140
0.9736
0.9631
0.9583
0.9583
0.9525
0.9531
0.9642
0.9789
0.9135
0.9730
0.9367
0.9531
0.9531
0.9379
0.9525
0.9584
0.9514
0.8902
0.9498
0.9484
0.9703
0.9703
0.9756
0.9644
0.9703
0.9644</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Bahdanau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ; and Bengio,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>arXiv preprint arXiv:1409</source>
          .
          <fpage>0473</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ; Van Merrie¨nboer, B.;
          <string-name>
            <surname>Gulcehre</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bahdanau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bougares</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Schwenk</surname>
            , H.; and Bengio,
            <given-names>Y.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Learning phrase representations using rnn encoder-decoder for statistical machine translation</article-title>
          .
          <source>arXiv preprint arXiv:1406</source>
          .
          <fpage>1078</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Chang, M.-W.;
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Endres</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Schindelin</surname>
            ,
            <given-names>J. E.</given-names>
          </string-name>
          <year>2003</year>
          .
          <article-title>A new metric for probability distributions</article-title>
          .
          <source>IEEE Transactions on Information theory 49</source>
          (
          <issue>7</issue>
          ):
          <fpage>1858</fpage>
          -
          <lpage>1860</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Fu</surname>
          </string-name>
          , T.-J.;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          -H.; and Ma, W.-Y.
          <year>2019</year>
          .
          <article-title>Graphrel: Modeling text as relational graphs for joint entity and relation extraction</article-title>
          .
          <source>In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <fpage>1409</fpage>
          -
          <lpage>1418</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Gardent</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shimorina</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Narayan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; and PerezBeltrachini,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Creating training corpora for nlg micro-planning. In 55th annual meeting of the Association for Computational Linguistics (ACL).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Gehring</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Auli,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Grangier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Yarats</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          ; and Dauphin,
          <string-name>
            <surname>Y. N.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Convolutional sequence to sequence learning.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>In Proceedings of the 34th International Conference on Machine Learning-</source>
          Volume
          <volume>70</volume>
          ,
          <fpage>1243</fpage>
          -
          <lpage>1252</lpage>
          . JMLR. org.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          2018.
          <article-title>Toward zero-shot entity recognition in task-oriented conversational agents</article-title>
          .
          <source>In Proceedings of the 19th Annual SIGdial Meeting on Discourse and Dialogue</source>
          ,
          <volume>317</volume>
          -
          <fpage>326</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Hoover</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Strobelt</surname>
          </string-name>
          , H.; and
          <string-name>
            <surname>Gehrmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>exbert: A visual analysis tool to explore learned representations in transformers models</article-title>
          . arXiv preprint arXiv:
          <year>1910</year>
          .05276.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cambria</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Marttinen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>P. S.</given-names>
          </string-name>
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <article-title>A survey on knowledge graphs: Representation, acquisition and applications</article-title>
          . arXiv preprint arXiv:
          <year>2002</year>
          .00388.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Jozefowicz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Schuster</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shazeer</surname>
          </string-name>
          , N.; and
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Exploring the limits of language modeling</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>arXiv preprint arXiv:1602</source>
          .
          <fpage>02410</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D. P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ba</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>arXiv preprint arXiv:1412</source>
          .
          <fpage>6980</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Lalithsena</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Kapanipathi,
          <string-name>
            <given-names>P.</given-names>
            ; and
            <surname>Sheth</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Harnessing relationships for domain-specific subgraph extraction: A recommendation use case</article-title>
          .
          <source>In 2016 IEEE International Conference on Big Data (Big Data)</source>
          ,
          <fpage>706</fpage>
          -
          <lpage>715</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>LeCun</surname>
          </string-name>
          , Y.;
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; and Hinton,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Deep learning</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>nature</source>
          <volume>521</volume>
          (
          <issue>7553</issue>
          ):
          <fpage>436</fpage>
          -
          <lpage>444</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          .; Han,
          <string-name>
            <given-names>J</given-names>
            .; and
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <year>2020</year>
          .
          <article-title>A survey on deep learning for named entity recognition</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering.</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Joshi,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ;
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ; and
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <string-name>
            <surname>V.</surname>
          </string-name>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          arXiv preprint arXiv:
          <year>1907</year>
          .11692.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Luan</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>He</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ostendorf</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and Hajishirzi,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <article-title>Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction</article-title>
          . arXiv preprint arXiv:
          <year>1808</year>
          .09602.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Luong</surname>
          </string-name>
          , M.-T.;
          <string-name>
            <surname>Pham</surname>
          </string-name>
          , H.; and
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C. D.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Effective approaches to attention-based neural machine translation</article-title>
          .
          <source>arXiv preprint arXiv:1508</source>
          .
          <fpage>04025</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Manica</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zipoli</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dolfi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Staar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ; Laino,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Bekas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ;
            <surname>Fujita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Toda</surname>
          </string-name>
          , H.; et al.
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <article-title>An information extraction and knowledge graph platform for accelerating biochemical discoveries</article-title>
          . arXiv preprint arXiv:
          <year>1907</year>
          .08400.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Marbach</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Costello</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          ; Ku¨ffner, R.; Vega,
          <string-name>
            <given-names>N. M.</given-names>
            ;
            <surname>Prill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            ;
            <surname>Camacho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            ;
            <surname>Allison</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. R.</surname>
          </string-name>
          ; Kellis,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ; Collins,
          <string-name>
            <surname>J. J.;</surname>
          </string-name>
          and Stolovitzky,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2012</year>
          .
          <article-title>Wisdom of crowds for robust gene network inference</article-title>
          .
          <source>Nature methods 9</source>
          (
          <issue>8</issue>
          ):
          <fpage>796</fpage>
          -
          <lpage>804</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Nadeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Sekine</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>A survey of named entity recognition and classification</article-title>
          .
          <source>Lingvisticae Investigationes</source>
          <volume>30</volume>
          (
          <issue>1</issue>
          ):
          <fpage>3</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>T. H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Grishman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Relation extraction: Perspective from convolutional neural networks</article-title>
          .
          <source>In Proceedings of the 1st Workshop on Vector Space Modeling for Natural Language Processing</source>
          ,
          <fpage>39</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Pasupat</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Zero-shot entity extraction from web pages</article-title>
          .
          <source>In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          ,
          <fpage>391</fpage>
          -
          <lpage>401</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Socher, R.; and
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C. D.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          2015.
          <article-title>Parsing english into abstract meaning representation using syntax-based machine translation</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <fpage>1143</fpage>
          -
          <lpage>1154</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Narasimhan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Salimans</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Improving language understanding by generative pre-training.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Riedel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>Modeling relations and their mentions without labeled text</article-title>
          .
          <source>In Joint European Conference on Machine Learning and Knowledge Discovery in Databases</source>
          ,
          <volume>148</volume>
          -
          <fpage>163</fpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Schwaller</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ; Laino,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Gaudin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Bolgar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Hunter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            ;
            <surname>Bekas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. A.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Molecular transformer: A model for uncertainty-calibrated chemical reaction prediction</article-title>
          .
          <source>ACS Central Science</source>
          <volume>5</volume>
          (
          <issue>9</issue>
          ):
          <fpage>1572</fpage>
          -
          <lpage>1583</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>Schwaller</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hoover</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Reymond</surname>
            ,
            <given-names>J.-L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Strobelt</surname>
            , H.; and Laino,
            <given-names>T.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>Unsupervised attention-guided atommapping</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q. V.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Sequence to sequence learning with neural networks</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          ,
          <volume>3104</volume>
          -
          <fpage>3112</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kaiser</surname>
          </string-name>
          , Ł.; and
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          ,
          <volume>5998</volume>
          -
          <fpage>6008</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <surname>Vig</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Belinkov</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Analyzing the structure of attention in a transformer language model</article-title>
          . arXiv preprint arXiv:
          <year>1906</year>
          .04284.
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; and Kudlur,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Order matters: Sequence to sequence for sets</article-title>
          .
          <source>arXiv preprint arXiv:1511</source>
          .
          <fpage>06391</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Parulian</surname>
            , N.; Han,
            <given-names>G</given-names>
          </string-name>
          .; Ma, J.;
          <string-name>
            <surname>Tu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; Zhang, H.; Liu,
          <string-name>
            <surname>W.</surname>
          </string-name>
          ; et al.
          <year>2020</year>
          .
          <article-title>Covid-19 literature knowledge graph construction and drug repurposing report generation</article-title>
          . arXiv preprint arXiv:
          <year>2007</year>
          .00576.
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H.; Allot</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Leaman</surname>
          </string-name>
          , R.; and
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Pubtator central: automated concept annotation for biomedical full text articles</article-title>
          .
          <source>Nucleic acids research</source>
          <volume>47</volume>
          (
          <issue>W1</issue>
          ):
          <fpage>W587</fpage>
          -
          <lpage>W593</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ye</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ; Han,
          <string-name>
            <given-names>X.</given-names>
            ;
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ;
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ;
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ;
            <surname>Zhou</surname>
          </string-name>
          , J.; and Sun,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Docred: A large-scale document-level relation extraction dataset</article-title>
          . arXiv preprint arXiv:
          <year>1906</year>
          .06127.
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          <string-name>
            <surname>Zaki</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Meira</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Data mining and analysis: fundamental concepts and algorithms</article-title>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Liu,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Lai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Zhou</surname>
          </string-name>
          , G.; and
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          <article-title>Relation classification via convolutional deep neural network</article-title>
          .
          <source>In Proceedings of COLING</source>
          <year>2014</year>
          ,
          <source>the 25th International Conference on Computational Linguistics: Technical Papers</source>
          ,
          <fpage>2335</fpage>
          -
          <lpage>2344</lpage>
          . Dublin, Ireland: Dublin City University and Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>He</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Liu,
          <string-name>
            <given-names>K.</given-names>
            ; and
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Extracting relational facts by an end-to-end neural model with copy mechanism</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          ,
          <fpage>506</fpage>
          -
          <lpage>514</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , Y.;
          <string-name>
            <surname>Zhong</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Angeli, G.; and
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C. D.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Position-aware attention and supervised data improve slot filling</article-title>
          .
          <source>In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <fpage>35</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ; Han,
          <string-name>
            <surname>S</surname>
          </string-name>
          .
          <article-title>-</article-title>
          K.; and
          <string-name>
            <surname>So</surname>
            ,
            <given-names>I.-M.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Architecture of knowledge graph construction techniques</article-title>
          .
          <source>International Journal of Pure and Applied Mathematics</source>
          <volume>118</volume>
          (
          <issue>19</issue>
          ):
          <fpage>1869</fpage>
          -
          <lpage>1883</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          2017.
          <article-title>Joint extraction of entities and relations based on a novel tagging scheme</article-title>
          .
          <source>arXiv preprint arXiv:1706</source>
          .
          <fpage>05075</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          2016.
          <article-title>Attention-based bidirectional long short-term memory networks for relation classification</article-title>
          .
          <source>In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          ,
          <fpage>207</fpage>
          -
          <lpage>212</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>