<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Simone Mellace , Vani K and Alessandro Antonucci⇤</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>IDSIA - Lugano (Switzerland)</string-name>
          <email>alessandro@idsia.ch</email>
          <email>simone@idsia.ch</email>
          <email>vanik@idsia.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>In: A. Jorge, R. Campos, A. Jatowt, A. Aizawa (eds.): Proceedings of the first AI4Narratives Workshop</institution>
          ,
          <addr-line>Yokohama</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>When coping with literary texts such as novels or short stories, the extraction of structured information in the form of a knowledge graph might be hindered by the huge number of possible relations between the entities corresponding to the characters in the novel and the consequent hurdles in gathering supervised information about them. Such issue is addressed here as an unsupervised task empowered by transformers: relational sentences in the original text are embedded (with SBERT) and clustered in order to merge together semantically similar relations. All the sentences in the same cluster are finally summarized (with BART) and a descriptive label extracted from the summary. Preliminary tests show that such clustering might successfully detect similar relations, and provide a valuable preprocessing for semi-supervised approaches.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Recent applications in the field of Natural Language
Processing (NLP) are exploiting data-driven techniques from the
general area of Machine Learning (ML). These are typically
Deep Learning (DL) systems based on multi-layer neural
networks fitted with the input text data, to be converted in
numerical objects by some embedding scheme. Such DL-NLP
systems are successful in extracting knowledge from natural
language and capturing the underlying narratives.</p>
      <p>As a matter of fact, most of these NLP efforts are focused
on a few mainstream applicative areas, such as biomedical
literature [Zhang et al., 2018; Lv et al., 2016] or news and
social media [Trieu et al., 2017; Ghosh and Shah, 2018]. Other
inputs such as literary text in the form of novels or short
stories received less attention [Wohlgenannt et al., 2016;
Volpetti et al., 2020]. This is unfortunate as literary texts
might exhibit high complexity in the narrative plots, while
also lacking explicit annotations, thus making the knowledge
extraction process very challenging. Handling such
complexities, helps in evaluating the models Natural Language
Understanding and creating benchmarks for these low-resource
domains. This could in turn be helpful for common sense
reasoning, reading comprehensions and enhance NLP
applications such as summary generation, machine translations and
question answering.</p>
      <p>Despite their astonishing applications in NLP, e.g., [Zhu et
al., 2019; Paulus et al., 2017], DL models are typically based
on discriminative functions with a huge number of
parameters, whose interpretation is often problematic. This prevents
both the explainability of the results and the possibility of
doing reasoning over the model entities. For this reason,
alternative approaches to NLP, based on so-called Knowledge
Graphs (KGs), i.e., relational ontologies providing
interlinked descriptions of the entities involved in a text, are also
popular. Despite the existence of techniques for automatic
KG extraction acting at the syntactic level [Tang et al., 2016;
Ruan et al., 2016], most of the approaches require
supervision in the form of manual annotations or access to
knowledge bases, such as UMLS1, for higher level descriptions.</p>
      <p>Of course these two orthogonal perspectives, say DL and
KGs, can be combined. DL models can be trained from KGs
[Socher et al., 2013; Li and Mao, 2019] and used for ML,
and, vice versa, DL models such as embeddings can be used
to predict missing links of the KG, classify relations, or align
entities from different KGs [Liu et al., 2019; Lin et al., 2015].</p>
      <p>
        Here, we follow such an integrated point of view, being
motivated by specific features of literary text understanding.
In fact, for this kind of text, the KG entities are typically the
characters in the plot, and no serious alignment issues
appear, while the classification of the relations becomes much
more challenging because of the lack of supervision. In other
domains the number of possible relations is typically limited
        <xref ref-type="bibr" rid="ref3">(e.g., in [Chen et al., 2010], few relations such as binding,
expression, protein interaction and few others)</xref>
        , while in the
literary case the possible relations between characters (e.g.,
Table 1) can be much more. Accordingly, we explore some
directions for an unsupervised approach to the identification
of relations in KGs obtained from literary texts. The goal is to
cluster semantically equivalent relational sentences including
      </p>
    </sec>
    <sec id="sec-2">
      <title>1https://uts.nlm.nih.gov</title>
      <p>descriptions of relations between the characters of a novel.
Our preliminary tests seem to be promising with respect to
the proper identification of similar relations, while also
giving directions about the most suitable clustering strategies as
well as further development of semi-supervised tools.</p>
      <p>The paper is organized as follows. In Section 2, we
summarize the existing literature in the field. Section 3 describes
our workflow, which is demonstrated by applicative examples
in Section 4. Conclusions and outlooks are in Section 5.
2</p>
      <sec id="sec-2-1">
        <title>Existing Work</title>
        <p>As discussed in the previous section, DL tools such as
sequence and self-attention models as well as transformers
(e.g., BERT, Xlnet, BART) have been widely and
successfully used in NLP for word and sentence encoding
[Peters et al., 2018; Devlin et al., 2018; Yang et al., 2019;
Lewis et al., 2019]. These models can be fine-tuned and used
for various tasks such as classification, summarization and
sentiment analysis. This also concerns KGs, where DL
models are used for embedding the triplet information and used
for tasks such as link predictions and KG completion [Lin et
al., 2019; Yao et al., 2019], while other researchers worked
on the training of embedding from KGs [Ji et al., 2015;
Wang et al., 2014; Lin et al., 2015; Bordes et al., 2013].</p>
        <p>None of these application was concerned with literary text.
Despite some attempts to apply ML and DL models in the
field [Worsham and Kalita, 2018; Short, 2019; Labatut and
Bost, 2019; K and Antonucci, 2019; Volpetti et al., 2020] to
analyze character relations, sentiments and visualizations, a
connection with KGs still remains under-explored. The goal
of this paper is to fill this gap by providing an unsupervised
alternative to the relational classifiers recently developed for
supervised tasks in [K et al., 2020].
3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Workflow</title>
        <p>Figure 1 depicts the workflow of the approach we propose
for the unsupervised identification of similar relations in the
KGs obtained from literary text. This involves a NLP part
for preprocessing (entity recognition, sentence tokenization,
detection of relational sentences and triplet generation)
corresponding to the red blocks, a DL abstraction level
(sentence embedding and summarization) corresponding to the
blue blocks, as well as classical ML techniques (characters
de-aliasing, sentence clustering, semi-supervised extension)
associated with green blocks. These steps in their sequential
order are described here below together with the main
challenges they present. The tool is available as a free software.2
Named Entity Recognition (NER). The very first step is
the identification of the entities to be associated with the KG
nodes. These are detected by a custom version of the Stanford
NER Tagger3 such that consecutive entities in a sentence (i.e.
words tagged as PERSON), are detected as a unique element
(e.g., Harry James Potter).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2https://github.com/IDSIA/novel2graph 3https://nlp.stanford.edu/software/CRF-NER.html</title>
      <p>Dealiasing. As a same character can be termed with
different aliases in the same novel, a de-aliasing might be
required. This issue has been already addressed in [K et al.,
2020], where a satisfactory solution based on ML and NLP
has been found. Here we adopt a similar strategy based on the
classical DBSCAN clustering (✏ = .3 and Levenshtein string
distances), together with a number of manual adjustments. In
our approach in fact, we first perform separate pre-clustering
over entities starting with the same letter (e.g., Hermione and
Hermione Granger are identified as a cluster while Harry,
Harry Potter and H. Potter as another one), and then adding
similar but unassigned names to a cluster (e.g., Granger
assigned to Hermione’s cluster and Potter to Harry). All the
occurrences of the aliases in the same cluster are finally
replaced by identifiers (e.g., CHAR0 replaces Harry, Potter,
Harry Potter and so on).</p>
      <p>Tokenization. Embeddings based on tranformers are based
on contextual information. Since, sentences are considered
as the simplest logical and meaningful unit that provides a
semantic intuition of the context, we rely on a segmentation
at this level.</p>
      <p>Relational Sentence Identification. Let us call relational
a sentence including two or more characters. We extract
relational sentences from the de-aliased and tokenized text, by
also evaluating whether or not the text between the two
character occurrences is a simple proposition or not (e.g., Harry
and Ron were having good time and Harry looked at Ron).
If this is the case we call the relation symmetric and we
generate two distinct input for the pipeline. Note also we only
use sentences containing exactly two characters and
excluding self-relations (e.g., Harry, I am Harry Potter).
Sentence Embedding. To identify the relations between
entities, we embed the relational sentences using Sentence
BERT (SBERT) [Reimers and Gurevych, 2019]. SBERT uses
a Siamese network structure [Schroff et al., 2015] to
reproduce meaningful encodings. The method was specifically
modelled for clustering and semantic search. SBERT adds
a pooling operation on top of BERT to derive these
embeddings. SBERT is fine tuned on SNLI [Bowman et al., 2015]
and MNLI [Williams et al., 2018] datasets with a three-way
soft-max classifier objective function for one epoch with the
default pooling strategy MEAN (computing the mean of all
output vectors).</p>
      <p>Sentence Clustering. Since, these embeddings encode
semantic and contextual information, sentences with similar
vector representations are supposed to share similar relations.
Hence, we adopt a simple clustering approach to group the
sentences with similar relations. The distances between the
vectors returned by SBERT are assumed to reflect the
semantic similarity between the corresponding sentences and hence
the relations included in these sentences. Classical clustering
methods such as k-means or DBSCAN can be therefore used
to create groups of sentences and hence triplets with the same
relation. We considered the Euclidean distance, as well as the
classical cosine distance. Even though clustering the entire
sentence may not explicitly cluster the relationships, the
sentences that fall into similar semantic spaces can provide us a
coarse-grained grouping of relations.</p>
      <p>CHAR0 and the
Philosopher’s Stone...
Cluster Summarization. After the clustering of the
relational sentences, we might want to represent these relations as
a summary of the sentences involved in the cluster. To achieve
that we adopt the BERT summarization pipeline based on the
BART [Lewis et al., 2019] model. This includes an encoder
like BERT and a decoder like GPT [Radford et al., 2019] and
it is trained on CNN/Daily Mail dataset with learning rate
3 · 10 5 (Adam optimizer). This performs extractive
summarization, giving most suitable representative sentences of each
cluster. Although training is not in-domain, as news articles
are also narratives, we use this for a preliminary set-up.
Triplet Generation. Once the extractive summary is
produced, for asymmetric relations we extract the phrase which
comes between the two reference characters. We then extract
only the verbs from these phrases, which are the considered
as part-of-speech tags that could convey some information
about the type of relations.</p>
      <p>From Unsupervised to Semi-Supervised Learning. The
overall procedure described in this section is purely
unsupervised. Yet, the clusters of relational sentences are described
by the summaries, first, and then labels generated by the
system. This might be the basis of a system where, part of those
clusters are manually inspected and their summaries/labels
validated or fixed by a human annotator. This would turn the
system into a semi-supervised one, where the annotated
clusters can be used as classifier of the relations.</p>
      <p>Book
HP
HP
LW
HP</p>
      <p>Sentence
Dumbledore smiled at the look of amazement on Henry’s face
Ron grinned at Henry
Brooke smiling at Meg as if everything had become possible him now
Henry stared as Dumbledore sidled back into the picture . . . gave him a small smile
For a first empirical validation of our pipeline we process, in a
single run, two novels, namely Harry Potter and the
Philosopher’s Stone (HP) by J. K. Rowling and Little Women (LW)
by Louisa May Alcott. 1307 suitable sentences out of 32365
are identified and grouped in 200 clusters (i.e. different
relations types). As the characters of the two books are distinct,
the system generates a KG with two disconnected
components (see Figure 2). Yet, the relation clustering is able to
detect similarities between sentences in the two books. E.g.,
sentences in Table 1 are related to smiling actions. For that
cluster the extractive summarization returns the first sentence
as a summary and, finally, the triplet generation mechanism
return smile as representative label.</p>
      <p>Concerning sentence clustering we considered both
DBSCAN and k-means algorithms both paired with Euclidean
and cosine distance. In the considered setup we did not found
significant differences with the two metrics. Regarding the
algorithms, an observed issue with DBSCAN was a sudden
transition from a huge number of single-sentence clusters to
very large clusters. Both these extreme scenarios prevent a
meaningful identification of relations. Yet, it was not possible
to automatically decide the number of clusters with k-means,
as the silhouette analysis returned monotone results.
Dumbledore</p>
      <p>Harry
Amy
Hagar</p>
      <p>Ron</p>
      <p>Snape</p>
      <p>Jo
Laurence</p>
      <p>Meg
Hermione</p>
      <p>Beth
Say Smile Look</p>
      <p>Others
An unsupervised approach to KG extraction from narrative
texts has been proposed. The procedure exploits transformer
models to detect similar relations in the triplets, then
generates summaries and representative labels for these clusters of
similar relations. This represent a pre-processing step for a
semi-supervised approach where the representative labels are
validated by human annotators and used as a relational
classifier. Validated clusters can define relational classifiers, while
the automatically generated labels are used for the others. As
a future work we want to apply our pipeline to a corpus of
literary texts and validate the clusters. This being a starting
point for the creation of a knowledge base for literary texts.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Bordes et al.,
          <year>2013</year>
          ]
          <string-name>
            <given-names>Antoine</given-names>
            <surname>Bordes</surname>
          </string-name>
          , Nicolas Usunier, Alberto Garcia-Duran,
          <string-name>
            <given-names>Jason</given-names>
            <surname>Weston</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Oksana</given-names>
            <surname>Yakhnenko</surname>
          </string-name>
          .
          <article-title>Translating embeddings for modeling multi-relational data</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>2787</fpage>
          -
          <lpage>2795</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Bowman et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Samuel</given-names>
            <surname>Bowman</surname>
          </string-name>
          , Gabor Angeli, Christopher Potts, and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>A large annotated corpus for learning natural language inference</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>632</fpage>
          -
          <lpage>642</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Chen et al.,
          <year>2010</year>
          ]
          <string-name>
            <given-names>Bin</given-names>
            <surname>Chen</surname>
          </string-name>
          , Xiao Dong, Dazhi Jiao, Huijun Wang, Qian Zhu, Ying Ding, and
          <string-name>
            <surname>David J Wild.</surname>
          </string-name>
          <article-title>Chem2Bio2RDF: a semantic framework for linking and data mining chemogenomic and systems chemical biology data</article-title>
          .
          <source>BMC bioinformatics</source>
          ,
          <volume>11</volume>
          (
          <issue>1</issue>
          ):
          <fpage>255</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Devlin et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          . BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>arXiv preprint arXiv:1810.04805</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Ghosh and Shah</source>
          , 2018]
          <string-name>
            <given-names>Souvick</given-names>
            <surname>Ghosh</surname>
          </string-name>
          and
          <string-name>
            <given-names>Chirag</given-names>
            <surname>Shah</surname>
          </string-name>
          .
          <article-title>Towards automatic fake news classification</article-title>
          .
          <source>Proceedings of the Association for Information Science and Technology</source>
          ,
          <volume>55</volume>
          (
          <issue>1</issue>
          ):
          <fpage>805</fpage>
          -
          <lpage>807</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Ji et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Guoliang</given-names>
            <surname>Ji</surname>
          </string-name>
          , Shizhu He, Liheng Xu, Kang Liu, and
          <string-name>
            <given-names>Jun</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <article-title>Knowledge graph embedding via dynamic mapping matrix</article-title>
          .
          <source>In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)</source>
          , pages
          <fpage>687</fpage>
          -
          <lpage>696</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[K and Antonucci</source>
          , 2019]
          <article-title>Vani K and Alessandro Antonucci. NOVEL2GRAPH: Visual summaries of narrative text enhanced by machine learning</article-title>
          .
          <source>In Text2Story@ ECIR</source>
          , pages
          <fpage>29</fpage>
          -
          <lpage>37</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>[K et al</article-title>
          .,
          <year>2020</year>
          ]
          <string-name>
            <surname>Vani</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simone Mellace</surname>
            , and
            <given-names>Alessandro</given-names>
          </string-name>
          <string-name>
            <surname>Antonucci</surname>
          </string-name>
          .
          <article-title>Temporal embeddings and transformer models for narrative text understanding</article-title>
          .
          <source>In Proceedings of Text2Story - Third Workshop on Narrative Extraction From Texts</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Labatut and Bost</source>
          , 2019]
          <string-name>
            <given-names>Vincent</given-names>
            <surname>Labatut</surname>
          </string-name>
          and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Bost</surname>
          </string-name>
          .
          <article-title>Extraction and analysis of fictional character networks: A survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>52</volume>
          (
          <issue>5</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>40</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>[Lewis</surname>
          </string-name>
          et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Mike</given-names>
            <surname>Lewis</surname>
          </string-name>
          , Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed,
          <string-name>
            <surname>Omer Levy</surname>
          </string-name>
          , Ves Stoyanov, and
          <string-name>
            <given-names>Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          . BART:
          <article-title>Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension</article-title>
          . arXiv preprint arXiv:
          <year>1910</year>
          .13461,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[Li and Mao</source>
          , 2019]
          <string-name>
            <given-names>Pengfei</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kezhi</given-names>
            <surname>Mao</surname>
          </string-name>
          .
          <article-title>Knowledgeoriented convolutional neural network for causal relation extraction from natural language texts</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <volume>115</volume>
          :
          <fpage>512</fpage>
          -
          <lpage>523</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>[Lin</surname>
          </string-name>
          et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Yankai</given-names>
            <surname>Lin</surname>
          </string-name>
          , Zhiyuan Liu, Maosong Sun, Yang Liu, and
          <string-name>
            <given-names>Xuan</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <article-title>Learning entity and relation embeddings for knowledge graph completion</article-title>
          .
          <source>In Twentyninth AAAI conference on artificial intelligence</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>[Lin</surname>
          </string-name>
          et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Bill</given-names>
            <surname>Yuchen Lin</surname>
          </string-name>
          , Xinyue Chen, Jamin Chen, and
          <string-name>
            <given-names>Xiang</given-names>
            <surname>Ren</surname>
          </string-name>
          .
          <article-title>Kagnet: Knowledge-aware graph networks for commonsense reasoning</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLPIJCNLP)</source>
          , pages
          <fpage>2822</fpage>
          -
          <lpage>2832</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Liu et al.,
          <year>2019</year>
          ] Weijie Liu, Peng Zhou,
          <string-name>
            <given-names>Zhe</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>Zhiruo</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Qi Ju, Haotang Deng, and
          <string-name>
            <given-names>Ping</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>K-bert: Enabling language representation with knowledge graph</article-title>
          .
          <source>In Proceedings of AAAI</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Lv et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Xinbo</given-names>
            <surname>Lv</surname>
          </string-name>
          , Yi Guan,
          <string-name>
            <given-names>Jinfeng</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jiawei</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Clinical relation extraction with deep learning</article-title>
          .
          <source>International Journal of Hybrid Information Technology</source>
          ,
          <volume>9</volume>
          (
          <issue>7</issue>
          ):
          <fpage>237</fpage>
          -
          <lpage>248</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Paulus et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Paulus</surname>
          </string-name>
          , Caiming Xiong, and
          <string-name>
            <given-names>Richard</given-names>
            <surname>Socher</surname>
          </string-name>
          .
          <article-title>A deep reinforced model for abstractive summarization</article-title>
          .
          <source>arXiv preprint arXiv:1705.04304</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Peters et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Peters</surname>
          </string-name>
          , Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (
          <issue>Long Papers)</issue>
          , pages
          <fpage>2227</fpage>
          -
          <lpage>2237</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [Radford et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Alec</given-names>
            <surname>Radford</surname>
          </string-name>
          , Jeffrey Wu, Rewon Child, David Luan,
          <string-name>
            <given-names>Dario</given-names>
            <surname>Amodei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          .
          <article-title>Language models are unsupervised multitask learners</article-title>
          .
          <source>OpenAI Blog</source>
          ,
          <volume>1</volume>
          (
          <issue>8</issue>
          ):
          <fpage>9</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>[Reimers and Gurevych</source>
          , 2019]
          <string-name>
            <given-names>Nils</given-names>
            <surname>Reimers</surname>
          </string-name>
          and
          <string-name>
            <given-names>Iryna</given-names>
            <surname>Gurevych</surname>
          </string-name>
          .
          <article-title>Sentence-BERT: Sentence embeddings using Siamese BERT-networks</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          , pages
          <fpage>3973</fpage>
          -
          <lpage>3983</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [Ruan et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Tong</given-names>
            <surname>Ruan</surname>
          </string-name>
          , Mengjie Wang, Jian Sun, Ting Wang, Lu Zeng, Yichao Yin, and
          <string-name>
            <given-names>Ju</given-names>
            <surname>Gao</surname>
          </string-name>
          .
          <article-title>An automatic approach for constructing a knowledge base of symptoms in chinese</article-title>
          .
          <source>In 2016 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)</source>
          , pages
          <fpage>1657</fpage>
          -
          <lpage>1662</lpage>
          . IEEE,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [Schroff et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Florian</given-names>
            <surname>Schroff</surname>
          </string-name>
          , Dmitry Kalenichenko, and
          <string-name>
            <given-names>James</given-names>
            <surname>Philbin</surname>
          </string-name>
          .
          <article-title>Facenet: A unified embedding for face recognition and clustering</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          , pages
          <fpage>815</fpage>
          -
          <lpage>823</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>[Short</source>
          , 2019]
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Short</surname>
          </string-name>
          .
          <article-title>Text mining and subject analysis for fiction; or, using machine learning and information extraction to assign subject headings to dime novels</article-title>
          .
          <source>Cataloging &amp; Classification Quarterly</source>
          ,
          <volume>57</volume>
          (
          <issue>5</issue>
          ):
          <fpage>315</fpage>
          -
          <lpage>336</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [Socher et al.,
          <year>2013</year>
          ] Richard Socher, Danqi Chen,
          <string-name>
            <surname>Christopher D Manning</surname>
            , and
            <given-names>Andrew</given-names>
          </string-name>
          <string-name>
            <surname>Ng</surname>
          </string-name>
          .
          <article-title>Reasoning with neural tensor networks for knowledge base completion</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>926</fpage>
          -
          <lpage>934</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Tang et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Zhiyuan</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Dong</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Zhiyong</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <article-title>Recurrent neural network training with dark knowledge transfer</article-title>
          .
          <source>In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pages
          <fpage>5900</fpage>
          -
          <lpage>5904</lpage>
          . IEEE,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [Trieu et al.,
          <year>2017</year>
          ]
          <article-title>Lap Q Trieu, Huy Q Tran, and MinhTriet Tran</article-title>
          .
          <article-title>News classification from social media using twitter-based doc2vec model and automatic query expansion</article-title>
          .
          <source>In Proceedings of the Eighth International Symposium on Information and Communication Technology</source>
          , pages
          <fpage>460</fpage>
          -
          <lpage>467</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [Volpetti et al.,
          <year>2020</year>
          ]
          <string-name>
            <given-names>Claudia</given-names>
            <surname>Volpetti</surname>
          </string-name>
          ,
          <string-name>
            <surname>Vani</surname>
            <given-names>K</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Antonucci</surname>
          </string-name>
          .
          <article-title>Temporal word embeddings for narrative understanding</article-title>
          .
          <source>In ICMLC 2020: Proceedings of the Twelfth International Conference on Machine Learning and Computing, ACM Press International Conference Proceedings Series. ACM</source>
          ,
          <year>2020</year>
          . ISBN:
          <fpage>978</fpage>
          -1-
          <fpage>4503</fpage>
          -7642-6.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>[Wang</surname>
          </string-name>
          et al.,
          <year>2014</year>
          ]
          <string-name>
            <given-names>Zhen</given-names>
            <surname>Wang</surname>
          </string-name>
          , Jianwen Zhang, Jianlin Feng, and
          <string-name>
            <given-names>Zheng</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Knowledge graph embedding by translating on hyperplanes</article-title>
          .
          <source>In AAAI</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>[Williams</surname>
          </string-name>
          et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Adina</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Nikita</given-names>
            <surname>Nangia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Samuel</given-names>
            <surname>Bowman</surname>
          </string-name>
          .
          <article-title>A broad-coverage challenge corpus for sentence understanding through inference</article-title>
          .
          <source>In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (
          <issue>Long Papers)</issue>
          , pages
          <fpage>1112</fpage>
          -
          <lpage>1122</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [Wohlgenannt et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Wohlgenannt</surname>
          </string-name>
          , Ekaterina Chernyak, and
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Ilvovsky</surname>
          </string-name>
          .
          <article-title>Extracting social networks from literary text with word embedding tools</article-title>
          .
          <source>In Proceedings of the Workshop on Language Technology Resources and Tools for Digital Humanities (LT4DH)</source>
          , pages
          <fpage>18</fpage>
          -
          <lpage>25</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <source>[Worsham and Kalita</source>
          , 2018]
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Worsham</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jugal</given-names>
            <surname>Kalita</surname>
          </string-name>
          .
          <article-title>Genre identification and the compositional effect of genre in literature</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Computational Linguistics</source>
          , pages
          <fpage>1963</fpage>
          -
          <lpage>1973</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [Yang et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Zhilin</given-names>
            <surname>Yang</surname>
          </string-name>
          , Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le.
          <article-title>Xlnet: Generalized autoregressive pretraining for language understanding</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>5754</fpage>
          -
          <lpage>5764</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [Yao et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Liang</given-names>
            <surname>Yao</surname>
          </string-name>
          , Chengsheng Mao, and
          <string-name>
            <given-names>Yuan</given-names>
            <surname>Luo</surname>
          </string-name>
          .
          <article-title>KG-BERT: BERT for knowledge graph completion</article-title>
          .
          <source>arXiv preprint arXiv:1909.03193</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [Zhang et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Yijia</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Hongfei Lin, Zhihao
          <string-name>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jian</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Shaowu Zhang, Yuanyuan Sun, and
          <string-name>
            <given-names>Liang</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>A hybrid model based on neural networks for biomedical relation extraction</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          ,
          <volume>81</volume>
          :
          <fpage>83</fpage>
          -
          <lpage>92</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [Zhu et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Xuelin</given-names>
            <surname>Zhu</surname>
          </string-name>
          , Biwei Cao, Shuai Xu, Bo Liu, and
          <string-name>
            <given-names>Jiuxin</given-names>
            <surname>Cao</surname>
          </string-name>
          .
          <article-title>Joint visual-textual sentiment analysis based on cross-modality attention mechanism</article-title>
          .
          <source>In International Conference on Multimedia Modeling</source>
          , pages
          <fpage>264</fpage>
          -
          <lpage>276</lpage>
          . Springer,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>