<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparative analysis of context representation models in the relation extraction task from biomedical texts*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ilseyar Alimova</string-name>
          <email>alimovaIlseyar@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Tutubalina</string-name>
          <email>elvtutubalina@kpfu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kazan Federal University</institution>
          ,
          <addr-line>Kazan</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper focuses on the task of extracting relations between entities in biomedical texts. This study aims to identify the most effective method for representing context between entities. We compare several context representation methods such as a bag of words representation, average word embeddings, sentence embedding, representations obtained by convolutional, recurrent neural networks, and bidirectional encoder representations from Transformers (BERT). We conduct a set of experiments on two benchmark corpora of patient electronic health records and scientific articles in English. As expected, thehighestclassificationresultswereobtainedwiththestate-oftheart neural architecture BERT.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>* Copyright © 2019 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0International (CC BY 4.0).
disease relation (CDR) corpus [42] includes annotations of scientific articles on a biomedical domain. This study
examines the relationship between drugs and their attributes and between chemicals and diseases.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        There are various approaches to the problem of identifying related entities in biomedical texts [
        <xref ref-type="bibr" rid="ref15 ref18 ref24 ref31">3,5,15,18,24,28,
38</xref>
        ]. Earlier works on relation extraction adopted frequency-based methods. This method calculates the frequency
of the entities occurrences within the given context length. If the resulting number is greater than the specified
threshold, then it is considered that the entities are related. The advantage of this approach is the simplicity of its
implementation, no need for linguistic analysis and labeled sampling. However, the significant drawback of this
approach is that it does not take into account the semantic interpretation of context between entities presented in a
text.
      </p>
      <p>
        The template-based approach is grounded in finding a match for linguistic patterns, represented as regular
expressions. Templates are generated automatically or manually based on context. The advantage of this approach
is that there is no need for an annotated corpus. However, a wide variety of contests generates a large number of
templates, which significantly reduces the quality of the system [
        <xref ref-type="bibr" rid="ref12 ref28">8,12,35</xref>
        ].
      </p>
      <p>
        The increase of annotated corpora of biomedical text number leads to experiments with machine learning
methods to the problem of the relation extraction [
        <xref ref-type="bibr" rid="ref19 ref2 ref20 ref29 ref32">2,19,20,31,36,39</xref>
        ]. According to this approach, the context is
encoded as the feature vectors. The most common features are:
• bag of words: a feature vector that consists of the words before, after, and between entities;
• part of speech tags: a feature vector consisting of parts of speech words before, after and between entities;
• distance between the entities: the number of the words between the entities, the number of indicator
wordsbetween the entities, for example, specific verbs that indicate the existence of a connection;
• shortest syntactic tree path: the encoded shortest path from one entity to another in the syntactic parse tree.
      </p>
      <p>
        Recent relation extraction approaches are based on neural networks, where context and entities are encoded
with word embeddings as an input [
        <xref ref-type="bibr" rid="ref10 ref25 ref30">10, 25, 37</xref>
        ]. Sahu et al. applied CNN for extracting relations from patients’
electronic health records [
        <xref ref-type="bibr" rid="ref30">37</xref>
        ]. The model utilized as input the whole sentence encoded with word embeddings. The
obtained vectors sequentially passed through convolutional and dense layers. The results show that CNN can extract
global features, which can give good context representation and improve the quality of the system. Lv X. et al.
adopted autoencoder for context representation [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. The experiments indicate that the proposed model is effective,
and the method of optimizing functions by the deep learning model has great potential. Dandala et al. employed
bidirectional long short term memory network with attention for extracting relations from electronic health records
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The proposed approach achieved 84% of F-measure.
      </p>
      <p>A review of the literature shows that machine learning models are the most widely used method, and the most
common method for representing context is a bag of words. There are no studies that utilize sentence embeddings
to solve the problem of context representations.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Context Representation Methods</title>
      <p>
        Let context be text between two entities. For the evaluation, we select several approaches for context
representation, ranging from the simplest methods, such as a bag of words and an average vector representation of
words, to more complex, such as a vector representation of sentences, convolutional, and recurrent neural networks.
A classifier takes the context representation between two entities as input and predicts whether it express a relation.
Bag of words (bow) is one of the first models of text presentation, proposed by Zelling Harris in 1954 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
Currently, a bag of words is actively used for text classification and information retrieval. According to this model,
the number of occurrences of each word from the dictionary is calculated for the text, where the dictionary is a set
of unique words of all the texts of the training dataset. The model does not take into account the word order in the
text, which is one of its main disadvantages. Besides, the final text representation vector has a large dimension.
      </p>
      <p>
        The averaged word embeddings (word2vec) is calculated by summing the embedding of each word in the
context divided by the total number of words in the context. Tomas Mikolov proposed the word embedding model
in 2013 [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. It is based on a neural network trained to predict a word by context on a large corpus of text, the
hidden states of which are later used as vectors for words. The advantage of this model is the ability to consider the
semantic meaning of the words. Thus, the vectors of words that are close in meaning will be close to each other in
the vector space. However, this property can be lost on the text level due to averaging vectors. Also, this
representation has a fixed dimension for all texts, equal to the length of the word embedding vector.
      </p>
      <p>Sentence embeddings (sent2vec) are one of the variations of word embedding representation model [33].
However, the neural network trains not only on separate words but also on word n-grams and the averaged
embeddings for the words in a sentence. Thus, the model can better represent the semantic meaning of the sentence
than a simple averaging of word embeddings.</p>
      <p>
        Convolutional neural network (CNN) is widely used for context modeling [
        <xref ref-type="bibr" rid="ref17 ref21 ref22">17,21,22</xref>
        ]. The network takes as an
input a matrix E consisting of context words encoded with word embeddings. We apply a standard convolutional
layer over the matrix E. It is followed by a global max-pooling layer to produce the text embedding:
,
where k ∈ Rv×d is a kernel matrix, v is the width of a kernel; B ∈ R(n−v)×d is a matrix composed of elements bij. The j
axis is computed using different parallel kernels. The max operation is applied alongside the i axis.
      </p>
      <p>Thus, each neuron on the next layer is connected not with all neurons, but only with a small localized subset of
neurons in the previous layer. This fact allows for identifying the most significant features for each of the input
matrix fragments. The pooling layer is used to reduce the size of the feature map. Most often, the function of
maxpooling or weighted average pooling is used.</p>
      <p>
        Recurrent neural network (RNN) is used to process sequential data such as time series or word sequences
[
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. The network utilizes information from the previous network states, which is one of the critical advantages of
this model. The model takes context words encoded with word embeddings as an input. At each step, the network
calculates the weights using the word embedding vector and the output obtained at the previous step. We use the
last cell state as the context representation. ys = cn,
      </p>
      <p>hi = RNN(wi,ci−1),
where cn is the RNN memory state after reading the entire input sequence; hi is the RNN output produced using wi
(a word embedding) and ci−1 (memory state from the previous time step) as inputs.</p>
      <p>
        BERT (Bidirectional Encoder Representations from Transformers) is a recent neural network model for NLP
presented by Google [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The model obtained state-of-the-art results in various NLP tasks, including question
answering, dialog systems, text classification, and sentiment analysis. BERT neural network based on bidirectional
attention-based transformer architecture [
        <xref ref-type="bibr" rid="ref33">40</xref>
        ]. One of the main model advantages is the ability to give it a row text
as the input. In our experiments, we calculated the averaged vector of each word in the context. We utilize a
biomedical version of BERT called BioBERT [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>
        We conduct experiments on two annotated corpora of biomedical texts: MADE [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and CDR [42]. The overall
corpora statistic is presented in Table 1.
Indication and Adverse Drug Events from Electronic Health Record Notes (MADE) corpus consist of 1089
anonymized electronic health records of patients with cancer [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Electronic records include an extract statement,
inspectionresults, and othernotes. The corpus containsnine typesof entities andseven typesofrelations. Annotated
entities can be divided into two groups: related to the disease or the drug. Entities of the first group: adverse drug
reaction (ADE), a reason to use the drag (Indication), the severity of the disease (Severity), and other symptoms
and diseases not included in previous groups (SSD). Entities related to drugs: name (Drug name), dose (Dose),
duration of taking a drug (Duration), frequency of taking a drug (Frequency), route of taking a drug (Route).
      </p>
      <p>The corpus includes seven types of relationships, 4 of which are between the name of the drug and its attributes:
• Drug name – Dose
• Drug name – Route
• Drug name – Frequency
• Drug name – Duration
• Drug name - Indication
• Drug name - ADE
• SSD - Severity, includes the relationship between the severity of the disease and all types of entities
includedin the group of diseases: ADE, Indication, SSD.</p>
      <p>Entities in relations can be found both in one sentence and indifferent ones. The corpus is divided into training
and test subsets.
4.2</p>
      <sec id="sec-3-1">
        <title>CDR corpus</title>
        <p>The CDR corpus was developed for the BioCreative V competition [42]. The corpus consists of abstracts of
scientific articles collected from the PubMed resource. The corpus annotations contain the entities denoting
diseases (Disease) and chemical preparations (Chemical), and the relations between these entities. The corpus is
divided into three subsets: training, test, and development. In this work, the training and development subsets are
combined into one common train subset; the model is evaluated on a test subset.
4.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Generation of negative examples for training</title>
        <p>
          Manual annotations in both corpora contain only positive examples, denoting related entities. It is necessary to
generate negative examples to train models for binary classification. For each entity, we obtained a set of candidate
entities following the rules from [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]: the number of characters between the entities is smaller than 1000, and the
number of other entities that may participate in relations and locate between the candidate entities is not more than
3. These restrictions allow to reduce infrequent negative pairs and mitigate the imbalanced class issues, while more
than 97% of the positive pairs remain in the MADE dataset, and 100% remain in CDR corpus.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <p>We applied word vectors trained on the texts of PubMed and PMC resource articles and Wikipedia texts [29]
for the average vector of context representations. The length of the vectors is 200. The vocabulary coverage is 93%
for CDR and 89% for the MADE corpus. For sentence, embeddings were obtained from the BioSentVec model,
pre-trained on the text corpus consisting of articles from the PubMed resource and electronic patient cards of the
MIMIC-III base [7]. The model is trained on bigrams, with a window size of 20 words, the length of the resulting
vectors is 700.</p>
      <p>
        We utilized freezed weights from the last layer of BioBERT model [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. BioBERT * was initialized with
General-domain BERT and in addition pre-trained on PubMed abstracts (PubMed) and PubMed Central full-text
articles (PMC) (version: BioBERT v1.0 (+ PubMed 200K + PMC 270K)).
      </p>
      <p>
        Following [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], we trained convolutional neural network with the following parameters: the number of layers
is 3, the size of the layer filters are 5, 4, 3, the number of epochs is 10, the batch size is 32, the weights for classes
is 0.7 for related entities, and 0.3 for unrelated entities. A recurrent neural network was trained with the following
parameters: the number of hidden states is 200, the dropout is 0.2, the number of epochs is 20, the size of the input
data block is 64, the weight for the classes is 0.75 for entities that have a connection and 0.25 for unrelated entities.
All implementation is based on Keras and TensorFlow libraries [
        <xref ref-type="bibr" rid="ref9">1,9</xref>
        ].
      </p>
      <p>We employed a support vector machine (SVM) as a classifier. The classifier takes as an input various context
representationssequentially. Theclassifierwasevaluatedwithstandardmetrics: precision(P),recall(R),F-measure (F).
The results are presented in Table 2.</p>
      <p>BERT .929 .882 .905 .473 .385 .424</p>
      <p>According to the results, all models outperformed the baseline results of the bag of words model, which obtained
69.3% and 36.7% F-measures on MADE and CDR corpora, respectively. The best method of context representation
is BERT for both corpora. This model achieved 90.5% and 42.4% of F-measures on MADE and CDR corpora,
respectively. The averaged sent2vec method performed the second results. For the CDR corpus, the difference
between sent2vec and BERT models is 1.9, while on the MADE corpus, the difference is 2.2%, which is more
significant. CNN outperformed the RNN on MADE and CDR corpora on 33.2%, while the result for CDR corpus
state on par. The highest results in terms of precision and recall for CDR corpus was achieved by averaged word
embeddings method (55.7% of precision) and recurrent neural network (51.6% of recall). On the MADE corpus,
the highest results of precision and recall were achieved with the BERT model (92.9% and 88.2%, respectively).</p>
      <p>The results show that the F-measure on the MADE corpus is higher than on the CDR corpus in common. Such
a difference in results could be due to the MADE corpus has significantly more examples of relations, which allows
the classifier to learn the parameters better and make a better classification.</p>
      <p>* This model is available at https://github.com/naver/biobert-pretrained.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we have investigated several methods for representing the context in the task of extracting
relations between biomedical entities. The study aims to identify the most effective methods of context
representation. The experiment results showed that the BERT model performed the highest results. In the future,
we plan to evaluate models considered in the article for the protein-protein relation extraction task.
Acknowledgments This research was supported by the Russian Foundation for Basic Research grant no.
190701115.
[28] Makoto Miwa, Rune Saetre, Yusuke Miyao, and Jun’ichi Tsujii, A rich feature vector for protein-protein
interaction extraction from multiple corpora, Proceedings of the 2009 Conference on Empirical Methods in
Natural Language Processing: Volume 1-Volume 1, Association for Computational Linguistics, 2009, pp.
121–130.
[29] SPFGH Moen and Tapio Salakoski2 Sophia Ananiadou, Distributional semantics resources for biomedical
text processing, Proceedings of LBM (2013), 39–44.
[30] Tsendsuren Munkhdalai, Feifan Liu, and Hong Yu, Clinical relation extraction toward drug safety
surveillance using electronic health record narratives: classical learning versus deep learning, JMIR public
health and surveillance 4 (2018), no. 2, e29.
[31] Yun Niu, David Otasek, and Igor Jurisica, Evaluation of linguistic features useful in extraction of interactions
from pubmed; application to annotating known, high-throughput and predicted interactions in i2d,
Bioinformatics 26 (2009), no. 1, 111–119.
[32] Stanley Chika ONYE, Arif AKKELEŞ, and Nazife DIMILILER, Review of biomedical relation extraction.
[33] Matteo Pagliardini, Prakhar Gupta, and Martin Jaggi, Unsupervised learning of sentence embeddings using
compositional n-gram features, arXiv preprint arXiv:1703.02507 (2017).
[34] Deepak Ravichandran and Eduard Hovy, Learning surface text patterns for a question answering system,
Proceedings of the 40th annual meeting on association for computational linguistics, Association for
Computational Linguistics, 2002, pp. 41–47.
[42] Chih-Hsuan Wei, Yifan Peng, Robert Leaman, Allan Peter Davis, Carolyn J Mattingly, Jiao Li, Thomas C
Wiegers, and Zhiyong Lu, Overview of the biocreative v chemical disease relation (cdr) task, Proceedings of
the fifth BioCreative challenge evaluation workshop, vol. 14, 2015.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Martín</given-names>
            <surname>Abadi</surname>
          </string-name>
          , Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis,
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Matthieu</given-names>
            <surname>Devin</surname>
          </string-name>
          , Sanjay Ghemawat, GeoffreyIrving, MichaelIsard, etal.,
          <source>Tensorflow: Asystemforlarge-scalemachinelearning, 12th {USENIX} Symposium on Operating Systems Design and Implementation</source>
          ({
          <source>OSDI} 16)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>265</fpage>
          -
          <lpage>283</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Syed</given-names>
            <surname>Toufeeq</surname>
          </string-name>
          <string-name>
            <surname>Ahmed</surname>
          </string-name>
          , Radhika Nair,
          <string-name>
            <given-names>Chintan</given-names>
            <surname>Patel</surname>
          </string-name>
          , and Hasan Davulcu,
          <article-title>Bioeve: bio-molecular event extraction from text using semantic classification and dependency parsing</article-title>
          ,
          <source>Proceedings of the Workshop on Current Trends in Biomedical Natural Language Processing: Shared Task, Association for Computational Linguistics</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>99</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Antti</given-names>
            <surname>Airola</surname>
          </string-name>
          , Sampo Pyysalo, Jari Björne, Tapio Pahikkala, Filip Ginter, and Tapio Salakoski,
          <article-title>All-paths graph kernel for protein-protein interaction extraction with evaluation of cross-corpus learning</article-title>
          ,
          <source>BMC bioinformatics 9</source>
          (
          <year>2008</year>
          ), no.
          <issue>11</issue>
          ,
          <issue>S2</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Omar</given-names>
            <surname>Alonso</surname>
          </string-name>
          , Jannik Strötgen,
          <article-title>Ricardo A Baeza-Yates, and Michael Gertz, Temporal information retrieval: Challenges and opportunities</article-title>
          .,
          <source>Twaw</source>
          <volume>11</volume>
          (
          <year>2011</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>William A Baumgartner</surname>
            ,
            <given-names>K Bretonnel</given-names>
          </string-name>
          <string-name>
            <surname>Cohen</surname>
            , and
            <given-names>Lawrence</given-names>
          </string-name>
          <string-name>
            <surname>Hunter</surname>
          </string-name>
          ,
          <article-title>An open-source framework for largescale, flexible evaluation of biomedical text mining systems</article-title>
          ,
          <source>Journal of biomedical discovery and collaboration 3</source>
          (
          <year>2008</year>
          ), no.
          <issue>1</issue>
          ,
          <issue>1</issue>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Christian</given-names>
            <surname>Blaschke</surname>
          </string-name>
          ,
          <article-title>Miguel A Andrade, Christos A Ouzounis, and Alfonso Valencia, Automatic extraction of biological information from scientific text: protein-protein interactions</article-title>
          .,
          <source>Ismb</source>
          , vol.
          <volume>7</volume>
          ,
          <issue>1999</issue>
          , pp.
          <fpage>60</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Qingyu</given-names>
            <surname>Chen</surname>
          </string-name>
          , Yifan Peng, and Zhiyong Lu,
          <article-title>Biosentvec: creating sentence embeddings for biomedical texts</article-title>
          ,
          <source>The 7th IEEE International Conference on Healthcare Informatics</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Yong</given-names>
            <surname>Suk</surname>
          </string-name>
          <string-name>
            <surname>Choi</surname>
          </string-name>
          ,
          <article-title>Tree pattern expression for extracting information from syntactically parsed text corpora</article-title>
          ,
          <source>Data Mining and Knowledge Discovery</source>
          <volume>22</volume>
          (
          <year>2011</year>
          ), no.
          <issue>1-2</issue>
          ,
          <fpage>211</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>François</given-names>
            <surname>Chollet</surname>
          </string-name>
          et al.,
          <article-title>Keras: The python deep learning library</article-title>
          ,
          <source>Astrophysics Source Code Library</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Bharath</surname>
            <given-names>Dandala</given-names>
          </string-name>
          , Venkata Joopudi, and Murthy Devarakonda,
          <article-title>Adverse drug events detection in clinical notes by jointly modeling entities and relations using neural networks</article-title>
          ,
          <source>Drug safety 42</source>
          (
          <year>2019</year>
          ), no.
          <issue>1</issue>
          ,
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Jacob</surname>
            <given-names>Devlin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          , and Kristina Toutanova,
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <source>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Andrew D Fox</surname>
          </string-name>
          , William A Baumgartner, Helen L Johnson, Lawrence E Hunter, and
          <string-name>
            <surname>Donna</surname>
            <given-names>K Slonim</given-names>
          </string-name>
          ,
          <article-title>Mining protein-protein interactions from generifs with opendmap</article-title>
          ,
          <source>Linking Literature, Information, and Knowledge for Biology</source>
          , Springer,
          <year>2010</year>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Zellig</surname>
            <given-names>S Harris</given-names>
          </string-name>
          , Distributional structure,
          <source>Word</source>
          <volume>10</volume>
          (
          <year>1954</year>
          ), no.
          <issue>2-3</issue>
          ,
          <fpage>146</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>HochreiterandJ.Schmidhuber</surname>
          </string-name>
          , LongShort-TermMemory,
          <source>NeuralComputation9</source>
          (
          <year>1997</year>
          ), no.
          <issue>8</issue>
          ,
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          , Based on TR FKI-
          <volume>207</volume>
          -95, TUM (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Minlie</surname>
            <given-names>Huang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Xiaoyan</given-names>
            <surname>Zhu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ming</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>A hybrid method for relation extraction from biomedical literature</article-title>
          ,
          <source>International journal of medical informatics 75</source>
          (
          <year>2006</year>
          ), no.
          <issue>6</issue>
          ,
          <fpage>443</fpage>
          -
          <lpage>455</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Abhyuday</surname>
            <given-names>Jagannatha</given-names>
          </string-name>
          , Feifan Liu, Weisong Liu, and
          <article-title>Hong Yu, Overview of the first natural language processing challenge for extracting medication, indication, and adverse drug events from electronic health record notes (made 1.0), Drug safety (</article-title>
          <year>2018</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Nal</surname>
            <given-names>Kalchbrenner</given-names>
          </string-name>
          , Edward Grefenstette, and
          <string-name>
            <given-names>Phil</given-names>
            <surname>Blunsom</surname>
          </string-name>
          ,
          <article-title>A convolutional neural network for modelling sentences</article-title>
          ,
          <source>arXiv preprint arXiv:1404.2188</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Halil</given-names>
            <surname>Kilicoglu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sabine</given-names>
            <surname>Bergler</surname>
          </string-name>
          ,
          <article-title>Adapting a general semantic interpretation approach to biological event extraction</article-title>
          ,
          <source>Proceedings of the BioNLP Shared Task 2011 Workshop</source>
          , Association for Computational Linguistics,
          <year>2011</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>182</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Mi-Young</surname>
            <given-names>Kim</given-names>
          </string-name>
          ,
          <article-title>Detection of gene interactions based on syntactic relations</article-title>
          ,
          <source>BioMed Research International</source>
          <year>2008</year>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Sun</surname>
            <given-names>Kim</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soo-Yong</surname>
            <given-names>Shin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>In-Hee</surname>
            <given-names>Lee</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soo-Jin</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Ram Sriram, and
          <string-name>
            <surname>Byoung-Tak</surname>
            <given-names>Zhang</given-names>
          </string-name>
          ,
          <article-title>Pie: an online prediction system for protein-protein interactions from text</article-title>
          ,
          <source>Nucleic acids research</source>
          <volume>36</volume>
          (
          <year>2008</year>
          ), no.
          <source>suppl_2</source>
          ,
          <fpage>W411</fpage>
          -
          <lpage>W415</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Yoon</surname>
            <given-names>Kim</given-names>
          </string-name>
          ,
          <article-title>Convolutional neural networks for sentence classification</article-title>
          ,
          <source>Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1746</fpage>
          -
          <lpage>1751</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Yann</surname>
            <given-names>LeCun</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoshua Bengio</surname>
          </string-name>
          , et al.,
          <article-title>Convolutional networks for images, speech, and time series, The handbook of brain theory and neural networks 3361 (</article-title>
          <year>1995</year>
          ), no.
          <issue>10</issue>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Jinhyuk</given-names>
            <surname>Lee</surname>
          </string-name>
          , Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim,
          <article-title>Chan Ho So, and Jaewoo Kang, Biobert: pre-trained biomedical language representation model for biomedical text mining</article-title>
          , arXiv preprint arXiv:
          <year>1901</year>
          .
          <volume>08746</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Florian</given-names>
            <surname>Leitner</surname>
          </string-name>
          ,
          <article-title>Scott A Mardis, Martin Krallinger, Gianni Cesareni, Lynette A Hirschman, and Alfonso Valencia, An overview of biocreative ii. 5, IEEE/ACM Transactions on Computational Biology and Bioinformatics (TCBB) 7 (</article-title>
          <year>2010</year>
          ), no.
          <issue>3</issue>
          ,
          <fpage>385</fpage>
          -
          <lpage>399</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Xinbo</surname>
            <given-names>Lv</given-names>
          </string-name>
          , Yi Guan,
          <string-name>
            <given-names>Jinfeng</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <surname>Jiawei Wu</surname>
          </string-name>
          ,
          <article-title>Clinical relation extraction with deep learning</article-title>
          ,
          <source>International Journal of Hybrid Information Technology</source>
          <volume>9</volume>
          (
          <year>2016</year>
          ), no.
          <issue>7</issue>
          ,
          <fpage>237</fpage>
          -
          <lpage>248</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Larry</given-names>
            <surname>Medsker and Lakhmi C Jain</surname>
          </string-name>
          ,
          <article-title>Recurrent neural networks: design and applications</article-title>
          , CRC press,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>Dietrich</given-names>
            <surname>Rebholz-Schuhmann</surname>
          </string-name>
          , Antonio Jimeno-Yepes,
          <string-name>
            <given-names>Miguel</given-names>
            <surname>Arregui</surname>
          </string-name>
          , and Harald Kirsch, Measuring predictioncapacityofindividualverbsfortheidentificationofproteininteractions,
          <source>Journalofbiomedicalinformatics</source>
          <volume>43</volume>
          (
          <year>2010</year>
          ), no.
          <issue>2</issue>
          ,
          <fpage>200</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Rune</surname>
            <given-names>Saetre</given-names>
          </string-name>
          , Kenji Sagae, and
          <article-title>Jun'ichi Tsujii, Syntactic features for protein-protein interaction extraction</article-title>
          .,
          <source>LBM (Short Papers)</source>
          <volume>319</volume>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>Sunil</given-names>
            <surname>Kumar</surname>
          </string-name>
          <string-name>
            <surname>Sahu</surname>
          </string-name>
          , Ashish Anand, Krishnadev Oruganty, and Mahanandeeshwar Gattu,
          <article-title>Relation extraction from clinical texts using domain invariant convolutional neural network</article-title>
          ,
          <source>arXiv preprint arXiv:1606.09370</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>Isabel</given-names>
            <surname>Segura-Bedmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Paloma</given-names>
            <surname>Martínez</surname>
          </string-name>
          , and
          <string-name>
            <surname>César de Pablo-Sánchez</surname>
          </string-name>
          ,
          <article-title>A linguistic rule-based approach to extract drug-drug interactions from pharmacological documents, BMC bioinformatics</article-title>
          , vol.
          <volume>12</volume>
          ,
          <string-name>
            <surname>BioMed</surname>
            <given-names>Central</given-names>
          </string-name>
          ,
          <year>2011</year>
          , p.
          <fpage>S1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [39]
          <string-name>
            <surname>Sofie</surname>
            <given-names>Van Landeghem</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yvan Saeys</surname>
          </string-name>
          , Bernard De Baets, and Yves Van de Peer,
          <article-title>Extracting protein-protein interactions from text using rich feature vectors and feature selection</article-title>
          ,
          <source>3rd International symposium on Semantic Mining in Biomedicine (SMBM</source>
          <year>2008</year>
          ),
          <source>Turku Centre for Computer Sciences (TUCS)</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>77</fpage>
          -
          <lpage>84</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [40]
          <string-name>
            <surname>Ashish</surname>
            <given-names>Vaswani</given-names>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
          <article-title>Łukasz Kaiser, and Illia Polosukhin, Attention is all you need</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>Dmitriy</given-names>
            <surname>Yur'yevich Vlasov</surname>
          </string-name>
          ,
          <article-title>Dmitriy Yevgen'yevich Pal'chunov, and Pavel Andreyevich Stepanov, Avtomatizatsiya izvlecheniya otnosheniy mezhdu ponyatiyami iz tekstov yestestvennogo yazyka, Vestnik Novosibirskogo gosudarstvennogo universiteta</article-title>
          .
          <source>Seriya: Informatsionnyye tekhnologii 8</source>
          (
          <year>2010</year>
          ),
          <source>no. 3.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>