<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Improved
man, A. M. Walker, G. Eisenhauer, P. Widener, services and an expanding collection of metabolites,
A. Clif, Reusability first: Toward fair workflows, Nucleic Acids Res</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.18653/v1</article-id>
      <title-group>
        <article-title>Text-to-Ontology Mapping via Natural Language Processing Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Uladzislau Yorsh</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander S. Behr</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Norbert Kockmann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Holeňa</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Biochemical and Chemical Engineering, TU Dortmund University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Information Technology</institution>
          ,
          <addr-line>CTU, Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Computer Science, Czech Academy of Sciences</institution>
          ,
          <addr-line>Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Leibniz Institute for Catalysis</institution>
          ,
          <addr-line>Rostock</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>44</volume>
      <issue>2015</issue>
      <fpage>444</fpage>
      <lpage>455</lpage>
      <abstract>
        <p>The paper presents work in progress attempting to solve a text-to-ontology mapping problem. While ontologies are being created as formal specifications of shared conceptualizations of application domains, diferent users often create diferent ontologies to represent the same domain. For better reasoning about concepts in scientific papers, it is desired to pick the ontology which best matches concepts present in the input text. We have started to automatize this process and attack the problem by utilizing state-of-the-art NLP tools and neural networks. Given a specific set of ontologies, we experiment with diferent training pipelines for NLP machine learning models with the aim to construct representative embeddings for the text-to-ontology matching task. We assess the final result through visualizing the latent space and exploring the mappings between an input text and ontology classes.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;text analysis</kwd>
        <kwd>language models</kwd>
        <kwd>fastText</kwd>
        <kwd>BERT</kwd>
        <kwd>matching text to ontologies</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>ontologies can focus on diferent sub-domains as well as
on diferent levels of abstraction. Choosing the ontology
The FAIR (Findable, Accessible, Interoperable and which best corresponds to an input text is an important
Reusable) research data management needs a consistent step towards reasoning about it.
data representation in ontologies, particularly for repre- In the reported work in progress, we focus on the latter
senting the data structure in the specific domain [ 1]. The problem. One of the possible ways to address the task is
application of ontologies varies from a domain-specific to consider it as matching input texts with an existing
vocabulary and a translation reference up to an environ- text collection. Such a formulation allows to employ
ment for logical reasoning and property inference. already existing rich text processing pipelines, as well as</p>
      <p>Despite their purpose of standardizing the knowledge powerful pretrained models.
conceptualization, there still may exist several ontologies
within the same domain [2]. Creating and managing an
ontology is a manual process often performed by many 2. Related Work
domain experts. As each expert works on diferent
problems, they also might have diferent conceptualizations 2.1. Entity linking
of their respective knowledge. However, approaches to
automate the knowledge conceptualization also face their
challenges, as a machine cannot easily create semantics
without human input (e.g. scientific theses, which are
created by humans). A constant demand for a
knowledge database expansion and utilizing of already available
knowledge leads to the problem of ontology alignment
and merging, which is a research field on their own.</p>
      <p>Another problem faced by domain experts is how to
choose a proper ontology for a certain task. Diferent</p>
      <sec id="sec-1-1">
        <title>The problem is closely related to the concept normaliza</title>
        <p>
          tion and entity linking tasks. The algorithms encountered
in this context include dictionary lookup [
          <xref ref-type="bibr" rid="ref7">3, 4</xref>
          ],
conditional random fields and tf-idf vector similarity [ 5], word
embeddings and syntactical similarity [6].
        </p>
        <p>The vector similarity approaches either employ tf-idf
vectors or dense word embeddings. The tf-idf vector is
a document vector of the size of the considered
vocabulary, where each element is the number of occurrences of
the term in a document, multiplied by the logarithmized
ITAT’22: Information technologies – Applications and Theory, Septem- reciprocal value of the number of the documents where
ber 23–27, 2022, Zuberec, Slovakia this term appears. These vectors are well-interpretable
$ yorshula@fit.cvut.cz (U. Yorsh); (high values indicate the rare term which appears in
paranloerxbaenrdt.ekro.bcekhmr@antnu@-dtour-tdmourtnmdu.dned(.Ade. S(.NB.eKhorc);kmann); ticular document often), but very sparse, which impedes
martin@cs.cas.cz (M. Holeňa) the performance of machine learning algorithms. On
© 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License contrary, word embeddings generated by representation
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org)
learning algorithms are dense, but provide no direct in- is the key diference from many other works, which rely
terpretation. on ground-truth either for training or evaluation.</p>
        <p>The mentioned systems share a common pipeline—
at the first step, they use an external algorithm to find Ontologies may significantly difer in size. This
potential concepts in a scientific text. After that, they can lead to very outbalanced datasets when generating
link proposals with concepts using retrieval techniques, them from ontologies.
such as dictionary lookup or vector distance. These dificulties should be considered in the first place
when choosing a solution method.</p>
        <sec id="sec-1-1-1">
          <title>2.2. Natural Language Processing</title>
        </sec>
        <sec id="sec-1-1-2">
          <title>3.2. Text Similarity Strategy</title>
          <p>Entity linking techniques relying on vector similarity
may either use tf-idf vectors or word embeddings. The Ontologies typically provide annotations for most of their
latter may be beneficial due to the dense vector structure classes and relations, potentially generating supervised
and an ability to be produced by high-capacity language datasets for ML algorithms. But before employing a text
models, trained on large corpora. similarity approach, we have to make several strong
asfastText [7] is a representation learning algorithm pro- sumptions:
ducing word-level embeddings. A neural network with
a single hidden layer is being trained to predict a word • The distribution of input texts is the same as the
given its context, and the learned word representations distribution of annotation texts. It means that
are then being used as word embeddings. the input sentences should follow the same
gen</p>
          <p>Another widely used representation learning algo- eral structure, length and vocabulary as ontology
rithm is BERT [8]. A deep sequence processing neu- annotations to avoid prediction skewing for
irrelral network is trained on two objectives—predicting a evant reasons.
masked word in a sentence and predicting the order of • The best matching ontology is the one which
protwo given sentences. vides annotations most similar to the input text.</p>
          <p>Compared to the fastText, BERT embeds the whole Since the considered methods are text-based, they
input sequence at once and produces contextual embed- will not rely on structures or hierarchies created
dings for each token—the same token in diferent con- by ontology classes and input text terms.
texts will be embedded diferently. This allows it to
achieve state-of-the-art results in text classification [ 8] For methods mentioned below in this subsection, we
and named entity recognition [9] tasks. Another benefit will employ fastText and BERT models trained on texts
of BERT is that its Transformer architecture demonstrates from related domains, which will serve as a backbone for
impressive transfer-learning capabilities [10], which can further processing. Following the notation introduced
be useful for fine-tuning the model for tasks laying out- in the Subsection 3.1, we consider a "hard" mapping
side pretraining data distribution.  : T ↦→  directly to the space of ontologies of interest.
3.2.1. Zero-shot classification</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Matching Texts to Ontologies</title>
      <sec id="sec-2-1">
        <title>The method consists of assigning an ontology consid</title>
        <p>3.1. Problem definition ering a similarity between annotation embeddings and
an embedding of an input text. The method is simple
Within the proposed framework, we define an ontology and does not require model fine-tuning, which allows to
 as a directed attributed multi-graph, where vertices quickly establish a baseline for other experiments. The
represent classes, edges represent relationships between common choices of similarity measures are Euclidean
them, and both vertices and edges can have attributes. or cosine distances – we choose the latter in our
experiGiven a set of specific ontologies  = {1, . . . , } ments. The reason is that for some embedding algorithms
and an input text  ∈ T, the task is to predict the ontol- vector length may be influenced by the input text size,
ogy that best matches the content of  . A predictor may so vectors corresponding to semantically close texts may
be either a "hard" mapping  : T ↦→  or a scoring func- generally point in the same direction but be dissimilar in
tion  : T ×  ↦→ R which allows to order ontologies terms of Euclidean distance.
by relevance.</p>
        <p>There are several complications of the task:
Given ontologies are the only source of supervision.</p>
        <p>No text-to-ontology mapping labels are provided. This
3.2.2. Supervised classification based on ontology</p>
        <p>annotations
This method relies on a supervision provided by
ontology annotation attributes. Given an ontology set , we
can generate a dataset of annotation-ontology label pairs
and use it for supervised training. Under the
aforementioned assumptions we can directly assign input texts to
ontologies using the trained model.
3.2.3. Negative sampling
This method extends the method above by adding a
"None" class, denoting that the input text does not relate
to any of given ontologies. The annotation dataset is
extended by:
• Sentences extracted from scientific papers from
unrelated domains and labeled with the "None"
label.
• Sentences extracted from papers from related
domains with a diferent objective during training.</p>
        <p>For related input texts, instead of maximizing the
model output scores for a ground truth class we
minimize the output scores for the "None" class.</p>
        <p>This method is intended to partially counter the
possible input distribution diference between
ontology annotations and scientific texts.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Experiments</title>
      <p>4.1. Setup
a scispaCy [20] model en_ner_bc5cdr_md. For the
remaining machine learning models, PyTorch
implementations were used. For 3D visualization method Uniform
Manifold Approximation and Projection for Dimension
Reduction (UMAP) [21], we used the implementation
described in [22].</p>
      <p>Due to the lack of ground truth matching data, we
assess the performance primarily through inspecting the
resulting input sentence-annotation pairs.
4.1.1. Text preprocessing
We employ the following text preprocessing pipeline
before constructing input embeddings:
1. *Split an input text into sentences with a spaCy
model.
2. *Filter valid sentences, which contain at least two
nouns and a verb.
3. *Filter out sentences with non-paired parenthesis
and ill-parsed formulas or composed terms.
4. (BERT) Tokenize with a tokenizer coming with
the model.
4. (fastText) Convert to lowercase and split into
words</p>
      <sec id="sec-3-1">
        <title>The points marked with an asterisk are meant to be applied to new sentences from scientific papers only.</title>
      </sec>
      <sec id="sec-3-2">
        <title>We conduct our experiments on a set of five ontologies</title>
        <p>related to the chemical domain (Table 1). The ontolo- 4.2. Text Similarity
gies NCIT, CHMO and Allotrope are considered to be the Zero-shot setup. We start with representation
learnclosest to it, while Chemical Entities of Biological Inter- ing of annotations using the fastText and BERT
algoest (CHEBI) has only a subset of relevant entities. The rithms and inspecting the embeddings produced. For
SBO was selected as it contains some general laboratory the dimensionality reduction, we use the UMAP
algoand computational contexts, which can be seen as some rithm with the number of neighbors set to 15, minimum
kind of a test, whether the tools used can also identify distance 0.5 and cosine metric. We have found that
3ontologies not fitting to the text content. dimensional embeddings preserve substantially more
in</p>
        <p>We also selected 28 scientific papers as inputs for as- formation (allowing to separate clusters that may be
insessment, consisting of 25 research and 3 review papers. separable in 2D). The result is illustrated in Figure 1, three
Those papers deal with the topic of methanation of CO2 example sentences together with annotations assigned
and consist in sum of 1,3M symbols. to them by fastText and BERT are in Table 3.</p>
      </sec>
      <sec id="sec-3-3">
        <title>We use the pretrained fastText model by [16] and the</title>
        <p>recobo/chemical-bert-uncased [17] checkpoint of
a BERT implementation [18] from the HuggingFace
repository. For preprocessing we use spaCy [19] with</p>
        <p>Ontology matching as text classification. As we
mentioned in the Subsection 3.2, another potential
strategy to solve the problem is to treat it as a classification
task. If the distributions of input texts and
corresponding ontologies are the same, we can train a classifier on
ontology annotations and apply it on input texts.</p>
        <p>We implement this by embedding ontology
annotations with BERT and training over them a shallow
fullyconnected multilayer perceptron (MLP) with a single
768-dimensional hidden layer. Due to the significant
difference in sizes between ontologies, we proportionally
oversample minority data points. The classifier reaches
0.987 validation accuracy after the single-shot validation
on the annotations from all the classes, which indicates
their good separability for diferent ontologies, cf.
Figures 2 and 3.</p>
        <p>However, if we preprocess input texts and embed them
in this way, the inspection will show that their
distribution significantly difers from the distribution of ontology
annotations. The visualizations in Figures 2 and 3 show a
dense separate cluster of sentences parsed from scientific
papers.</p>
        <p>Negative sampling. As an attempt to counter the
issue, we introduced scientific texts into training data. We
sampled 400 scientific texts from the chemical domain (as
positive examples) and 400 from unrelated domains (as
negatives). During training, the model is being trained
on two objectives:
1. Cross-entropy loss if the input is an ontology
annotation (same as before)
2. Binary cross-entropy loss if the input is a sentence
from a scienticfi paper. The model minimizes the
probability of a special "Negative" class output
for a related scientific text, and maximises it for
unrelated.</p>
      </sec>
      <sec id="sec-3-4">
        <title>In this setting we train the head over BERT until con</title>
        <p>vergence first, leaving the backbone frozen. Considering
only ontology annotations and leaving aside sampled
sentences, the model reaches 0.984 validation accuracy,
which is very similar to the performance of the classifier
described above.</p>
        <p>After that, we fine-tune the whole BERT model. The
model reaches 0.958 validation accuracy after single-shot
validation on the combined annotation and paper
sentence dataset, with the confusion matrix on Figure 4. As
(a) The progress of training and validation accuracy
during training. The blue (above) and orange
(below) lines indicate the training and validation
accuracy respectively.
(a) 3-dimensional projection of the embeddings
produced by the fine-tuned BERT
(b) Annotation from ontologies and sentences from
14 scientific papers embedded by BERT
we will show later, mixing sampled sentences in from
both relevant and irrelevant scientific texts allowed to
improve classification accuracy over the classifier on top
of BERT.</p>
        <p>Despite the good separability of individual ontologies
and the additional optimization criterion, the UMAP
embeddings look similar to the previous setup in terms of
clustering input sentences into a separate subspace.</p>
        <p>It is worth to note that the classifier and negative
sampling models produce softmax scores, which can be
interpreted as a class probability distribution. However,
neural networks tend to be overconfident in their
outputs [23], so additional calibration is needed before using
the outputs for relevance estimation.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Statistical results. To compare the models, we con</title>
        <p>duct the Friedman test first to check if the models perform
the same. We perform a stratified split of the validation
dataset with the ontology annotations into 50 samples
(b) 3-dimensional projection of the activities of the
hidden layer of the MLP trained over BERT</p>
        <p>The Friedman test resulted in the null hypothesis
rejection on the significance level of 5%. To further compare
the models, we perform the Wilcoxon signed-rank test on
each pair of models. We make the following assumptions
about the algorithms:
• For a larger  the NN classifier can work the
same or better than the 1NN.
• The neural network model can fit training data
the same or better than the NN.
• The negative sampling results in a non-decrease
or an improvement in the model generalization.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Hypothesis H2 (Null for NN models): The 10NN mod</title>
        <p>els perform the same as their 1NN variants.</p>
      </sec>
      <sec id="sec-3-7">
        <title>While the 1NN is a common setting for many NLP</title>
        <p>systems, it may produce complex decision boundaries
and lead to overfitting. We test a larger  versus one to
determine whether this is an issue in our setup.</p>
      </sec>
      <sec id="sec-3-8">
        <title>Hypothesis H3 (Null for neural network classifier) : The</title>
        <p>NN classifier performs the same as the NN models both
on BERT/fastText embeddings.</p>
      </sec>
      <sec id="sec-3-9">
        <title>The assumption behind this hypothesis is that a neural network as a universal approximator can fit data better than a nearest-neighbour classifier.</title>
      </sec>
      <sec id="sec-3-10">
        <title>Hypothesis H4 (Null for the fine-tuned model with nega</title>
        <p>tive sampling): The fine-tuned BERT with negative
sampling performs the same as other considered models.</p>
      </sec>
      <sec id="sec-3-11">
        <title>We suppose that additional sampled sentences would allow to improve the model performance and help to avoid overfitting when fine-tuning the whole model instead of head only.</title>
      </sec>
      <sec id="sec-3-12">
        <title>Hypothesis H5 (Null for the rest): In each remaining</title>
        <p>pair, both models have the same performance.</p>
        <p>We indicate the relative model performance on
Figure 5. Considering the 5% significance level, the test
rejected all the null hypotheses except the H2, which was
rejected only for the fastText embeddings. To explain
that, we can note that there is a relatively sharp boundary
between individual classes on UMAP embeddings. If it
holds so for the original space, larger  may suppress
outlier noise but decrease classification accuracy near it.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion and Further</title>
    </sec>
    <sec id="sec-5">
      <title>Research</title>
      <sec id="sec-5-1">
        <title>We are not aware of other works on unsupervised textto-ontology mappings, so we are not able to discuss them and compare the proposed approach with previous methods.</title>
        <p>The reported work in progress revealed that the
distribution of the scientific texts substantially difers from
the one of ontology annotations. In spite of the high
classification accuracy both for the annotations from the
considered ontologies and the sentences of the additional
800 scientific papers, this leads to mapping into separate
subsets of the embedding space. This is true even for the
most sophisticated of the three investigated settings –
with the BERT fine-tuned using both the ontology
annotations and scientific texts from (un-)related domains.</p>
        <p>To avoid such a loss of generality, the future research
could include an intermediate step of entity recognition.
Using such recognized entities instead of raw text can
help to separate the information in scientific papers that
is directly related to concepts from ontologies and
unrelated words, sentences and other parts of text not
eliminated during preprocessing.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>nik 94</source>
          (
          <year>2022</year>
          )
          <fpage>852</fpage>
          -
          <lpage>863</lpage>
          . doi:https://doi.org/10.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          1002/cite.202100177. [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Hirschman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Valencia</surname>
          </string-name>
          , J. Fluck,
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Evaluation</given-names>
            <surname>Workshop</surname>
          </string-name>
          (
          <year>2007</year>
          )
          <fpage>149</fpage>
          -
          <lpage>151</lpage>
          . [4]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Morgan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Cohen</surname>
          </string-name>
          , J. Fluck,
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>P.</given-names>
            <surname>Ruch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Divoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Fundel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Leaman</surname>
          </string-name>
          , J. Haken- [17]
          <article-title>Bert for chemical industry</article-title>
          ,
          <year>2022</year>
          . URL: https://
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>W. W.</given-names>
            <surname>Lau</surname>
          </string-name>
          , H. Liu,
          <string-name>
            <given-names>C.-N.</given-names>
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schuemie</surname>
          </string-name>
          , K. B. [18]
          <string-name>
            <surname>Bert</surname>
          </string-name>
          ,
          <year>2022</year>
          . URL: https://huggingface.co/docs/
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>gene normalization</surname>
          </string-name>
          ,
          <source>Genome Biol 9 Suppl</source>
          <volume>2</volume>
          (
          <year>2008</year>
          ) [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Honnibal</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Montani</surname>
          </string-name>
          , spaCy 2: Natural language
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>S3. understanding with Bloom embeddings</article-title>
          , convolu[5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Leaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Islamaj</given-names>
            <surname>Dogan</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Lu,</surname>
          </string-name>
          <article-title>DNorm: disease tional neural networks and incremental parsing,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>name normalization with pairwise learning to rank</article-title>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>Bioinformatics</source>
          <volume>29</volume>
          (
          <year>2013</year>
          )
          <fpage>2909</fpage>
          -
          <lpage>2917</lpage>
          . [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Neumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          , W. Ammar, Scis[6]
          <string-name>
            <given-names>İ.</given-names>
            <surname>Karadeniz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Özgür</surname>
          </string-name>
          ,
          <article-title>Linking entities through paCy: Fast and robust models for biomedical natu-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <article-title>an ontology using word embeddings and syntactic ral language processing</article-title>
          ,
          <source>in: Proceedings of the 18th</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          re-ranking,
          <source>BMC Bioinformatics 20</source>
          (
          <year>2019</year>
          )
          <article-title>156</article-title>
          . URL: BioNLP Workshop and Shared Task, Association for
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          https://doi.org/10.1186/s12859-019-2678-8. doi:10.
          <string-name>
            <surname>Computational</surname>
            <given-names>Linguistics</given-names>
          </string-name>
          , Florence, Italy,
          <year>2019</year>
          , pp.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>1186/s12859-019-2678-8</source>
          .
          <fpage>319</fpage>
          -
          <lpage>327</lpage>
          . URL: https://aclanthology.org/W19-5034. [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Grave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , T. Mikolov, En- doi:10.18653/v1/
          <fpage>W19</fpage>
          -5034.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <article-title>riching word vectors with subword information</article-title>
          , [21]
          <string-name>
            <given-names>L.</given-names>
            <surname>McInnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Healy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Melville</surname>
          </string-name>
          , Umap: Uniform
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <source>arXiv preprint arXiv:1607.04606</source>
          (
          <year>2016</year>
          ).
          <article-title>manifold approximation and projection for dimen[8</article-title>
          ]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , sion reduction,
          <year>2018</year>
          . URL: https://arxiv.org/abs/
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <article-title>BERT: Pre-training of deep bidirectional transform-</article-title>
          <year>1802</year>
          .03426. doi:
          <volume>10</volume>
          .48550/ARXIV.
          <year>1802</year>
          .
          <volume>03426</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <article-title>ers for language understanding</article-title>
          , in: Proceedings [22]
          <string-name>
            <given-names>L.</given-names>
            <surname>McInnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Healy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Saul</surname>
          </string-name>
          , L. Grossberger, Umap:
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <article-title>of the 2019 Conference of the NAACL, Associ- Uniform manifold approximation and projection,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <article-title>ation for Computational Linguistics</article-title>
          , Minneapo-
          <source>The Journal of Open Source Software</source>
          <volume>3</volume>
          (
          <year>2018</year>
          )
          <fpage>861</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>lis</surname>
          </string-name>
          , Minnesota,
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          . URL: https: [23]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gal</surname>
          </string-name>
          , Uncertainty in Deep Learning,
          <source>Ph.D. thesis,</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          University of Cambridge,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>