<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Long Tailed Entity Extraction of Model Names using Distant Supervision</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Swayatta Daw</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vikram Pudi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data Sciences and Analytics Center IIIT Hyderabad</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <fpage>28</fpage>
      <lpage>38</lpage>
      <abstract>
        <p>We introduce the task of long-tailed detection of model entities from scientific documents. We use distant supervision using an external Knowledge Base (KB) to generate synthetic training data and use a simple entity replacement technique to improve performance significantly by addressing the problem of overfitting in small sized datasets for supervised NER baselines. We introduce strong baselines for this task which are evaluated on our annotated gold standard dataset. We also release the distantly supervised silver labels generated using the KB. We introduce this model as part of a starting point for an end-to-end automated framework to extract relevant model names and link them with their respective cited papers from research documents. We believe this task will serve as an important starting point to map the research landscape in a scalable manner, needing minimal human intervention.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Long-Tailed Entity</kwd>
        <kwd>Entity Extraction</kwd>
        <kwd>NER</kwd>
        <kwd>Information Extraction</kwd>
        <kwd>Scientific Literature</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Long tailed entities are named entities which rarely occur in text documents. For these types of
entities, the task of Named Entity Recognition (NER) is non-trivial. Recent approaches have
aimed at solving the problem of NER using supervised training using deep learning models.
However, supervised learning techniques require a large amount of token-level labelled data
for NER tasks. Annotating a large number of tokens can be time-consuming, expensive and
laborious. For real-life applications, the lack of labelled data has become a bottleneck on adopting
deep learning models to NER tasks.</p>
      <p>
        Most scientific named entities can be classified as long-tailed entities because of the rarity
and domain-specificity of their occurrence. Recent work on NER in scientific documents has
been concentrated around detecting biomedical named entities [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] or scientific entities like
tasks, methods and datasets [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
        ]. Some papers like [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] focus on the detection of a single
specific entity-type (like dataset names) from scientific documents. Although previous work
has focused on identifying methods [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] as named entities, but what constitutes a method can
have a significant variance when it comes to human annotated data. The authors [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] report the
Kappa score of 76.9% for inter-annotator agreement in the SciERC dataset, which is widely used
as a benchmark for scientific entity extraction.
      </p>
      <p>
        NER has traditionally been treated as a sequence labelling problem, using CRF [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and HMM
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Recent approaches have used deep learning-based models [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to address this task, which
require a large amount of labelled data to train. The high cost of labelling remains the main
challenge to train such models on rare long tailed entity types, where availability of labelled data
is scarce. In order to address the label scarcity problem, several methods like Active Learning
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Distant Supervision [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10, 11, 12</xref>
        ], Reinforcement Learning-based Distant Supervision [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]
have been proposed. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] focused on detecting dataset mentions from scientific text and used
data augmentation to overcome the label scarcity problem. In this paper, we leverage an external
Knowledge Base and a large scale unlabelled corpora for our distantly supervised approach,
using a simple entity replacement technique to prevent overfitting. In this paper, we introduce
the task of detection of model entity names from scientific documents. Papers with Code (PwC 1)
is a community driven corpus that serves to automatically list models that solve particular
subtasks, with links to the scientific research paper that introduced the model. Our aim is to
build a similar but automated end-to-end pipeline that detects model names from scientific
papers and benchmarks them against other similar models that solve the same task. We believe
the task introduced in this paper (extraction of model names from scientific documents) to be a
significant step forward towards the whole pipeline. This task is non-trivial mainly due to the
lack of availability of token-level high-quality labelled data which is required for training deep
learning models and the shortage of human annotated gold standard dataset for evaluation.
      </p>
      <p>To address the above bottlenecks, we present a simple yet efective technique leveraging an
external Knowledge Base and a large unlabelled corpora (both of which are cheap and easy
to obtain) to generate our training dataset. We believe this simple technique can be easily
extended to any other domain given the availability of a domain-specific Knowledge Base and
unlabelled text corpora. Utilising this training set, we are able to establish a strong baseline for
this task using a standard BERT-CRF model. In order to evaluate our performance for this task,
we present a high quality human annotated gold standard evaluation dataset.</p>
      <p>
        Using our trained models, we create an automated framework of detecting model names of
related work from research papers. We define related work as prior research work done by the
scientific community for the same or a similar related task that has been investigated by the
original paper. Our pipeline contains two steps: Firstly, we build a sentence intent classifier that
classifies whether a citation sentence contains information regarding related work or not. Then
we extract model names from the positively labelled sentences using our trained NER model
and link them to their respective citation mentions using a string distance based technique,
introduced by [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. We believe this framework is a starting point to efectively map the entire
research landscape in a scalable manner.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Annotation</title>
      <p>In order to create our whole set of gold labels for the evaluation test-set, we randomly sample
abstracts from a large set of arxiv Research papers2. We also introduce randomly sampled
papers from DBLP citation dataset to add to the diversity of train and test-set selection.
1https://github.com/paperswithcode/paperswithcode-data
2https://www.kaggle.com/Cornell-University/arxiv</p>
      <p>Our labeling model was built upon SciBERT
(Beltagy et al., 2019), a pre-trained language model
based on BERT (Devlin et al., 2019) but trained
on a large corpus of scientific text.</p>
      <p>There are two
models in the transformers, which can handle multilingual posts –
multilingual-BERT[19] and XLM-Roberta [57].</p>
      <p>For example, KG-BART encoded the graph structure of KGs with
knowledge embedding
algorithms like TransE (Bordes et al., 2013), and then took the
informative entity embeddings as
auxiliary input (Liu et al., 2021).</p>
      <p>Considering our end goal of automating a high precision framework of extracting related
model names and to minimise ambiguity, we consider only named models as model entities for
this task. Few examples are - BERT+BiLSTM+CRF, KG-BERT (with overlap), LSTM + Attention,
DeepWalk. We consider both single named entities and a combination of multiple model named
entities for annotation. A few example sentences with model entities are displayed in Figure 1.</p>
      <p>We consider strict span matches for the model entity names, and do not consider any partial
matches or synonyms. We also consider plural variants of entity names as matches.</p>
      <p>We aim to minimise ambiguity by considering only those model named entities that we can
verify about in Google Scholar and Semantic Scholar. We follow the process of identifying a
candidate model name and reviewing the existing Computer Science literature to verify whether
it is a model name entity or not by identifying its usage in the literature. A simple criteria that
we use is to observe if the model(or a variant of the model) has been mentioned in a Results
table and compared with baselines/other related models, in previous literature. Only after this
thorough review, we annotate a named entity as a model name. We discard any sentence if
a model named entity within the sentence does not follow the defining criteria. Hence, we
believe we reduce the ambiguity suficiently enough to allow for a single annotator for the
entire annotation process. All the annotations has been done by a graduate NLP researcher
who is also a co-author of this paper. The overall statistics of the training and test set has been
provided in Table 2.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Training Set Creation with Entity Replacement</title>
      <p>For the unlabelled corpus, we use the arxiv dataset containing 2˜27,000 abstracts from various
domains of Computer Science. We use the Papers with Code (PwC) corpus as a reference
Knowledge Base to obtain a total of 14,748 model entity names. We use this list of named
entities and create a set of distant silver labels by extracting the corresponding sentences out
of the arxiv dataset that contain the same entity mention. We aim for exact match while also
considering plural forms of the entity words. We obtain a total of 7800 sentences that contain a
model named entity mention.</p>
      <p>We plot a model entity vs frequency of occurrence in the entire corpus of our obtained
sentences. We provide the plot in Figure 2. We notice that the distribution is long-tailed in
nature, which is consistent with our hypothesis about scientific named entities as discussed in
the Introduction section. This means that there are certain pre-dominant popular models that
occur most frequently in the literature. The distribution tapers down and takes a long-tailed
form, where most of the entities have a much significant lower number of occurrence in the
literature. This can be attributed to the wide-spread use of certain models (like CNN), in the
existing Computer Science research literature.</p>
      <p>However, such a skewed distribution is unsuitable for training supervised custom NER models.
The models tend to memorise and overfit for certain named entities. Hence, we use a simple
entity replacement technique to deal with this bottleneck. More specifically, we detect the
entity span of the model named entity in the occurring sentence. Then, we replace this entity
with another entity from the entire set of model entities obtained from the Knowledge Base.
We execute the process keeping the number of entity distributions to atmost 2 to maintain
uniformity. After the entire process is completed, the entire train sentences set is set with
uniformly distributed entities.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Distantly Supervised NER Model</title>
      <p>
        • BiLSTM + CRF: This BiLSTM-CRF model captures the contextual representations and
encodes them into a bidirectional hidden state using BiLSTM. The CRF layer models
the dependency among a sequence of tokens by considering the entire sequence label
probability distribution.
• SciBERT + CRF: This model contains pre-trained SciBERT [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] embeddings trained on
large scientific corpus. The SciBERT-embeddings are passed onto a CRF layer that models
each sequence probability distribution.
• BERT+CRF: This consists of a pretrained BERT-model and a CRF layer to model
sequencelevel dependencies.
      </p>
      <p>We evaluate the models on the gold test labels. We find that SciBERT model in combination
with CRF provides the best performance. We also find that the entity replacement technique
is particularly efective when dealing with long tailed entity distributions. We find that this
simple technique ofers a significant boost in performance across all models. It is efective in
countering the overfitting bottleneck and successfully prevents memorisation of named entities.
We rely only on distant labels to obtain strong performance on gold labels. We illustrate our
best performing model in Figure 3.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Sentence Intent Classifier</title>
      <p>We train a classifier to detect whether a sentence contains relevant information regarding
models that solve a similar task as specified in the target research paper. For a target scientific
document, we define a relevant model name as a model that the author has cited, which solves
a task that is similar or relevant to the original task that the target paper is solving. To create an
0.4
0.9</p>
      <sec id="sec-5-1">
        <title>B-Model</title>
      </sec>
      <sec id="sec-5-2">
        <title>I-Model O</title>
        <sec id="sec-5-2-1">
          <title>I-Model O O</title>
        </sec>
        <sec id="sec-5-2-2">
          <title>O B-Model O O</title>
          <p>1.9
0.7
0.1</p>
          <p>CRF
0.9
1.3
0.4
Sci-BERT
0.4
0.1
1.9
0.7
0.6
1.6
0.8
0.5
1.3
We
present</p>
          <p>SDP</p>
          <p>
            LSTM
a
novel
network
automatically labelled dataset, we iterate over all sentences in the research corpora. If a sentence
contains the words - ‘Related Work’ or ‘Previous Work’ or ‘Baseline’, then we take 15 sentences
occurring after it. We assign positive labels to sentences containing model entity mentions by
referring to our KB. We consider the maximal span for entity matching between our unlabelled
text and KB. For creating negative samples, we randomly sample from all sentences and make
sure the above words are absent and it also does not contain model entity mentions. We keep
an equal distribution of positive and negative labels. An example of a positive and negative
label is shown in Figure 4.
5.1. Training the classifier
The most commonly used approach of averaging BERT embeddings or using the output of the
ifrst token (the [CLS] token) yields subpar sentence representations [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ]. Hence, we choose
Sentence-BERT [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ], a modification of the pretrained BERT network that uses siamese and
triplet network structures to derive semantically meaningful sentence embeddings that can
          </p>
          <p>The authors have introduced a probabilistic
framework based on Hidden Markov Random Fields
(HMRFs) for semi-supervised clustering that
combines the constraint-based and distance based
approaches in a unified framework.</p>
          <p>All processing units perform the same
computation, specified by equation (1), and are
locally connected to their three neighbours.
be compared using dot product. It takes a sentence as input and returns the corresponding
sentence-level representation as output. We use Sentence-BERT to encode the sentences and use
Logistic Regression as our binary classifier to train it on 15,518 labelled sentences, containing
both citation and non-citation sentences. Positive and negative samples are equally distributed.
The sentence dataset size is kept small to avoid compromising on the quality of the labels. The
train-test split followed is 75-25. The testset accuracy (which, again, consists of both citation
and non-citation sentences) is 86.41%.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Entity Citation Linker</title>
      <p>
        For the entity citation linking, we iterate between all possible extracted entities and citation
combination and get their closeness score, which is the string distance between an entity and
the citation occurrence. We first take all the citations and keep the closest entity per citation.
Then, we take all the entities and keep the closest citations per entity. This linking process is
able to accurately link most of the extracted entities with their closest citations, as demonstrated
by [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
    </sec>
    <sec id="sec-7">
      <title>7. Pipeline Formation</title>
      <p>We show two end-to-end pipelines in this paper. First, we show the entire training process for
both model entity extraction and sentence intent classification. We use the same unlabelled
corpora and Knowledge Base(KB) for both training processes. The automatic data labelling</p>
      <p>Sentences
The optimized 4-layer BiLSTM model was then calibrated
and validated for multiple prediction horizons.</p>
      <p>Furthermore, case studies show that SIMCLDA
can effectively predict candidate lncRNAs
for renal cancer.</p>
      <p>Longformer's attention mechanism is a
drop-in replacement for the standard self-attention.</p>
      <p>Bi-LSTM
SIMCLDA
Longformer</p>
      <p>MODEL
MODEL
MODEL
Unlabelled
corpora</p>
      <p>Sentences
Knowledge</p>
      <p>Base</p>
      <p>The authors have</p>
      <p>used...</p>
      <p>All processing
units...</p>
      <p>Entity Replacement
The optimized 4-layer BiLSTM model</p>
      <p>was then calibrated and ...</p>
      <p>The optimized 4-layer TransE model
was then calibrated and ...</p>
      <p>Weak Labels
Distantly labelled
TrDaiinstSanetnlytelnacbeeslled
TrDaiinstSanetnlytelnacbeeslled
Train Sentences
Sentence-BERT</p>
      <p>Classification Layer
Predicted
Labels</p>
      <p>The authors
have used ...</p>
      <p>All processing
units ...</p>
      <p>Positive(+)
Negative(-)</p>
      <p>B-Model I-Model O O O O</p>
      <p>CRF Layer
SciBERT Embeddings
Distantly Labelled</p>
      <p>Training Data
process using the external KB followed by entity replacement is shown in Figure 5. This
transformed dataset is used as distantly supervised training labels for input to the SciBERT-CRF
NER model. Also, the sentence intent classifier approach is illustrated, where the KB is utilised
to obtain weak binary labelled sentences from ’Related Work’ section to train the classifier.</p>
      <p>For the automated framework, we use a two stage pipeline. It takes a scientific research paper
as input and obtains the citation sentences from it. Using the trained sentence intent classifier,
we segregate the sentences into positive and negative labels. Only positively labelled sentences
are passed into the next stage of the pipeline. We use the trained SciBERT-CRF NER model to
extract entity mentions from the sentences. The model mentions are then linked with their
respective citations using our Entity-Citation linker. The framework is illustrated in Figure 6.</p>
      <p>Citation
Sentences</p>
      <p>
        The authors use CNN [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
layer on top of BERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
      </p>
      <p>
        embeddings
The ImdB dataset [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is
popular for sentiment
classification
      </p>
      <p>Trained Sentence Intent</p>
      <p>Classifier
(Sentence-BERT + Binary</p>
      <p>Classifier)
Predicted Labels
(green denotes +,
red denotes -)</p>
      <p>
        The authors use
CNN [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] layer on
top of BERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
embeddings
The ImdB dataset
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is popular for
sentiment
classification
      </p>
      <p>Entity-Citation</p>
      <p>
        Linker
The authors use CNN [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] layer on top of
      </p>
      <p>
        BERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] embeddings.
      </p>
      <p>
        B-Model
The authors use CNN [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] layer on
top of BERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] embeddings
SciBERT Embeddings +
      </p>
      <p>CRF
Trained NER Model
(Model Name Extractor)</p>
      <p>
        Predicted Entities ( with
citation link)
CNN [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
BERT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
      </p>
    </sec>
    <sec id="sec-8">
      <title>8. Error Analysis</title>
      <p>We conduct error analysis for model entity extraction, sentence intent classification and entity
citation linking. Some precision error is introduced into the model because for the training
set we consider the maximum span of each entity and the I-Model entity occurrence (a token
that lies inside a named entity) is high. We find in our evaluation dataset, the number of
B-Model entities is massively more, which leads to the model misclassifying an O as an I for
few sentences.</p>
      <p>Also, due to the usage of citation sentences in the evaluation dataset, our model recognises
the citation marker occurring right after the entity as an I-Model. Also, most of the citation
sentences in the evaluation dataset has a large number of named entities occurring adjacently,
as seen in many citation contexts. The model, which is trained on sentences from abstracts
only, is unable to recognise all of them as entities sometimes.</p>
      <p>For the sentence intent classification, our classifier often recognises sentences containing
dataset names as a positive label. This can be attributed to the fact that citation sentences that
refer to diferent datasets often have a similar structure to those citing model names of prior
work. Lastly, for the entity citation linker, sometimes an entity that is associated with a citation
marker occurs in the initial part of a sentence and its not the closest to the citation. This can
lead to missed out or incorrect linking.</p>
    </sec>
    <sec id="sec-9">
      <title>9. Implementation details</title>
      <p>We use PyTorch framework to implement our NER model. We use the pre-trained SciBERT
tokenizer and embeddings as input to a dropout layer with a dropout probability of 0.5 to prevent
overfitting. We use a learning rate of 1e-5 and train all models for 20 epochs. We pass the
output from the dropout layer through a linear layer with input dimension same as the hidden
dimension of SciBERT embeddings (768) and output dimension same as the number of labels
(4). We train the BiLSTM-CRF model for 20 epochs. We annotate the evaluation dataset in the
standard CoNLL BIO format. For Sentence-BERT, we use pretrained models available in Pytorch.
We use the DBLP corpus consisting of 4˜3K papers as our unlabelled research corpora to obtain
the distant labels for training the classifier. For the Knowledge Base (KB), we use PwC public
data corpus. For the CRF layer, we use allennlp3 models library. We use regular expressions to
extract citation sentences from papers that are written in the Springer LNCS/LNAI format.
10. Conclusion and future work
We have introduced a novel task of long-tailed model entity recognition from scientific
documents. We test our gold standard evaluation set on multiple baselines. We also find that a simple
strategy of entity replacement works well on small labelled datasets for distant supervision. We
hope to extend this technique to diferent types of entities with low labelled data availability.
We integrate our model in the automated pipeline framework to extract model names from
scientific research documents and link them to their respective citations. For future work, we
aim to utilise this pipeline on a large research corpus to obtain a map of benchmarked model
names linked with their respective papers on a much larger scale. We believe our work will
serve as an important starting point for mapping the entire research landscape of computer
science.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>V.</given-names>
            <surname>Kocaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Talby</surname>
          </string-name>
          ,
          <article-title>Biomedical named entity recognition at scale</article-title>
          , CoRR abs/
          <year>2011</year>
          .06315 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2011</year>
          .06315. arXiv:
          <year>2011</year>
          .06315.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ostendorf</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <article-title>Hajishirzi, Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction</article-title>
          ,
          <source>in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Brussels, Belgium,
          <year>2018</year>
          , pp.
          <fpage>3219</fpage>
          -
          <lpage>3232</lpage>
          . URL: https://aclanthology.org/D18-1360. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D18</fpage>
          -1360.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          , M. van Zuylen,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hajishirzi</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Beltagy</surname>
          </string-name>
          ,
          <article-title>Scirex: A challenge dataset for documentlevel information extraction</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2020</year>
          . arXiv:
          <year>2005</year>
          .00512.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mesbah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lofi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Torre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bozzon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.-J.</given-names>
            <surname>Houben</surname>
          </string-name>
          ,
          <article-title>Tse-ner: An iterative approach for long-tail entity extraction in scientific publications</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2018</year>
          , pp.
          <fpage>127</fpage>
          -
          <lpage>143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          , P. cheng Li,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Q.</surname>
          </string-name>
          <article-title>Cheng, Long-tail dataset entity recognition based on data augmentation</article-title>
          ,
          <source>in: EEKE@JCDL</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Laferty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pereira</surname>
          </string-name>
          ,
          <article-title>Conditional random fields: Probabilistic models for segmenting and labeling sequence data</article-title>
          , in: ICML,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H. L.</given-names>
            <surname>Chieu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ng</surname>
          </string-name>
          ,
          <article-title>Named entity recognition with a maximum entropy approach</article-title>
          , in: CoNLL,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sun</surname>
          </string-name>
          , J. Han,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>A survey on deep learning for named entity recognition</article-title>
          , ArXiv abs/
          <year>1812</year>
          .09449 (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Goldberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Grant</surname>
          </string-name>
          ,
          <article-title>A probabilistically integrated system for crowdassisted text labeling and extraction</article-title>
          ,
          <source>J. Data and Information Quality</source>
          <volume>8</volume>
          (
          <year>2017</year>
          ). URL: https://doi.org/10.1145/3012003. doi:
          <volume>10</volume>
          .1145/3012003.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Li</surname>
          </string-name>
          , J. Han,
          <article-title>Pattern-enhanced named entity recognition with distant supervision</article-title>
          ,
          <source>in: 2020 IEEE International Conference on Big Data (Big Data)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>818</fpage>
          -
          <lpage>827</lpage>
          . doi:
          <volume>10</volume>
          .1109/BigData50022.
          <year>2020</year>
          .
          <volume>9378052</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Er</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. Zhang,</surname>
          </string-name>
          <article-title>BOND: bert-assisted opendomain named entity recognition with distant supervision</article-title>
          , CoRR abs/
          <year>2006</year>
          .15509 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2006</year>
          .15509. arXiv:
          <year>2006</year>
          .15509.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hedderich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lange</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Klakow</surname>
          </string-name>
          ,
          <article-title>ANEA: distant supervision for low-resource named entity recognition</article-title>
          ,
          <source>CoRR abs/2102</source>
          .13129 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2102.13129. arXiv:
          <volume>2102</volume>
          .
          <fpage>13129</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Nooralahzadeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Lønning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Øvrelid</surname>
          </string-name>
          ,
          <article-title>Reinforcement-based denoising of distantly supervised NER with partial annotation</article-title>
          ,
          <source>in: Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo</source>
          <year>2019</year>
          ),
          <article-title>Association for Computational Linguistics</article-title>
          , Hong Kong, China,
          <year>2019</year>
          , pp.
          <fpage>225</fpage>
          -
          <lpage>233</lpage>
          . URL: https://aclanthology.org/D19-6125. doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D19</fpage>
          -6125.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Zhang,</surname>
          </string-name>
          <article-title>Distantly supervised NER with partial annotation learning and reinforcement learning</article-title>
          ,
          <source>in: Proceedings of the 27th International Conference on Computational Linguistics</source>
          , Association for Computational Linguistics, Santa Fe, New Mexico, USA,
          <year>2018</year>
          , pp.
          <fpage>2159</fpage>
          -
          <lpage>2169</lpage>
          . URL: https://aclanthology.org/C18-1183.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ganguly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Pudi</surname>
          </string-name>
          ,
          <article-title>Competing algorithm detection from research papers</article-title>
          ,
          <source>in: Proceedings of the 3rd IKDD Conference on Data Science</source>
          ,
          <year>2016</year>
          , CODS '16,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2016</year>
          . doi:
          <volume>10</volume>
          .1145/2888451.2888473.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cohan</surname>
          </string-name>
          ,
          <article-title>Scibert: A pretrained language model for scientific text</article-title>
          , arXiv preprint arXiv:
          <year>1903</year>
          .
          <volume>10676</volume>
          (
          <year>2019</year>
          ). URL: https://www.aclweb.org/anthology/D19-1371/.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          , Sentence-BERT:
          <article-title>Sentence embeddings using Siamese BERTnetworks</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          ,
          <article-title>Association for Computational Linguistics</article-title>
          , Hong Kong, China,
          <year>2019</year>
          , pp.
          <fpage>3982</fpage>
          -
          <lpage>3992</lpage>
          . URL: https://www.aclweb.org/anthology/D19-1410/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>