<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deep Bayesian Semi-Supervised Active Learning for Sequence Labelling</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tom´aˇs Sˇabata</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juraj Eduard P´all</string-name>
          <email>palljuraj1@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Holenˇa</string-name>
          <email>martin@cs.cas.cz</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Information Technology, Czech Technical University in Prague</institution>
          ,
          <addr-line>Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Mathematics and Physics, Charles University</institution>
          ,
          <addr-line>Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Computer Science of the Czech Academy of Sciences</institution>
          ,
          <addr-line>Prague, Czech republic</addr-line>
        </aff>
      </contrib-group>
      <fpage>80</fpage>
      <lpage>95</lpage>
      <abstract>
        <p>In recent years, deep learning has shown supreme results in many sequence labelling tasks, especially in natural language processing. However, it typically requires a large training data set compared with statistical approaches. In areas where collecting of unlabelled data is cheap but labelling expensive, active learning can bring considerable improvement. Sequence learning algorithms require a series of token-level labels for a whole sequence to be available during the training process. Annotators of sequences typically label easily predictable parts of the sequence although such parts could be labelled automatically instead. In this paper, we introduce a combination of active and semi-supervised learning for sequence labelling. Our approach utilizes an approximation of Bayesian inference for neural nets using Monte Carlo dropout. The approximation yields a measure of uncertainty that is needed in many active learning query strategies. We propose Monte Carlo token entropy and Monte Carlo N-best sequence entropy strategies. Furthermore, we use semi-supervised pseudo-labelling to reduce labelling effort. The approach was experimentally evaluated on multiple sequence labelling tasks. The proposed query strategies outperform other existing techniques for deep neural nets. Moreover, the semi-supervised learning reduced the labelling effort by almost 80% without any incorrectly labelled samples being inserted into the training data set.</p>
      </abstract>
      <kwd-group>
        <kwd>Active Learning</kwd>
        <kwd>Semi-supervised Learning</kwd>
        <kwd>Bayesian Inference</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Sequence Labelling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Deep learning is achieving state-of-the-art performance in image or video
processing, audio processing or natural language processing. However, without using
a pretrained model, deep learning typically requires a large amount of data. To
c 2019 for this paper by its authors. Use permitted under CC BY 4.0.
Se2mi-SupTeormvia´sˇsedSˇaAbacttai,vJeuLraejaErndiunagrdinPSa´lel,qaunedncMe aLrtaibneHlloinlegnˇa
obtain unlabelled input for deep networks in video processing, cameras and other
sensors are increasingly available. In natural language processing, a lot of
unlabelled inputs can be obtained for almost no cost by gathering them from web
sites. Unfortunately, labelling such data is very time consuming and expensive.</p>
      <p>In this situation, we can benefit from semi-supervised learning using a large
unlabelled dataset along with a small labelled one. Another option is to use
active learning wherein each iteration, a part of an annotation budget is spent
on labelling the most informative unlabelled samples. The model is retrained
including those new samples and the process repeats. The annotation budget is
significantly lower than the total number of available unlabelled samples.</p>
      <p>Although active learning is a promising way to benefit from unlabelled data,
the most common query strategy, uncertainty sampling, requires a measure of
uncertainty. In sequence labelling, the measure can be easily defined for
statistical models, such as hidden Markov models or conditional random fields (CRF),
as they provide a probability of the labelled sequence or a marginal probability
distribution for each element of the sequence. For neural networks, defining an
uncertainty measure is more complicated since the soft-max activation function,
typically used in the last network layer, does not correspond to a real
uncertainty of network predictions. To overcome this issue, one can use a Bayesian
neural network or include a statistical model, such as CRF, as the last layer of
the network.</p>
      <p>In sequence labelling, query strategies can be divided into two groups. The
first group computes the uncertainty of the sequence predicted by a model.
Query strategies of the second group compute uncertainties of separated tokens
and then aggregate them to express the uncertainty of the whole sequence.</p>
      <p>Querying the most informative sequence means that the annotator has to
label every token of the sequence. This is expensive and often not necessary
because some tokens can be very reliably annotated automatically. This situation
can be found in many natural language processing (NLP) tasks, where some
words can be assigned to only one category and we can predict that without
knowing the context. A similar situation can be found in a video where two
consecutive frames often contain the same or similar information and labelling
all frames might be inefficient.</p>
      <p>In this paper, we propose an active learning algorithm for sequence labelling
with deep neural networks that queries labels of the most informative tokens
whereas other labels are labelled automatically.</p>
      <p>In the following section, we summarize approaches addressing this topic. In
section 3, we define the architecture of our sequence labelling models. In section
4, we describe details of the proposed algorithm. The algorithm is evaluated
with experiments on tasks from natural language processing and the results are
shown in section 5.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Sequence labelling models have been used in many areas such as part of speech
tagging (POS) or named entity recognition (NER) [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], handwritten
recognition [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], protein secondary structure prediction [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], video analysis [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ] or facial
expression dynamic modeling [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In the early years, probabilistic models were
the most frequent approach. The most commonly used among them are Hidden
Markov models, dynamic Naive Bayesian classifiers, maximum entropy Markov
models or Conditional Random fields.
      </p>
      <p>
        With the increasing amount of data and computational power, and with
formulating new network topologies, deep networks are more and more popular in
sequence labelling. This is especially true for long short term memory networks
(LSTM), which deal well with vanishing gradient problem and are able to
incorporate context far from the predicted token. One of the state-of-the-art
topologies in sequence labelling is the bi-directional LSTM network (BI-LSTM) [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ]
or an extended version with a CRF layer on top (BI-LSTM-CRF) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
Another interesting topology specific to language processing uses an additional
layer (LSTM [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] or CNN [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]) as a character-level embedding for words.
      </p>
      <sec id="sec-2-1">
        <title>Active Learning in Sequence Labelling was studied intensively for proba</title>
        <p>
          bilistic models [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ]. Query strategies used in AL can be categorized into several
groups. Uncertainty Sampling that selects the most uncertain samples, Query by
Committee selects samples in which a committee disagree the most, Expected
Gradient Length selects samples that would conduct the greatest change to the
current model or Fisher information strategy that selects samples that
minimize the model variance. These strategies differ in computational complexity
and model requirements. The most commonly used strategy, uncertainty
sampling, requires the model to return confidence of its predictions. Furthermore, to
avoid querying samples that are rather outliers than representative samples, the
informativeness of the sample is weighted by its average similarity to all other
samples. The technique is called information density [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ].
        </p>
        <p>
          In active learning for sequence labelling, the most informative sequence is
labelled. The sequence is then added to the training set, the model is retrained
and the process repeats. This requires the whole sequence to be labelled at
once. In contrast, Tomanek [
          <xref ref-type="bibr" rid="ref40">40</xref>
          ] introduced the SeSAL algorithm, where parts of
sequences can be labelled automatically. That algorithm was designed for HMMs
and CRFs.
        </p>
        <p>
          Active Learning in Connection with Deep Learning Although active
learning has been applied to many ML tasks, application to deep learning is
marginal compared to probabilistic modelling. One of the main problems in deep
active learning is that many query strategies require some uncertainty estimate,
however, most kinds of deep neural networks rarely support it. In literature, we
can find several approaches approximating the model posterior: variational
inference [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], probabilistic back-propagation [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], Monte-Carlo (MC) dropout [
          <xref ref-type="bibr" rid="ref16 ref6">6, 16</xref>
          ]
Se4mi-SupTeormvia´sˇsedSˇaAbacttai,vJeuLraejaErndiunagrdinPSa´lel,qaunedncMeaLrtaibneHlloinlegnˇa
or mixture density networks [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. With such approximations of uncertainty, active
learning has been used in connection with deep learning in image classification [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]
or text classification [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. In the sequence labelling area, active deep learning was
successfully used for NER. In [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ], a CNN-CNN-LSTM network was used
together with active learning in a setup where the whole queried sequence had to
be labelled at once.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Underlying Models</title>
      <p>
        A sequence labelling model assigns categorical labels to all members (tokens)
of a sequence of observed values. In general, it considers the optimal label for
a given token to be dependent on the choices of nearby tokens. The problem is
often simplified through the assumption that the sequence of labels is a Markov
chain. With that simplification, the problem can be modelled with a
probabilistic graphical model such as a hidden Markov model [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] or a conditional random
field. Although the probabilistic models work well on many sequence labelling
tasks [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], the Markov property assumption might be too restrictive and
unrealistic for problems where a wider context is needed to label tokens correctly. This
can be overcome by considering dependencies of higher-order but the
computational complexity is growing exponentially with the order which makes these
models unusable for real-world problems. Deep learning neural networks can help
to overcome the issue of wider context.
      </p>
      <p>
        Deep Learning Models. In sequence labelling, various kinds of neural
networks are used. These networks are typically designed for a specific task. This is
particularly true for their first layers that extract features. In NLP, a character
level embedding layer extracts low-level features from the text. In video analysis,
feature vectors are extracted using pretrained convolutional networks. After the
first layer, a layer that incorporates contextual information from neighbouring
elements is plugged in. The most commonly used layers on this level are LSTM
cells [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] or gated recurrent unit (GRU) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Last, a sequence decoder layer is
used to predict the final sequence. Both context independent layers, for
example a fully connected dense layer (BI-LSTM-FCN) (Figure 1a), and contextual
layers, for example conditional random fields (BI-LSTM-CRF) (Figure 1b), can
be used.
      </p>
      <p>Moreover, to avoid over-fitting, a dropout regularization technique can be
used. In our experiments, we use dropout for non-recurrent connections (solid
lines in Figure 1). The dropout enabled for each layer allows to estimate
prediction uncertainties, as the following section describes.</p>
      <p>
        Bayesian neural networks aim to tackle several drawbacks of neural networks
such as overconfidence about their predictions or tendency to overfitting. In
classification, the prediction probabilities obtained from the soft-max function
are often erroneously interpreted as model confidence [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It means, the model
      </p>
      <p>Y2
Fully connected
neural network
o2(back) o2(for)</p>
      <p>LSTM
Backward</p>
      <p>LSTM
Forward
X2
Y2</p>
      <p>
        CRF2
(a) BI-LSTM with fully connected neural network on the top.
. . .
. . .
. . .
. . .
. . .
. . .
. . .
. . .
. . .
can be uncertain despite high values of the soft-max function and these values
require correct calibration [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] before using them as confidences. The main idea of
Bayesian neural networks is placing probabilistic distribution over nets’ weights
[
        <xref ref-type="bibr" rid="ref24 ref26">24, 26</xref>
        ]. However, the approach introduces two issues, intractable inferences and
computation costs. Although stochastic variational inference [
        <xref ref-type="bibr" rid="ref14 ref18 ref27 ref30">14,18,27,30</xref>
        ] solves
the problem with intractable inference, the number of parameters is doubled and
it requires more time to converge.
Se6mi-SupTeormvia´sˇsedSˇaAbacttai,vJeuLraejaErndiunagrdinPSa´lel,qaunedncMe aLrtaibneHlloinlegnˇa
      </p>
      <p>
        Gal &amp; Ghahramani introduced Monte Carlo dropout [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. They have shown
that dropout or various other stochastic regularization techniques can be used
to obtain an approximation of Bayesian inference. Consider a sequence of input
vectors denoted x to which a sequence of labels denoted y is assigned. A training
set containing pairs hx, yi is denoted T . Consider a neural net with parameters
ω that uses dropout at every layer for the training. Using dropout during testing
can be seen as sampling from a model’s approximate posterior. This leads to
approximate variational inference in which a tractable distribution qθ∗(ω)
minimizes the Kullback-Leibler (KL) divergence [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] to the true model posterior
p(ω|T ) given a training set T . The prediction uncertainty can be approximated
by marginalization over the approximate posterior using Monte Carlo
integration:
      </p>
      <p>Z
p(y = c|x, T ) =
p(y = c|x, ω)p(ω|T )dω
t=1</p>
      <p>
        R
≈ R1 X p(y = c|x, ωˆt),
where ωˆt ∼ qθ∗(w), R is the number of Monte Carlo runs, and where qθ(w)
denotes the Dropout distribution [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>Monte Carlo dropout does not affect the model training complexity, however,
each point has to be inferred repeatedly to obtain prediction uncertainty.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Active Learning Strategies</title>
      <p>
        Query strategies for sequence labelling models can be divided into several
frameworks such as uncertainty sampling (US), query by committee (QbC), expected
gradient length (EGL) or information density (ID) [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]. In this section, we
describe some of the strategies and propose how they can be used together with
the introduced models. The most informative samples are considered to be found
by maximising a particular utility function:
x∗ = arg max φ(x).
      </p>
      <p>x
The probability of the sequence y in a by model M given the input sequence
x is denoted PM (y|x). The set of labelled sequences is denoted L and set of
unlabelled sequences is denoted U .
4.1</p>
      <sec id="sec-4-1">
        <title>Query Strategies Utility Functions</title>
        <p>
          Query strategies of the uncertainty sampling framework select the sequences
that have the most uncertain label. The uncertainty measure can be expressed
in several ways. Least confidence query strategy [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] selects the sequence with
the lowest probability of the most likely sequence:
φLC (x) = 1 − PM (y∗|x),
where y∗ is the most likely sequence. For CRF, the most likely sequence and
its probability can be found using the Viterbi algorithm. For neural nets, the
probability of the most likely sequence can be approximated by an empirical
probability based on Monte Carlo dropout, which will be denoted PMMC(y|x). This
empirical distribution is calculated by counting the occurrences of the sequence
y for input sequence x in several forward passes through the network, where each
forward pass has a different dropout mask. These counts are normalized to sum
to 1.
        </p>
        <p>
          Margin query strategy [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] selects samples where the first and the second
most likely sequences have the most similar probabilities. Finding the second
most likely sequence in case of probabilistic graphical models requires an updated
version of the Viterbi algorithm called N-best Viterbi algorithm [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ]. For neural
nets, the distribution PMMC(y|x) can be used to find the probability of the second
most likely path.
        </p>
        <p>The margin query strategy utility function is defined as:</p>
        <p>φM (x) = − PM (y1∗|x) − PM (y2∗|x) ,
where y1∗ and y2∗ are first and second the most likely sequences.</p>
        <p>
          Token entropy query strategy [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ] uses the Shannon entropy of the model’s
posteriors:
        </p>
        <p>K
HM (l) = − X PM (yl = k) log PM (yl = k),</p>
        <p>k
where yl is label of the sequence in time l and K is the number of possible labels,
over its labellings to define the utility function for selecting the most uncertain
sequence:
where L is the length of the sequence. The utility function is normalized by the
length of the sequence. Omitting this normalization, the strategy would lead to
querying long sequences as they contain more information. The unnormalized
utility function is called total token entropy.</p>
        <p>
          Whereas the marginal probability for CRF can be calculated using forward
and backward scores, those scores are not available for neural networks. We
propose an approximation called Monte Carlo approximation token entropy,
which uses the idea of Bayesian inference with Monte Carlo [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]:
        </p>
        <p>PMMC(yl = k) log PMMC(yl = k)).</p>
        <p>
          Sequence entropy query strategy computes the entropy of probabilities of
all possible sequences. This strategy is unfeasible for long sequences as the
number of possible sequences grows exponentially with the length of the sequence.
Furthermore, it is not possible to obtain probabilities of a particular sequence
Se8mi-SupTeormvia´sˇsedSˇaAbacttai,vJeuLraejaErndiunagrdinPSa´lel,qaunedncMeaLrtaibneHlloinlegnˇa
directly in neural network based models. For probabilistic graphical models, the
strategy can be approximated with the N-best sequence entropy [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]:
where N = {y1∗, ..., yN∗ } is set of N most likely sequences found by N-best Viterbi
algorithm [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] and C1 normalizes the probabilities to sum to 1.
        </p>
        <p>While the probabilities of N most likely sequences can be obtained directly
in probabilistic models, this cannot be done in neural networks. Therefore, we
propose a Monte Carlo approximation of the sequence entropy:</p>
        <p>PMMC (yˆ|x)logPMMC (yˆ|x),
(1)
where N MC = {y1, y2, . . . } is set of all sequences predicted by Monte Carlo
sampling and C2 normalizes the probabilities to sum to 1.</p>
        <p>
          In the query by committee framework, a committee of models C = {M (1), ..., M (C)},
representing different hypotheses, is maintained during the whole process of
learning. The committee is used to query the sequence over which the
members are most in disagreement about how to label it. The committee is usually
trained using bagging. In each round, the labelling set is sampled with
replacement to create a unique training set L(C) that is used to train model M (C). The
committee prediction is obtained by models voting. In the context of deep neural
networks, maintaining a committee is too expensive for practical use. Although
dropout can be considered as a form of bagging [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], we do not deal with the
framework in the paper.
        </p>
        <p>US and QbC strategies are prone to querying outliers as they are often
uncertain for the model and the committee of models often disagrees about them.
The framework, called information density(ID), can be used to avoid this
problem. ID uses a base utility function φB(x) and weights it by samples‘
representativeness. All above defined utility functions can be used as base utility
functions. ID utility function is defined:
φID(x) = φB(x) ×
1 |U| β</p>
        <p>X sim(x, x(u)) ,
|U| u=1
(2)
where sim(x, x(u)) is a chosen similarity function for two sequences and β a
parameter that controls a relative importance of the representativeness term.
The similarity measure differs from task to task.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Token-Level Semi-Supervised Active Learning</title>
        <p>
          In the standard AL approach, the annotator has to label the whole sequence
although the sequence can contain subsequences that do not add too much value
to the utility function. If the model is sufficiently learned, these subsequences can
be easily annotated automatically using model inference. The decision whether
a token can be labelled automatically can rely on some kind of model confidence
[
          <xref ref-type="bibr" rid="ref40">40</xref>
          ] or the disagreement about the most probable paths [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ].
        </p>
        <p>We propose to use a combination of active and semi-supervised learning. For
models with CRF layer on the top, marginal probability represents the model
prediction confidence. Otherwise, the Monte Carlo dropout estimates the model
prediction confidence. First, the most informative sequence is found with a
chosen query strategy. Tokens in which the model is confident are automatically
labelled using semi-supervised learning and the rest is given to an annotator.
The labelled sequence is added to the training set and the process repeats.
Details of the approach are described in Algorithm 1. The confidence threshold has
to be chosen according to the model, problem type and query strategy.</p>
        <p>Algorithm 1: Sequential semi-supervised AL framework</p>
        <p>Input:
L: labelled set
U : unlabelled set
φ(·): query strategy utility function
θ: confidence threshold
M : model type
begin
train model m of type M on data set L
while stopping criterion is not met do
// Find the most informative sequence from U
x∗ = argmaxx∈U φ(x)
// label the sequence with the model or query the annotator
yˆ = m(x∗)
for i = 1 to length of x∗ do
if Pm(yi = yˆi|x∗) &gt; θ then</p>
        <p>yi∗ = yˆi
else
end</p>
        <p>yi∗ = query(xi∗)
end
L = L ∪ hx∗, y∗i
U = U \ x∗
retrain model m on L
end
end
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>To evaluate the performance of the proposed approach, we have chosen three
different sequence labelling problems: named entity recognition(NER), part of
Se1m0i-SupTeormvia´sˇsedSˇaAbacttai,vJeuLraejaErndiunagrdinPSa´lel,qaunedncMe aLrtaibneHlloinlegnˇa
speech tagging (POS) and chunking. Experiments were performed with two
sequential models: BI-LSTM-FCN with Monte Carlo dropout and BI-LSTM-CRF,
and various query strategies designed for each of those models.</p>
      <p>In the paper, we report two experiments. The first experiment tests proposed
query strategies against random sampling and least confident query strategies
as a baseline. The second experiment is using sequential semi-supervised active
learning framework to reduce the labelling effort. Our primary aim was reducing
the amount of labelled data required for training, rather than labelling
performance. Therefore, we did not extensively optimize hyper-parameters such as
learning rate, batch size or momentum.
5.1</p>
      <sec id="sec-5-1">
        <title>Experiment Design</title>
        <p>
          The experiments were performed on the publicly available benchmark dataset
CoNLL 2003 [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ]. The dataset provides a predefined training set and two testing
sets for POS, NER and Chunking. We report performance for the testing set A.
The training set was randomly divided into a labelled set and an unlabelled set
in the ratio 1:9.
        </p>
        <p>
          Both models use GloVe embeddings [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ] where each word is represented by
a vector of length 300. The models contain two LSTM hidden layers with size
100 and dropout with probability 0.4 applied to all layers. The last layer of the
BI-LSTM-FCN is linear and uses the soft-max activation function. The number
of forward passes for computing PMC was set to 500.
        </p>
        <p>First, each model was trained on the labelled training set with 10% of the
original size for 30 epochs. This model was used for experiments with all query
strategies. For each query strategy, the model was used to find the most likely
paths and their scores together with tokens prediction confidences for all
unlabelled sequences. The most informative sequences were selected and annotated
until the annotation budget was exhausted. We have defined the annotation
budget of one AL cycle in two ways: the number of labelled sequences and the total
number of annotated tokens. In the first scenario, 100 sequences were selected
and annotated, whereas, in the second scenario, sequences were annotated until
the total number of annotated tokens reached 1000. With the updated labelled
training set, the model was updated by iterative training for one epoch, then
the new score was calculated. This active learning cycle was repeated 20 times.
In the second experiment, samples were sorted according to their confidences,
and the threshold value was chosen to achieve 0% or 1% of incorrectly labelled
samples.</p>
        <p>
          Early results showed that proposed query strategies are prone to select
outliers. Therefore, the information density wrapping strategy was used for all of
them. Each sequence was represented by the average of embedding vectors. The
representativeness of the sequence was computed as an average cosine distance
to all other sequences in the unlabelled dataset. The cosine distance is claimed
to be an efficient similarity measure of the linguistic or semantic similarity of
corresponding words for the chosen embedding [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. In the results, we use the
names of base query strategies for clarity.
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Results</title>
        <p>In experiments, we studied the achieved performance in terms of F-measure
(specifically F1 score) and accuracy by particular query strategies and the
number of tokens that can be labelled automatically by semi-supervised learning.
We report the macro-averaged F1 score that is calculated:</p>
        <p>F 1macro =
1</p>
        <p>X F 1q,
|Q| q∈Q
where Q is the set of all possible labels and F 1q is the F1 score for the class
labeled q considered as the positive class and all remaining classes as the negative
class.</p>
        <p>These scores were compared with the models learned on the whole labelled
dataset and models learned on a small labelled dataset that was later used for
active learning.</p>
        <p>Query Strategies Comparison Query strategies were compared in two AL
scenarios: an unlimited number of tokens and a limited number of tokens. Table
1a shows that the strategies MC total token entropy and MC sequence entropy
outperforms other strategies in both F1 score and accuracy. The MC total token
entropy, however, required more tokens to be labelled. In NER, it queried almost
twice as many tokens. In the scenario with a limited number of tokens, the MC
sequence entropy dominates over other strategies except in the Chunking.</p>
        <p>Table 1b shows that for the BI-LSTM-CRF model, the least confident and
total token entropy query strategies have shown better results compared to the
token entropy query strategy. We conclude that total token entropy query
strategy dominates in the scenario with an unlimited number of tokens, whereas
least confident achieves better results in the scenario with a limited number of
tokens. The sequence entropy query strategy is missing as our implementation
was lacking n-best Viterbi algorithm.</p>
        <p>Moreover, Table 1b shows that the MC sequence entropy query strategy is
the best among the compared strategies in the NER and POS tasks during the
whole AL loop if the number of annotated tokens is limited in each cycle of the
AL loop (Figures 2a and 2b).</p>
      </sec>
      <sec id="sec-5-3">
        <title>Active Learning in Combination with Semi-supervised Learning Last,</title>
        <p>we studied a possible reduction of the labelling effort using semi-supervised
learning. We report how many tokens were automatically labelled if the threshold is
set to not allow errors propagate into the training dataset and if 1% of errors
are allowed. The results in Table 2 indicate that the BI-LSTM-CRF model has
a more reliable uncertainty measure for the marginal distribution than the
BILSTM-FCN model. It can reduce the labelling effort up to almost 80% without
any incorrectly labelled samples being inserted into the training data set. The
labelling effort is reduced up to almost 84% with 1% of incorrectly labelled samples
being inserted into the training dataset.
Se1m2i-SupTeormvia´sˇsedSˇaAbacttai,vJeuLraejaErndiunagrdinPSa´lel,qaunedncMe aLrtaibneHlloinlegnˇa
MC Sequence Entropy
Random
MC Total Token Entropy
Least Confident
MC Token Entropy</p>
        <p>NER
0
2500
5000
7500
0.775
0.770
0.765
1
F0.760
0.755
0.750</p>
        <p>Anno1ta0t0e0d0tokens12500 15000 17500 20000</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6 Conclusions and Future Work</title>
      <p>In this paper, we presented an application of Monte Carlo dropout, an
approximation of Bayesian inference for deep neural networks, to active learning
strateSe1m4i-SupTeormvia´sˇsedSˇaAbacttai,vJeuLraejaErndiunagrdinPSa´lel,qaunedncMe aLrtaibneHlloinlegnˇa
gies developed for probabilistic graphical models. We proposed two not yet used
adaptations of token entropy and sequence entropy query strategies suitable for
LSTM-type deep neural networks. Moreover, we tested a combination of active
and semi-supervised learning for sequence labelling for that network.</p>
      <p>The proposed query strategies have shown a substantial improvement over
the until now used strategy in sequence labelling with deep neural networks,
least confident. The proposed strategies outperformed the least confident in all
three considered sequence labelling tasks in case of the network without a CRF
layer. This is particularly true, if the annotation budget is limited for each active
learning batch, which is a typical real-world situation.</p>
      <p>The combination of active and semi-supervised learning allows us to achieve
up to 80% labelling cost reduction for the BI-LSTM-CRF model. The uncertainty
measure based on Monte Carlo dropout, however, still needs improvement to
achieve labelling effort reduction comparable with BI-LSTM-CRF. To this end,
we would like to study uncertainty measures provided by other approaches to
Bayesian recurrent neural networks.</p>
      <p>Although uncertainty sampling has shown to be applicable to deep neural
networks, other active learning frameworks have not been enough studied. In
the future, we would like to study, in the context of sequence labelling and deep
neural networks, active learning based on expected gradient length. In addition
to this, we would like to apply deep active learning to sequence labelling in video
processing, where context is also very important information.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>The work has been supported by the grant 18-18080S of the Czech Science
Foundation (GACR).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Burkhardt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siekiera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kramer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Semi-supervised bayesian active learning for text classification</article-title>
          .
          <source>In: Bayesian Deep Learning</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , Van Merri¨enboer,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Bahdanau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          :
          <article-title>On the properties of neural machine translation: Encoder-decoder approaches</article-title>
          .
          <source>arXiv preprint arXiv:1409.1259</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Uncertainty-aware learning from demonstration using mixture density networks with sampling-free variance modeling</article-title>
          .
          <source>In: 2018 IEEE International Conference on Robotics and Automation (ICRA)</source>
          . pp.
          <fpage>6915</fpage>
          -
          <lpage>6922</lpage>
          . IEEE (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sebe</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garg</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>T.S.</given-names>
          </string-name>
          :
          <article-title>Facial expression recognition from video sequences: temporal and static modeling</article-title>
          .
          <source>Computer Vision and image understanding 91(1-2)</source>
          ,
          <fpage>160</fpage>
          -
          <lpage>187</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Culotta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Reducing labeling effort for structured prediction tasks</article-title>
          .
          <source>In: AAAI</source>
          . vol.
          <volume>5</volume>
          , pp.
          <fpage>746</fpage>
          -
          <lpage>751</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gal</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghahramani</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Dropout as a bayesian approximation: Representing model uncertainty in deep learning</article-title>
          .
          <source>In: international conference on machine learning</source>
          . pp.
          <fpage>1050</fpage>
          -
          <lpage>1059</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gal</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Islam</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghahramani</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Deep bayesian active learning with image data</article-title>
          .
          <source>In: Proceedings of the 34th International Conference on Machine Learning-Volume 70</source>
          . pp.
          <fpage>1183</fpage>
          -
          <lpage>1192</lpage>
          . JMLR. org (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Graves</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Practical variational inference for neural networks</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <fpage>2348</fpage>
          -
          <lpage>2356</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Graves</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Offline handwriting recognition with multidimensional recurrent neural networks</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <fpage>545</fpage>
          -
          <lpage>552</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pleiss</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weinberger</surname>
            ,
            <given-names>K.Q.</given-names>
          </string-name>
          :
          <article-title>On calibration of modern neural networks</article-title>
          .
          <source>In: Proceedings of the 34th International Conference on Machine Learning-Volume 70</source>
          . pp.
          <fpage>1321</fpage>
          -
          <lpage>1330</lpage>
          . JMLR. org (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hern</surname>
          </string-name>
          <article-title>´andez-</article-title>
          <string-name>
            <surname>Lobato</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adams</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Probabilistic backpropagation for scalable learning of bayesian neural networks</article-title>
          .
          <source>In: International Conference on Machine Learning</source>
          . pp.
          <fpage>1861</fpage>
          -
          <lpage>1869</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakhutdinov</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          :
          <article-title>Improving neural networks by preventing co-adaptation of feature detectors</article-title>
          .
          <source>arXiv preprint arXiv:1207.0580</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9(8)</source>
          ,
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hoffman</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paisley</surname>
          </string-name>
          , J.:
          <article-title>Stochastic variational inference</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          <volume>14</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1303</fpage>
          -
          <lpage>1347</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bidirectional lstm-crf models for sequence tagging</article-title>
          .
          <source>arXiv preprint arXiv:1508</source>
          .
          <year>01991</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Kendall</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gal</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>What uncertainties do we need in bayesian deep learning for computer vision</article-title>
          ? In
          <source>: Advances in neural information processing systems</source>
          . pp.
          <fpage>5574</fpage>
          -
          <lpage>5584</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cha</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>G.G.</given-names>
          </string-name>
          :
          <article-title>Mmr-based active machine learning for bio named entity recognition</article-title>
          .
          <source>In: Proceedings of the Human Language Technology Conference of the NAACL, Companion Volume: Short Papers</source>
          . pp.
          <fpage>69</fpage>
          -
          <lpage>72</lpage>
          . Association for Computational Linguistics (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welling</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Auto-encoding variational bayes</article-title>
          .
          <source>arXiv preprint arXiv:1312.6114</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Krogh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larsson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Von Heijne</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sonnhammer</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          :
          <article-title>Predicting transmembrane protein topology with a hidden markov model: application to complete genomes</article-title>
          .
          <source>Journal of molecular biology 305(3)</source>
          ,
          <fpage>567</fpage>
          -
          <lpage>580</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Kullback</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <source>Information theory and statistics. Courier Corporation</source>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Lafferty</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>F.C.</given-names>
          </string-name>
          :
          <article-title>Conditional random fields: Probabilistic models for segmenting and labeling sequence data (</article-title>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Lample</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ballesteros</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subramanian</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kawakami</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Neural architectures for named entity recognition</article-title>
          .
          <source>arXiv preprint arXiv:1603.01360</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          , E.:
          <article-title>End-to-end sequence labeling via bi-directional lstm-cnns-crf</article-title>
          .
          <source>arXiv preprint arXiv:1603.01354</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>MacKay</surname>
          </string-name>
          , D.J.:
          <article-title>A practical bayesian framework for backpropagation networks</article-title>
          .
          <source>Neural computation 4(3)</source>
          ,
          <fpage>448</fpage>
          -
          <lpage>472</lpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Nadeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sekine</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A survey of named entity recognition and classification</article-title>
          .
          <source>Lingvisticae Investigationes</source>
          <volume>30</volume>
          (
          <issue>1</issue>
          ),
          <fpage>3</fpage>
          -
          <lpage>26</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Neal</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          :
          <article-title>Bayesian learning for neural networks</article-title>
          , vol.
          <volume>118</volume>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Paisley</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Variational bayesian inference with stochastic search</article-title>
          .
          <source>arXiv preprint arXiv:1206.6430</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.: Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In: Empirical Methods in Natural Language Processing (EMNLP)</source>
          . pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          (
          <year>2014</year>
          ), http://www.aclweb.org/anthology/D14-1162
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Rabiner</surname>
            ,
            <given-names>L.R.:</given-names>
          </string-name>
          <article-title>A tutorial on hidden markov models and selected applications in speech recognition</article-title>
          .
          <source>Proceedings of the IEEE</source>
          <volume>77</volume>
          (
          <issue>2</issue>
          ),
          <fpage>257</fpage>
          -
          <lpage>286</lpage>
          (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Rezende</surname>
            ,
            <given-names>D.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohamed</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wierstra</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Stochastic backpropagation and approximate inference in deep generative models</article-title>
          .
          <source>arXiv preprint arXiv:1401.4082</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Sˇabata</surname>
          </string-name>
          , T.,
          <string-name>
            <surname>Borovicka</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holena</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>K-best viterbi semi-supervized active learning in sequence labelling</article-title>
          . CEUR workshop proceedings pp.
          <fpage>144</fpage>
          -
          <lpage>152</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Sang</surname>
            ,
            <given-names>E.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Meulder</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Introduction to the conll-2003 shared task: Languageindependent named entity recognition</article-title>
          .
          <source>arXiv preprint cs/0306050</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Scheffer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Decomain</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wrobel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Active hidden markov models for information extraction</article-title>
          .
          <source>In: International Symposium on Intelligent Data Analysis</source>
          . pp.
          <fpage>309</fpage>
          -
          <lpage>318</lpage>
          . Springer (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Schuster</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paliwal</surname>
            ,
            <given-names>K.K.:</given-names>
          </string-name>
          <article-title>Bidirectional recurrent neural networks</article-title>
          .
          <source>IEEE Transactions on Signal Processing</source>
          <volume>45</volume>
          (
          <issue>11</issue>
          ),
          <fpage>2673</fpage>
          -
          <lpage>2681</lpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Seshadri</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sundberg</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          :
          <article-title>List viterbi decoding algorithms with applications</article-title>
          .
          <source>IEEE transactions on communications 42(234)</source>
          ,
          <fpage>313</fpage>
          -
          <lpage>323</lpage>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Settles</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Active learning literature survey</article-title>
          .
          <source>Science</source>
          <volume>10</volume>
          (
          <issue>3</issue>
          ),
          <fpage>237</fpage>
          -
          <lpage>304</lpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <surname>Settles</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Craven</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An analysis of active learning strategies for sequence labeling tasks</article-title>
          .
          <source>In: Proceedings of the conference on empirical methods in natural language processing</source>
          . pp.
          <fpage>1070</fpage>
          -
          <lpage>1079</lpage>
          . Association for Computational Linguistics (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yun</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipton</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kronrod</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anandkumar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Deep active learning for named entity recognition</article-title>
          .
          <source>Proceedings of the 2nd Workshop on Representation Learning for NLP</source>
          (
          <year>2017</year>
          ). https://doi.org/10.18653/v1/w17-2630, http://dx.doi. org/10.18653/v1/
          <fpage>W17</fpage>
          -2630
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          39.
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mansimov</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakhudinov</surname>
          </string-name>
          , R.:
          <article-title>Unsupervised learning of video representations using lstms</article-title>
          .
          <source>In: International conference on machine learning</source>
          . pp.
          <fpage>843</fpage>
          -
          <lpage>852</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          40.
          <string-name>
            <surname>Tomanek</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hahn</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Semi-supervised active learning for sequence labeling</article-title>
          .
          <source>In: Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP: Volume 2 -</source>
          Volume 2. pp.
          <fpage>1039</fpage>
          -
          <lpage>1047</lpage>
          . ACL '
          <volume>09</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>