<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RECKONition: a NLP-based system for Industrial Accidents at Work Prevention</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>(Discussion Paper)</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrizia Agnello</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silvia Maria Ansaldi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emilia Lenzi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessio Mongelluzzo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Roveri</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>INAIL - Dipartimento Innovazioni Tecnologiche</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Extracting patterns and useful information from Natural Language datasets is a challenging task, especially when dealing with data written in a language diferent from English, like Italian. Machine and Deep Learning, together with Natural Language Processing (NLP) techniques have widely spread and improved lately, providing a plethora of useful methods to address both Supervised and Unsupervised problems on textual information. We propose RECKONition, a NLP-based system for Industrial Accidents at Work Prevention. RECKONition, which is meant to provide Natural Language Understanding, Clustering and Inference, is the result of a joint partnership with the Italian National Institute for Insurance against Accidents at Work (INAIL). The obtained results showed the ability to process textual data written in Italian describing industrial accidents dynamics and consequences.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Workplace Safety</kwd>
        <kwd>Industry 4</kwd>
        <kwd>0</kwd>
        <kwd>Natural Language Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Despite the ever-growing awareness and the advances in the technology and the procedure, the
problem of accidents at work represents a relevant issue from both the social and economic
point of view. In more detail, accidents at work refer to damages to the health of the worker
induced by accidents correlated with the working activities. Unfortunately, such accidents can
result into serious damages, such as temporary or permanent reduction or loss of working
capacity, or even death. In Italy, this issue is particularly perceived as very relevant since in
last few years the number of accidents at work is larger than 500.000 per year, among which
more that 1000 resulted in the death of the worker. Addressing this issue in Italy is the primary
goal of INAIL, which is National Insurance Institute for Industrial Accidents and Occupational
Diseases. More specifically, INAIL’s objectives are the reduction of injuries, the protection
of workers performing hazardous jobs, and the facilitation of the return to work of people
injured at workplace. To achieve these relevant and challenging objectives, INAIL explores
novel technological and societal solutions and paradigms to provide an integrated system of
protection, ranging from preventive actions at the workplace to medical services and financial
assistance.</p>
      <p>
        Many Natural Language Processing (NLP) applications and solutions can be found in the
literature to tackle tasks like Sentiment Analysis or Text Classification, for instance [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However,
when it comes to the Italian language, and injury prevention in particular, there are no suited
solutions available. In this path, the aim of this paper is to introduce RECKONition, a novel
Natural Language Processing-based system to extract knowledge from textual descriptions of
accidents at work aiming at defining industrial accidents preventive actions at the workplace.
The RECKONition system comprises three diferent modules, i.e., Association Rule Generation,
Textual Description Clustering and Textual Description Inference, able to extract relationships
between accident events and create groups of homogeneous accident descriptions representing
the input for the definition of preventive actions. The RECKONition system has been successfully
applied to a real-world database storing the textual descriptions of accidents occurring in
Northern Italy from 2013-2018.
      </p>
      <p>The paper is organized as follows. In Section 2 introduce RECKONition system and its three
algorithmic modules. In Section 3, 4, and 5 respectively, we detail our implementation of the
Association Rules Mining, Natural Language Clustering, and Natural Language Prediction modules.
Section 6 summarizes our work and discuss on the possible next steps and improvements</p>
    </sec>
    <sec id="sec-2">
      <title>2. The RECKONition system: extracting knowledge from textual descriptions of accidents at work</title>
      <p>The architecture of the RECKONition system, which is described in Fig. 1, comprises three
diferent modules: Association Rule Generator, Textual Description Clustering and Textual
Description Inference. All these models receive in input a set of textual descriptions of accidents
at work and provide in output association rules to discover relationships between terms, clusters
of textual descriptions highlighting groups of accidents at work sharing similarities in their
descriptions, and next sentence predictions from the textual description of the accidents. All
the outputs of the RECKONition system represent valuable tools for the INAIL expert to gain
knowledge about occurred accidents at work and define prevention policies. The three modules
of RECKONition system are detailed in the next sections.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Association Rules Mining</title>
      <p>
        Association Rules (ARs) allow to discover relations between variables in large datasets and
to graphically show such dependencies in a convenient representation which can be easily
interpreted by humans [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] (we show in Figure 2 an example of an AR representation).
!(HELMET)
(INJURY)
      </p>
      <p>Let us define our database  as a collection of transactions from the transactions set  : a
transaction is defined as an element of our database, i.e., a sentence containing a textual
description of the work dynamics leading to an accident. Hence, for each transaction we can
define a set of words or a transformation of words, called items, which belong to the items set 
and constitute our transaction. In this scenario, an association rule  ⇒  indicates that the
occurrence of item  is usually observed when also  is found in the transaction, for ,  ⊂ .</p>
      <p>Association Rules mining procedure is based on three measures, namely Support, Confidence ,
and Lift, whose values define whether a rule is meaningful or not. The support of an item  can
be defined as () =  () = |{|∀∈ ∧∈}| ,namely, the percentage of transactions where
| |
item  is found. Support can also be defined for a rule, e.g., ( ⇒ ), to represent the
percentage of transactions in the database where both  and  are found. With the confidence
of a rule we model the conditional probability of observing the consequent item having observed
the antecedent, so it can hence be defined as  ( ⇒ ) =  () . The third measure, i.e.,
 ()
the lift, is used to characterize the relationship between the antecedent and the consequent of
the association rule. It is defined as  ( ⇒ ) =  () () =  (|) , a lift greater than 1
 ()
 ()
identifies a positive relationship between the items, i.e., the conditional probability of observing
 given that  is found is greater than the probability of observing  in the database.</p>
      <p>
        Apriori algorithm is a well-known association rules mining method in the literature, which
leverages the measures we introduced to explore the itemset and build association rules from
them [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This method only focuses on the subset of those items belonging to the Frequent
Itemset (FIS), i.e, those items whose support is greater than a minimum threshold .
However, in some settings (e.g., medical descriptions) those items whose occurrence frequency
is very low can be the ones carrying more semantical importance (e.g., names of diseases or
even specific injuries). Following this intuition, we provided a Python implementation 1 to the
algorithm FISinFIS Apriori proposed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], where Positive Association Rules (PARs) and Negative
Association Rules (NARs) are built from both the Frequent Itemset and the Infrequent Itemset.
FISinFIS Apriori algorithm sets a minimum threshold on the Inverse Document Frequency
| |
(IDF), defined as  () = log |{|∈ ∧∈}| , to filter those items which are too frequently
1The code is available at https://github.com/AlessioMongelluzzo/FISinFIS_Apriori_Python
used throughout the database. Moreover, we introduced in our implementation an additional
threshold to the IDF as an upper bound for too rarely used items. The rationale behind this
decision is to avoid typos and missing values placeholders which constitute a non-negligible
part of our dataset. The results coming from the application of the mining algorithm to our
work accidents descriptions database allows RECKONition to detect those industrial properties
or actions that are more likely related to an injury or, with negative rules, to their prevention,
we show in Figure 3 an AR graph example from a mock experiment performed on few sentences
we wrote for this purpose.
      </p>
      <p>!(FINGER)
(GLOVES)</p>
      <p>(WEAR)</p>
    </sec>
    <sec id="sec-4">
      <title>4. Natural Language Clustering</title>
      <p>
        Textual Description Clustering [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is performed to highlight diferences and similarities among
accident descriptions, and to group together the similar ones in order to facilitate the
identification of appropriate interventions according to the situation described. In the following sections
we will show the two diferent approaches used by our system. We will see both of them use
the K-Medoids algorithm [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], but difer in the operations performed on the dataset.
      </p>
      <sec id="sec-4-1">
        <title>4.1. Tags Occurrence Clustering</title>
        <p>A first approach is based on an ontology describing the context of interest (in the case of
RECKONition, the metallurgical company) and uses TAGs representing ontological classes to
integrate contextual knowledge in the clustering analysis.</p>
        <p>1
Dataset choice
+ pre-processing</p>
        <p>Dataset
split</p>
        <p>Training set:
90% of the
samples
2 Occurrences 3 100</p>
        <p>iCd1oe0nu0tnitwfic+oarftdiirosstn TAdGepf-ianWiirtsiOoRnD
Test set:
10% of the
samples
4</p>
        <p>TAG - WORD
substitution
in Test set
5</p>
        <p>TF - IDF
calculation</p>
        <p>Clustering
ready
dataset
6</p>
        <p>Clusters
Clustering
algorithm</p>
        <p>
          K
As you can see in Figure 4, we calculate the occurrences of each word, we identify the hundred
most frequent ones, and we define TAG - WORD pairs in the training set; we then perform the
hundred substitution in the test set. In this way, words representing similar concepts are read
by the algorithm as the same term. As can also be seen from the diagram, before calculating
the occurrences, all the necessary pre-processing operations were carried on the entire dataset,
and in particular the stop_words were eliminated. Almost all the remaining words resulted
to be descriptive of the context then, and it was possible to associate them with a TAG. In the
rare case in which this did not happen, or the substitution introduced ambiguity, the terms
remained unchanged. Once performed the substitution, we calculate the Time frequency
Inverse Document Frequency (TF-IDF) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] of each term in each description. After step 5 we
than have a dataset ready for clustering and it is composed by 6662 terms as features, whose
values correspond to the calculated TF-IDF.
        </p>
        <p>Once we have the dataset, we perform K-Medoids. Here we report some example of the
clusters obtained by setting the  = n_clusters parameter equal to 30, and we show how we
are able to separate the diferent situations described.</p>
        <sec id="sec-4-1-1">
          <title>Cluster 12: medoid: [USING HOSE] elements: [USING HOSE]; [USING DRILL]; [USING HAMMER]; [USING CUTTER];...</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>Cluster 17: medoid: [GOT OFF FORKLIFT] elements: [GOT OFF FORKLIFT]; [RIDE FORKLIFT]; [MOVE TRUCK]; [RIDING TRUCK];...</title>
          <p>What is important to notice is that the algorithm is not only able to group together descriptions
containing the exact same words, but it is also able to identify the ones concerning similar
incidents in the analysed context. In this perspective, Clusters 12 represents accidents involving
working tools, while Clusters 17 represents accidents occurring on board mobile machinery.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Transformers-Based Clustering</title>
        <p>
          A diferent clustering approach adopted by RECKONition leverages the novel Encoder-Decoder
models with Attention and self-Attention mechanism to build homogenous groups of accidents
descriptions [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. These techniques comprise several stacked encoder-decoder layers, which
can model diferent syntactic structures with self-Attention modules together with feed-forward
neural networks, resulting in milions of trainable parameters: the base version of the
Bidirectional Encoder Representations from Transformers (BERT) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] features 12 encoder-decoder
layers, 768 hidden nodes for each layer and 12-attention-heads, resulting in 110M parameters,
hence, requiring a very large amount of data to be properly trained, as well as time and
computational capacity. RECKONition uses a pre-trained BERT model publicly available2, which
was trained on 13GB of Italian textual data from Wikipedia and OPUS corpora3, and performs
ifne-tuning of such model on a subset of the accidents-at-work dataset provided by INAIL. The
2The pre-trained model can be found at https://huggingface.co/dbmdz/bert-base-italian-cased
3https://opus.nlpl.eu/
ifne-tuned model is then validated on a hold-out set which is used to perform clustering on
unseen textual descriptions: during the forward step of a new sentence, the hidden states of the
last decoder layer in the model are extracted and used as numerical features representing the
sentence. This procedure is performed for each description in the validation set, so to build a
new dataset with shape  × (768 * ), being 768 the number of hidden states of a BERT layer,
 the number of sentences in the validation set, and  the number of length of the tokenized
sentences. RECKONition leverages Incremental Principal Component Analysis (IPCA) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] to
reduce the dimensionality of the representative features to a number of principal components
so that at least 85% of the variability of data is explained, and performs K-Medoids clustering
for values of the n_clusters  in the range [
          <xref ref-type="bibr" rid="ref2">2, 100</xref>
          ]. Unlike the clustering approach described in
Section 4.1, this method can efectively find sentences with highly correlated syntactic
structures. We show in the following two examples of clusters obtained with the Transformers-based
clustering approach applied to a sample dataset.
        </p>
        <p>MEDOID: "RIDING THE FORKLIFT "
["RIDING THE FORKLIFT", "RIDING THE CAR", "RIDING THE FORKLIFT ", "RIDING THE CAR", "RIDING
THE CAR", "RIDING THE BIKE", "RIDING THE FORKLIFT", "REPAIRING THE FORKLIFT"]
MEDOID: "CLIMBING THE LADDER"
["CLIMBING THE LADDER", "CLIMBING THE WALL", "CLIMBING THE LADDER", "CLIMBING ON THE
ROOF", "CLIMBING THE LADDER ", "CLIMBING THE LADDER ", "REMOVING HEAR DEFENDER"]</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Natural Language Prediction</title>
      <p>
        There is plenty of experimental settings where a Language Model (LM) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], i.e., a model that
is able to process textual sequences [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], can be efectively applied. These range from feature
extraction (as we did with Trasformers-based clustering in Section 4.2), to text classification,
question answering, next sentence prediction, natural language inference, or even text
generation. RECKONinition comprises a custom LM, which was properly built and trained on the
accidents dataset provided by INAIL, with the purpose of performing next sentence prediction
from the textual description of the accidents dynamics to those of the consequences on the
workers involved. The high-level architecture of the LM is shown in Figure 5. We emphasize
that the tokenization step [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], which allows to move from symbolical (textual) to numerical
representation, requires to define a vocabulary size to fix the number of tokens, which our
model sets to 5000 items. The embedding layer is then used to project the sparse representation
obtained by the tokenization step to a dense space with size 128. The sequence learning ability of
the model resides in the n layer shown in Figure 5, which is composed of two Bidirectional Long
Short-Term Memory (LSTM) with 100 units each [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], a flattening layer, two dense layers
with 50 neurons each and ReLU activations, and a final Dropout [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] layer with a 50% drop rate
before the output dense layer with 5000 neurons (one for each item) and sigmoidal activation.
This model can be trained in a supervised fashion [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] minimizing the Binary Cross Entropy
loss function with Adaptive Moment Estimation (ADAM) [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. This is achieved by providing the
descriptions of the accidents dynamics as input features, together with the corresponding efects
and injuries on the workers involved as input targets. For a new unseen sentence describing a
working scenario, the model will output a probability vector on the items in the vocabulary
defining those which better describe the consequence of the input description.
s
t t(s)
e e(t)
n
n(e)
d
s'=d(n(e))
      </p>
      <p>We show in in the following two examples of textual prediction given some input descriptions,
where the output sentence results in a scattered sequence of items due to the limitation on the
vocabulary size.</p>
      <p>Description: "the worker was climbing the ladder"
Prediction: ["the", "to", "its", "fall", "off", "ankle", "sprained"]
Description: "she was driving her car to go home"</p>
      <p>Prediction: ["she", "her", "to", "the", "cervical", "trauma"]</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion and conclusion</title>
      <p>We introduced and described RECKONition system and its application scenario on a real world
dataset. The system is currently operating to understand and extract knowledge from the
accidents descriptions database and support the industries towards a safer workplace. The
results obtained and the continuous dialogue with INAIL’s partners have highlighted the
unsupervised ability of RECKONition to both cluster work activities and to identify relationship
between accidents and consequences. As a next step, RECKONition will be deployed in real
factories where proper sensors can produce inputs to the system, allowing real-time predictions
to detect hazardous activities and at the same time improving data quality and performances of
the system itself.</p>
      <p>Acknowledgements
This work has been funded by INAIL within the BRiC/2018, ID09 framework, project RECKON.
The authors wish to acknowledge all the other researchers involved in the project: Francesco
Braghin (PoliMI, DMEC), Enrico Cagno (PoliMI, DIG), and their research groups. We also thank
Cinzia Frascheri (IAL) and Irene Tagliaro (API-TECH).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Badjatiya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Varma</surname>
          </string-name>
          ,
          <article-title>Deep learning for hate speech detection in tweets</article-title>
          ,
          <source>in: Proceedings of the 26th international conference on World Wide Web companion</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>759</fpage>
          -
          <lpage>760</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Imieliński</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Swami</surname>
          </string-name>
          ,
          <article-title>Mining association rules between sets of items in large databases</article-title>
          ,
          <source>in: Proceedings of the 1993 ACM SIGMOD international conference on Management of data</source>
          ,
          <year>1993</year>
          , pp.
          <fpage>207</fpage>
          -
          <lpage>216</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Srikant</surname>
          </string-name>
          , et al.,
          <article-title>Fast algorithms for mining association rules</article-title>
          ,
          <source>in: Proc. 20th int. conf. very large data bases, VLDB</source>
          , volume
          <volume>1215</volume>
          ,
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <year>1994</year>
          , pp.
          <fpage>487</fpage>
          -
          <lpage>499</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mahmood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shahbaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guergachi</surname>
          </string-name>
          ,
          <article-title>Negative and positive association rules mining from text using frequent and infrequent itemsets</article-title>
          ,
          <source>The Scientific World Journal</source>
          <year>2014</year>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rajaraman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ullman</surname>
          </string-name>
          , Mining of massive datasets cambridge university press,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.-S.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Jun</surname>
          </string-name>
          ,
          <article-title>A simple and fast algorithm for k-medoids clustering</article-title>
          ,
          <source>Expert systems with applications 36</source>
          (
          <year>2009</year>
          )
          <fpage>3336</fpage>
          -
          <lpage>3341</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Van</given-names>
            <surname>Merriënboer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gulcehre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bougares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwenk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Learning phrase representations using rnn encoder-decoder for statistical machine translation</article-title>
          ,
          <source>arXiv preprint arXiv:1406.1078</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>arXiv preprint arXiv:1706.03762</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Van</given-names>
            <surname>Merriënboer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gulcehre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bougares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwenk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Learning phrase representations using rnn encoder-decoder for statistical machine translation</article-title>
          ,
          <source>arXiv preprint arXiv:1406.1078</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ducharme</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vincent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Janvin</surname>
          </string-name>
          ,
          <article-title>A neural probabilistic language model</article-title>
          ,
          <source>The journal of machine learning research 3</source>
          (
          <year>2003</year>
          )
          <fpage>1137</fpage>
          -
          <lpage>1155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <article-title>Sequence to sequence learning with neural networks</article-title>
          ,
          <source>arXiv preprint arXiv:1409.3215</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Webster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kit</surname>
          </string-name>
          ,
          <article-title>Tokenization as the initial phase in nlp</article-title>
          ,
          <source>in: COLING 1992 Volume 4: The 15th International Conference on Computational Linguistics</source>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <article-title>Long short-term memory</article-title>
          ,
          <source>Neural computation 9</source>
          (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jaitly</surname>
          </string-name>
          , A.-r. Mohamed,
          <article-title>Hybrid speech recognition with deep bidirectional lstm</article-title>
          , in: 2013 IEEE workshop
          <article-title>on automatic speech recognition and understanding</article-title>
          , IEEE,
          <year>2013</year>
          , pp.
          <fpage>273</fpage>
          -
          <lpage>278</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>N.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <article-title>Dropout: a simple way to prevent neural networks from overfitting</article-title>
          ,
          <source>The journal of machine learning research 15</source>
          (
          <year>2014</year>
          )
          <fpage>1929</fpage>
          -
          <lpage>1958</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>C. M. Bishop</surname>
          </string-name>
          ,
          <source>Pattern recognition and machine learning</source>
          , springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic optimization</article-title>
          ,
          <source>arXiv preprint arXiv:1412.6980</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>