<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multidomain Contextual Embeddings for Named Entity Recognition</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Word Em-</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ponti cal Catholic University of Rio Grande do Sul (PUCRS)</institution>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>434</fpage>
      <lpage>441</lpage>
      <abstract>
        <p>Neural Networks are widely used for Named Entity Recognition due to their capability of extracting features from texts automatically and integrating them with sequence taggers. Pretrained Language Models are also extensively used for NER, as their product, Word Embeddings, are key elements for improving the performance of NER systems. A novel type of embeddings, called Contextual Word Embeddings, can adapt according to the context it is inserted, something traditional word embeddings could not do. These contextual embeddings have proven to be superior to traditional embeddings for NER. In this work, we show the results of our network, which uses Neural Networks in conjunction with a contextual language model, on corpora composed of texts belonging to rarely tested textual genres, such as o cial police documents and clinical notes, as proposed by a task in IberLEF 2019.</p>
      </abstract>
      <kwd-group>
        <kwd>Neural Networks Named Entity Recognition beddings</kwd>
        <kwd>Flair Embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Named Entity Recognition (NER), a task in the eld of Natural Language
Processing (NLP), consists of nding proper nouns in a given text and to classify
them on di erent prede ned categories [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Modern approaches for the NER
task utilize Neural Networks (NN) that automatically learn the features from
raw text and making the use of manually constructed rules obsolete [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The use of vector representation of words (word embeddings) have helped to
increase the quality of NER systems. These embeddings can be created through
the training of a Language Model (LM) on a corpus of raw texts [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] - [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
Modern LMs create contextualized embeddings in runtime, as opposed to retrieving
information from static word-vector dictionaries as do traditional LMs. One of
      </p>
      <p>
        Recent State-of-the-Art results for sequential labeling tasks has used Deep
Learning strategies. The Long-Short Term Memory (LSTM) networks has shown
cutting-edge results on this type of task due to the fact that they allow to analyze
the context in which a word is inserted in two directions of a given sentence:
forward and backward [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        On this work we present our system for Portuguese NER proposed by the
Shared Task \Portuguese Named Entity Recognition and Relation Extraction
Tasks (NerRelIberLEF2019) " on IberLEF 2019 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Iberian Languages
Evaluation Forum (IberLEF) is a forum that has that aims to develop the Natural
Language Processing for Iberian American languages. On this year of 2019 the
shared task NerRelIberlef presented three challenges (tasks) for researchers on
the NLP area: a task about NER and other two about Relation Extraction. In
what follows, this work describes our approach and resources used on the
participation for the NER task, to which we will refer to Task 1. Our system makes use
of a BiLSTM Neural Network fed by a composition of other two language
models: Flair Embeddings and Word Embeddings. A nal layer of a CRF classi er
returns the token classi cation.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Bidirectional Encoder Representations from Transformers (BERT) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] achieved
the State-of-the-Art using a complex network based around Transformer Neurons
[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Recent works for NER in the English language showed excellent results with
the use of contextualized word embeddings. BERT was recently surpassed by Flair
Embeddings [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which utilizes contextualized embeddings from a character-level
Language Model.
      </p>
      <p>
        Embeddings from Language Models (ELMo) is an approach that creates
embeddings based on a deep BiLSTM network, this enables ELMo to analyze
the context into which any given word is inserted. ELMo is also character based,
enabling the model to create vector representations words it was not trained on
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        In the context of Named Entity Recognition for Portuguese, there is a
proposal for the use of a Deep Neural Network (DNN) called CharWNN [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
CharWNN is a speci c type of DNN utilized for sequence tagging, extracting features
from word and character level.
      </p>
      <p>
        Another approach also using Neural Networks was presented by [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. It uses
LSTM networks that has a Conditional Random Fields (CRF) classi er as its
last layer, based on the network proposed by Lample et al. (2016) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Table 1
shows the F1-measure of the cited works in this section.
Multidomain Contextual Embeddings for Named Entity Recognition
Recurrent Neural Networks (RNN) are currently considered to be the standard
networks for NLP [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Of these networks, Long-Short Term Memory (LSTM)
networks stand out in abundance of use. LSTM networks are a variation of the
RNN with a re ned architecture that shows better results for sequential tasks
[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. The LSTM architecture uses functions that determine whether previously
added information should be kept, modi ed or discarded in order to better relate
to newly added information. An important variation of LSTM networks are
BiLSTM networks. These are composed of two LSTM working in parallel. One
of the networks deals with a \forward" data sequence (F orward LSTM) and the
other with a \backwards" data sequence (Backwards LSTM). This makes it so
the network has a greater learning capacity [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Conditional Random Fields (CRF) is a classi er used for the construction of
probabilistic models with the goal of segmenting and labeling sequential data.
This type of classi er has been widely used on tasks of Part-of-Speech Tagging
and NER [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Recent works about NER in the Portuguese language adopt CRF as
a nal component of the system in order to give the token it's nal classi cation
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Neural Embeddings or Word Embeddings are ways to represent words in
ndimensional vector spaces. Recently, this type of representation has been widely
used in the eld of NLP [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Among these, Word2Vec [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] stands out. Word2Vec
is freely available and is based on RNN, being able to learn the representation
of words in high-dimensional vector spaces [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <p>
        Word2Vec is divided into two architectures, Continuous Bag of Words (CBOW)
and Skip-Gram. CBOW uses a context as an input for the network and then the
desired word gets returned, on the other hand, using Skip-Gram architecture,
we use a word as an input and then the context in which the word is presented
is returned [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        For our system, we used the 300-dimensional Word2Vec Skip-Gram
(W2VSKPG) language model made available by the Interinstitutional Center for
Computational Linguistics of Sa~o Paulo University (NILC). The embeddings are
available in their website (http://nilc.icmc.usp.br/embeddings).
Flair Embeddings is a recent language modelling architecture that works on
both the character and word levels. It goes beyond traditional word embedding
architectures like Skip-gram and CBOW, as it takes into account the
characterlevel morphological features of words as well as the more traditional contextual
information. Because of this, the authors consider Flair's embeddings to be
Contextual String Embeddings [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The Flair Model used in our system, called FlairBBP, was trained with a
corpus of over 4 billion tokens and is available for use in our GitHub page
(https://github.com/jneto04/ner-pt).
4</p>
    </sec>
    <sec id="sec-3">
      <title>Neural Network for NER</title>
      <p>
        Our NER model is the product of one of our previous works, where we trained
a Neural Network for this task. We used a BiLSTM-CRF Neural Network that
has been previously used for NER in English and German [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
BiLSTMCRF network was trained using a structure of embeddings concatenation called
Stacking Embeddings. That is, each one of our tokens were represented by a
compilation of two types of embeddings: Flair Embeddings (FlairBBP) and Word
Embeddings (W2V-SKPG). The matrix w of the equation 4 shows our stacking
of embeddings. The Neural Network used for this work is composed of two layers,
a Character Language Model and a Sequence Labeling Model. First, all of the
tokens are passed to the Character Language Model of the Stacking Embeddings,
which then returns a vector r for each input token.
      </p>
      <p>w =</p>
      <p>wF lairBBP
wW 2V SKP G</p>
      <p>This vector r is then passed to the \Sequence Labeling Model" where a
BiLSTM network receives the vectors r and pass it output to the CRF classi er
that returns the token classi cation.</p>
      <p>Multidomain Contextual Embeddings for Named Entity Recognition</p>
      <p>The training of the network was done with a corpus using the First HAREM
(https://www.linguateca.pt/HAREM/), considering the categories Person, Place,
Organization, Time and Value.</p>
      <p>
        Table 2 shows the hyperparameters used on training.
The system was evaluated using three di erent test corpora: a Police Dataset, a
Clinical Dataset, and a General Dataset (Created from SIGARRA [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and the
second HAREM [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]). Table 3 presents the results provided by Task 1's
coordination team using the CoNLL-2002 script [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>
        Out of all of the submitted systems that participated in the shared Task
1[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], our system achieved the best F1-measure for the General Dataset
(Overall). We attribute that to the Stacking Embeddings and to the test corpus that
was used for this part of the task. The General Dataset is composed of two
corpora: SIGARRA and Second HAREM (Relation Version) that are relatively
close structurally and linguistically to the HAREM collection with which our
system was trained.
      </p>
      <p>
        Our results for the Police Dataset were competitive, having achieved a 2.8%
lower F1-measure score than the best system for this dataset. Our good
performance with this test dataset is also attributed to the embeddings and to the
fact that the texts used to build the Police Dataset are very well structured, like
those used to build the HAREM collection [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In the case of the Clinical Dataset, the di erence between the F1-measure of
our system and the system with highest F1-measure of this dataset is 12.98%.
We believe that this is due to the fact that the Clinical Dataset's unusual
structure and language. It is composed of clinical notes containing abbreviations,
medical terms and various other particularities found in texts from hospital
environments. Texts like this di er greatly from the traditional style for the NER
task in Portuguese, and these di erences were not taken into account during the
system's training.
6</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>This paper presented our proposed approach for \Task 1: Named Entity
Recognition" in the NerRelIberLEF Shared Task in IberLEF 2019. Our approach
involved the use of a BiLSTM-CRF that receives a compilation of highly
representational embeddings: FlairBBP + W2V-SKPG. As such, we understand
that our results come from the representational power of the Flair Embeddings
architecture in representing a natural language.</p>
      <p>As a future work, we plan to train our system with more corpora, as we
believe this will yield even better metrics.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>We thank CNPQ for their nancial support.
Multidomain Contextual Embeddings for Named Entity Recognition</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Akbik</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blythe</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vollgraf</surname>
          </string-name>
          , R.:
          <article-title>Contextual string embeddings for sequence labeling</article-title>
          .
          <source>In: Proceedings of the 27th International Conference on Computational Linguistics</source>
          . pp.
          <volume>1638</volume>
          {
          <issue>1649</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. do Amaral,
          <string-name>
            <given-names>D.O.F.</given-names>
            ,
            <surname>Vieira</surname>
          </string-name>
          , R.:
          <article-title>Nerp-crf: uma ferramenta para o reconhecimento de entidades nomeadas por meio de conditional random elds</article-title>
          .
          <source>Linguamatica</source>
          <volume>6</volume>
          (
          <issue>1</issue>
          ),
          <volume>41</volume>
          {
          <fpage>49</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. de Castro,
          <string-name>
            <given-names>P.V.Q.</given-names>
            ,
            <surname>da Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.F.F.</given-names>
            ,
            <surname>da Silva</surname>
          </string-name>
          <string-name>
            <surname>Soares</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Portuguese named entity recognition using lstm-crf</article-title>
          .
          <source>In: International Conference on Computational Processing of the Portuguese Language</source>
          . pp.
          <volume>83</volume>
          {
          <fpage>92</fpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Collobert</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karlen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavukcuoglu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuksa</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Natural language processing (almost) from scratch</article-title>
          .
          <source>Journal of machine learning research 12(Aug)</source>
          ,
          <volume>2493</volume>
          {
          <fpage>2537</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Collovini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Consoli</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Terra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quaresma</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Souza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Claro</surname>
            ,
            <given-names>D.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glauber</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xavier</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          :
          <article-title>Portuguese named entity recognition and relation extraction tasks at iberlef</article-title>
          <year>2019</year>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Freitas</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carvalho</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goncalo</surname>
            <given-names>Oliveira</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Mota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Second harem: advancing the state of the art of named entity recognition in portuguese</article-title>
          . In: quot; In Nicoletta Calzolari; Khalid Choukri; Bente Maegaard; Joseph Mariani; Jan Odijk; Stelios Piperidis; Mike Rosner; Daniel Tapias (ed)
          <source>Proceedings of the International Conference on Language Resources and Evaluation (LREC</source>
          <year>2010</year>
          )
          <article-title>(Valletta 17-</article-title>
          23 May de 2010)
          <article-title>European Language Resources Association</article-title>
          .
          <source>European Language Resources Association</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Graves</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jaitly</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohamed</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          r.:
          <article-title>Hybrid speech recognition with deep bidirectional lstm</article-title>
          .
          <source>In: 2013 IEEE workshop on automatic speech recognition and understanding</source>
          . pp.
          <volume>273</volume>
          {
          <fpage>278</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>La</surname>
            <given-names>erty</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>F.C.</given-names>
          </string-name>
          :
          <article-title>Conditional random elds: Probabilistic models for segmenting and labeling sequence data (</article-title>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lample</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ballesteros</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subramanian</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kawakami</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Neural architectures for named entity recognition</article-title>
          .
          <source>arXiv preprint arXiv:1603.01360</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Dependency-based word embeddings</article-title>
          .
          <source>In: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          :
          <string-name>
            <given-names>Short</given-names>
            <surname>Papers</surname>
          </string-name>
          <article-title>)</article-title>
          .
          <source>vol. 2</source>
          , pp.
          <volume>302</volume>
          {
          <issue>308</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient estimation of word representations in vector space</article-title>
          .
          <source>arXiv preprint arXiv:1301.3781</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>3111</volume>
          {
          <issue>3119</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Milidiu</surname>
            ,
            <given-names>R.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duarte</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cavalcante</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>: Machine learning algorithms for portuguese named entity recognition. Inteligencia Arti cial</article-title>
          . Revista Iberoamericana de Inteligencia Arti cial
          <volume>11</volume>
          (
          <issue>36</issue>
          ) (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Murdoch</surname>
            ,
            <given-names>W.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szlam</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Automatic rule extraction from long short term memory networks</article-title>
          .
          <source>arXiv preprint arXiv:1702.02540</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iyyer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gardner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>arXiv preprint arXiv:1802</source>
          .
          <volume>05365</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Pires</surname>
            ,
            <given-names>A.R.O.</given-names>
          </string-name>
          :
          <article-title>Named entity extraction from Portuguese web text</article-title>
          .
          <source>Master's thesis</source>
          , Faculdade de Engenharia da Universidade de Porto, Porto, Portugal (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Pirovani</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Oliveira</surname>
          </string-name>
          , E.:
          <article-title>Portuguese named entity recognition using conditional random elds and local grammars</article-title>
          .
          <source>In: LREC</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Sang</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erik</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Introduction to the conll-2002 shared task: languageindependent named entity recognition</article-title>
          .
          <source>In: Proceedings of CoNLL-2002</source>
          . pp.
          <volume>155</volume>
          {
          <issue>158</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Santos</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zadrozny</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Learning character-level representations for part-ofspeech tagging</article-title>
          .
          <source>In: Proceedings of the 31st International Conference on Machine Learning (ICML-14)</source>
          . pp.
          <year>1818</year>
          {
          <year>1826</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Santos</surname>
            ,
            <given-names>C.N.d.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guimaraes</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Boosting named entity recognition with neural character embeddings</article-title>
          .
          <source>arXiv preprint arXiv:1505.05008</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Tai</surname>
            ,
            <given-names>K.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Improved semantic representations from tree-structured long short-term memory networks</article-title>
          .
          <source>arXiv preprint arXiv:1503.00075</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>5998</volume>
          {
          <issue>6008</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , D.,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Chinese comments sentiment classi cation based on word2vec and svmperf</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>42</volume>
          (
          <issue>4</issue>
          ),
          <year>1857</year>
          {
          <year>1863</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>