<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Nominal Coreference Annotation in IberEval2017: the case of FORMAS group</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marlo Souza</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafael Glauber</string-name>
          <email>rglauber@dcc.ufba.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leandro Souza de Oliveira</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cleiton Fernando Lima Sena</string-name>
          <email>cflsena2@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniela Barreiro Claro</string-name>
          <email>dclaro@ufba.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Mathematics and Statistics, Federal University of Bahia - UFBA</institution>
          ,
          <addr-line>Av. Adhemar de Barros, S/N, Ondina - Salvador-BA</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>M Souza</institution>
          ,
          <addr-line>R Glauber, L. S. de Oliveira, C F L Sena, D B Claro</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>92</fpage>
      <lpage>101</lpage>
      <abstract>
        <p>This work describes the participation of the FORMAS group from Federal University of Bahia (UFBA) in the Shared Task on Collective Elaboration of a Coreference Annotated Corpus for Portuguese Texts for IberEval 2017. As such, it describes the creation of a corpus annotated with coreference information for the Portuguese language. We discuss the choices adopted oin the annotation process, as well as the results obtained and their possible application to the development of methods and systems focusing on the processing of texts in portuguese.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Anaphora and coreference resolution are well-established problems in the
literature of Computational Linguistics [
        <xref ref-type="bibr" rid="ref11 ref14 ref17 ref9">17, 11, 9, 14</xref>
        ]. While there is some
terminological confusion regarding the use of anaphora resolution and coreference
identi cation, in this work we adopt these terms to be similar and to refer to the
problem commonly known as anaphora resolution in the computational
linguistics literature. We consider the problem of nominal coreference resolution as the
problem concerning the identi cation of two (or more) nominal phrases which
refer to the same discourse entity in the domain of discourse [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>This work describes the creation of a corpus annotated with coreference
information in the context of the Shared Task on Collective Elaboration of a
Coreference Annotated Corpus for Portuguese Texts for IberEval 2017. Our team was
composed of ve researchers, with three main annotators.</p>
    </sec>
    <sec id="sec-2">
      <title>The problem of identifying coreference chains</title>
      <p>
        It has been pointed out in the literature that the problem of coreference
resolution and the guidelines for corpus annotation for this task are underspeci ed, or
that, at least, there are some terminological inadequacies in their de nition [
        <xref ref-type="bibr" rid="ref13 ref18">18,
13</xref>
        ]. Thus, as a rst step in our group's annotation e ort, we tried to establish
a common understanding of the phenomenon and the di culties regarding the
annotation process.
      </p>
      <p>
        Commonly, two noun phrases (NPs) are said to co-refer if they \refer to the
same entity" [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This de nition requires of coreferring noun phrases the
properties that (i) they refer directly to an (unique, unambiguous) entity and (ii) this
entity is identi able from the context of the noun phrases. These requirements,
however, are true only for a small set of noun phrases in a text.
      </p>
      <p>
        On the notion of referring used in this work, while the semantic/philosophical
logic notion of referring is a relation between a linguistic expression and an
object, it is clear that this relation is usually too restrictive to explain the notion
we are interested in this work. Otherwise, NPs contained in counterfactual or
hypothetical statements, as (1) below, would be non-referring, even if the noun
phrases `a car' and `it' are naming the same entity in the universe of the discourse,
i.e. a hypothetical car. As such, in this work we adopt a broader notion of
referring, which holds between two linguistic expressions. In this case it is not
problematic to say that the pronoun `it' in sentence (1) refers to the same entity
as the NP `a car'. The we adopt the de nition that two NPs co-refer if they refer
to the same entity introduced in the universe of discourse, i.e. have the same
discourse referent in the nomenclature of functional grammar theory [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>(1) If I had a car, I would drive it to the coast.</p>
      <p>
        Notice that noun phrases may have several semantic functions in a sentence,
according to Poesio [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and not all noun phrases refer directly to an entity. To
understand which kind of NP may be of importance to the annotation task, we
must investigate further the uses of NPs in the language. Some of the functions
a NP may have in a sentence are:
{ Referring: a NP is said to be referring if it refers directly to some discourse
entity, as the NP \a car" in sentence (1).
{ Quanti cational: a noun phrase may acts as a quanti cation over the
domain of discourse bounding the interpretation of the predicate to a set of
discourse entities denoted by the NP. For example, consider the NP \Every
TV network" in sentence (2) below. This NP act as a quanti cation ranging
over all discourse entities which are considered to be TV networks. As such,
the (logical) meaning of the sentence (2) may be expressed by the logical
expression (2').
      </p>
      <p>(2) Every TV network reported its pro ts.</p>
      <p>
        (2') 8x:(T v network(x) ! 9y:(reported(x; y) ^ prof its of (y; x)))
Notice that, as Van Deemter and Kibble [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] point out, we cannot simply
take the quanti cational NP \Every TV network" in sentence (2) to directly
refer to the class of all TV network entities in the universe of discourse,
otherwise the (anaphoric) relation between the reference of this NP and the
pronoun `its' cannot be properly established. For instance, if we take `its'
to co-refer with 'Every TV network`, the sentence (2") below would be a
paraphrase of (2).
      </p>
      <p>
        (2") Every TV network reported every TV network's pro ts.
{ Predicative: a NP can express properties of an object, and it can not refer
to a speci c entity as in the case of `a preacher' in sentence (3) provided by
Poesio [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], which describes a property (i.e. a predicate) of the entity referred
by `Kim'. As such, the (logical) meaning of sentence (3) can be expressed by
the logical expression (3').
      </p>
      <p>(3) Kim is a preacher.</p>
      <p>(3') preacher(Kim)
{ Expletive: some languages require the presence of certain verbal arguments,
such as French or English in which null subjects are not allowed in certain
types of sentences. In these languages, non-referring NPs may be used as
llers, to occupy a required syntactical position in the sentence. This is the
case of the pronoun 'It' in sentence (4) for the English language and similarly
the pronoun `Il' in sentence (5) for the French language.</p>
      <p>(4) It's two o'clock.</p>
      <p>(5) Il est deux heures.</p>
      <p>
        From this discussion, it is clear that only those NPs that refer to some entities
in the domain of discourse are of interest to the annotation process, namely those
fuunctioning as referring NPs and as quanti cational NPs. As Poesio [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] points
out, however, it is not always clear how to classify a given noun phrase in a
sentence according to their semantic function, even for a human [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. More yet,
usually there are di erent ways to interpret a given NP depending on the adopted
linguistic theory.
      </p>
      <p>Another aspect of identi cation of coreferent NPs concerns how to delimit
when the referents of two NPs can be considered equal. Notice that the relation
between the entities referred by two distinct NPs may not be that of identity,
and yet it is arguable the case that the two NPs to corefer. Let's consider the
case of the case of the NPs \The house" and \the bathroom" in the sentence (6)
below.</p>
      <p>(6) The house is great, but the bathroom is too dark and humid.</p>
      <p>
        The referents of these two NPs are related by a meronymy relation, i.e. the
bathroom to which the second NP refers is a part of the house to which the rst
refers. As such, the NPs refers to the same entity, but to di erent parts of it.
This kind of coreference, which were subject to anotation in the MUC-7 task [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
is often called associative coreference.
      </p>
      <p>
        In the annotation task described in this work, we do not consider associative
coreferences. We aim to annotate only the cases of coreference established by
means of referring NPs and by quanti cational NPs. However, as Van Deemter
and Kibble [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] point out, it is not always clear how to establish the reference
of a quanti cational NP. On one hand, if we establish that the referent of a
quanti cational NP, such as \Every TV network" in sentence (7) we lose the
coreference relation between this NP and the pronoun \its" in the same sentence.
4
If we take the meaning to be a single TV Network bound to the context of
the quanti cation, we may not establish the connection between "Every TV
network" and the pronoun \they" in the sentence (7) below.
      </p>
      <p>(7)Every TV network reported its pro ts. They are required by the
government to do so.</p>
      <p>This is not an easy problem to x. Particularly, depending on the context,
either option in de ning the referent of the NP may be more suitable. Since
our aim in this annotation is to maximize the annotation of coreferent NPs, we
establish that either possibility may be taken by the annotator, as long as it is
done consistently throughout the text. In other words, the annotator may choose
either that the referent of \Every TV network" is the same as \it" or the same
as \they", as long as the annotator does not change the referent while analyzing
the text.</p>
      <p>Regarding possible di culties relating change over time, for descriptors like
\the president of Brasil" for which the reference is dependent on a temporal
context, we adopted the same strategy to that of the MUC-7 conference, i.e.
\two markables should be recorded as coreferential if the text asserts them to
be coreferential at ANY TIME" [7, p. 11], as long as it is clear by context that
the reference of the NPs is intended to be the same.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The corpus</title>
      <p>To perform the annotation task, we composed a corpus of thirty encyclopedic
texts written in portuguese language, taken from the Wikipedia 1. Wikipedia
texts are an important resource for languages with scarce computational
linguistic resources, since they compose a corpus of a signi cant size, usually coupled
with important annotation, such as the domain classi cation, cross-references
between pages in di erent languages, etc.</p>
      <p>
        Wikipedia corpora have been widely used in the NLP literature for its
availability, structure and existing metadata. For the portuguese language,
particularly, the Wikipedia Corpus has been used in Ontology Learning [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], Open
Information Extraction [
        <xref ref-type="bibr" rid="ref12 ref2 ref21 ref22">21, 2, 22, 12</xref>
        ], Named Entity Recognition [
        <xref ref-type="bibr" rid="ref19 ref3">3, 19</xref>
        ], as well
as been subject of the Pagico - Shared Task on information retrieval [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], among
many others.
      </p>
      <p>The texts that compose the corpus used in the annotation process were
randomly selected from the Wikipedia dump of March 26 of 2017 (03.26.2017) using
the Wikipedia Extractor tool2. To select the texts we have established the
following criteria:
1. the text must have approximatively 1200 (between 1100 and 1400) words;
2. the text must not concern physical or mathematical theories, nor contain
mathematical formulas as gures;
1 http://pt.wikipedia.org
2 Available at: http://medialab.di.unipi.it/wiki/Wikipedia_Extractor
3. the topic of the text must not be wiki meta-information, such as discussion
pages;
4. the text must be a running text discussion of a topic, excluding thus any
Wikipedia page containing lists (e.g. List of awards received by Justin
Timberlake).</p>
      <p>The rst requirement was made to conform to the shared task speci cation of
texts containing 1200 words. The second requirement is justi ed by the fact that
complex mathematical formulas are not easy to parse by text processors and, in
fact, are excluded from the text in the extraction using the Wikipedia Extractor.
Since the pages describing mathematical and physical theories commonly rely in
a great amount of mathematical formulas to describe their topics, the extracted
text becomes poorly informative and di cult to process. The third and fourth
requirements are made to exclude uninteresting pages which are become common
in the corpus considering the restriction of texts with size of 1200 words.
4</p>
    </sec>
    <sec id="sec-4">
      <title>The annotation tool</title>
      <p>
        Following the methodology established for the shared task, the annotation
process consisted of two steps. In the rst step, an initial automatic annotation
of the corpus was performed by the organizing team, in which each NP in a
text is identi ed and those NPs participating in a coreference chain are grouped
together. In the second step, each annotation team performed the manual
correction of the initial annotation, both of the problems of delimitation of NPs
and the identi cation of coreference chains. To perform this manual correction,
the annotation teams used the tool CorrefVisual [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], provided by the organizing
team. Regarding our experience with the tool and the annotation process, we
o er some considerations for the task and the resulting corpus.
      </p>
      <p>First, it was our impression, in the annotation process, that the visual
grouping of the elements in the same coreference chain, as provided by the annotation
tool, does indeed help the veri cation that all noun phrases in the same chain are
actually coreferent. When the number of coreference chains grew, however, the
navigation through all these groupings became a hindrance in the annotation.
Since the tool only provided manual navigation or text search, deciding whether
a given NP should be included in some existing chain usually involved navigating
through several coreference chains, which became very time-consuming.</p>
      <p>Regarding the noun phrase identi cation and delimitation, while the tool
allowed the correction of the boundaries of noun phrases, in some cases the noun
phrases were not even partially identi ed by the tool. Since the annotation tool
did not allow the creation of new noun phrases, only to alter the boundaries
of those already identi ed, some coreference relations were not possible to be
annotated. This was a common occurrence in the presence of compound noun
phrases such as \os g^eneros Sambucus e Viburnum" (the genera Sambucus and
Virbunum), as the tool normally presented only either the option with the
complete NP \os g^eneros Sambucus e Viburnum" or two separated noun phrases
\Sambucus" and \Viburnum".
6</p>
      <p>While it is our belief that the annotation tool did indeed help the annotation
process, reducing the amount of labor involved in it, we also believe that the
amount of restrictions imposed by the tool to the annotators may have an impact
in the quality of the resulting corpus.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Annotation evaluation</title>
      <p>The agreement in the annotation was measured by means of the Kappa statistics.
Four texts were annotated by all three annotators for this comparison. The
agreement for each text and among the group is depicted in Table 1.</p>
      <p>To understand these results, we performed a quantitative analysis of the
texts to determine the reason for the high deviation in the agreement among
the texts. The hypothesis was that the higher agreement was achieved in
simpler texts, while the lowest agreements were achieved on more complex ones.
For this quantitative analysis, we evaluated the number of noun phrases in the
text (#NPs), as well as the statistics of the annotation, such as number of
identi ed chains (#chains), number of NPs that had a coreference relationship
with another (#correferent), the average size of the coreference chains in the
text (Avg size of chains) and the size of the biggest identi ed chain (Size of
biggest chain) based on the annotation of each annotator of the group. The
resulting data are depicted in Tables 2, 3 and 4</p>
      <p>Notice that in text 1, the text with higher agreement, the amount of identi ed
coreference chains in small (for all annotators) compared to the others and the
Text #NPs #chains #correferents Avg size of chains Size of biggest chain
text 1 422 17 85 5.0 32
text 2 452 29 129 4.45 45
text 3 407 28 130 4.64 22
text 4 479 75 295 3,93 30
number of NPs that participate in a coreference chain is also slightly smaller
than for the other texts, which seems to agree with our hypothesis.</p>
      <p>While the increase in the number of identi ed chains in text 4 did reduce
the agreement, this behavior was not observed in text 2, which has the lowest
agreement and almost the same number of identi ed chains as text 3. Also, it is
not clear that the increase from 25 coreference chains with 106 coreferent NPs
in text 1 to 42 chains with 191 coreferent NPs in text 3 (for annotator 1) could
explain such a signi cant variation in agreement between the two texts.</p>
      <p>Analyzing text 3, we identi ed that this text su ers from extreme poor
writing quality, being, in fact, a translation from the English language with several
only partially translated expressions within it. In this context, both the NP
delimitation as well as the coreference determination became compromised, which
explain the low value in agreement for the annotators. As such, we consider that
the text should be treated as an outlier and non representative of the result of
the annotation.</p>
      <p>Regarding the general results for each annotator, notice that annotators 1
and 2 are more coherent with each other in their annotations, identifying similar
number of coreference chains and coreferent NPs for all texts, while annotator
3 deviates more from the other two. One possible explanation for such behavior
is that, apart from the inherent di culty of establishing reference of NPs, the
notion of coreference adopted in the work may not have been a consensus for all
annotators.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Coreference information in Open Information</title>
    </sec>
    <sec id="sec-7">
      <title>Extraction</title>
      <p>
        Open Information Extraction (Open IE) is the area that studies methods for
extracting information from fragment texts without any previous constraint on
the kind of relations to be identi ed. It was introduced by Banko et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] with
8
      </p>
      <p>M Souza, R Glauber, L. S. de Oliveira, C F L Sena, D B Claro
the system TextRunner and has ourished into an active area of research in
Natural Language Processing.</p>
      <p>
        Early Open IE method relied on using domain-independent extraction
patterns to identify relation instantiations [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. As a result, several extractions made
by this systems have low quality, the result of what Etzioni et al.[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] call
incoherent and uninformative extractions. According to Etzioni et al.[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], incoherent
extractions are those which have no meaningful interpretation, while
uninformative extractions are those in which critical information for the interpretation
of the expression is missing. These authors claim that methods that tackle the
problem of reducing these low quality extractions compose the second generation
of Open IE systems.
      </p>
      <p>Notice that resolving coreference is an essential challenge to guarantee the
minimization of uninformative extractions of Open IE systems. The reason for
that is that outside its discursive context, pronouns and other descriptors of
discourse entities have no clear referent. Let's analyze the text (6) below.
(6) Mariana's car is more reliable than that of Louis. She takes very good care
of it. The car was revised this week</p>
      <p>Typical examples of extracted information from Open IE systems,
considering the text of (6) would be the tuples represented in (6*),(6**) and (6***).
(6*) (Mariana's car, is more reliable, that of Louis)</p>
      <p>(6**)(She, takes good care of, it)
(6***) (The car, was revised, this week)</p>
      <p>Without its linguistic context, the extractions (6**) and (6***) are
uninformative, since it is not clear to which entities the pronouns \She" and `'it" refer
to in sentence (6**), nor to which car (Mariana's or Louis') the NP \The car"
refers to in sentence (6***).</p>
      <p>With the construction of a corpus of coreference, we aim to allow the
development of Open IE systems for the Portuguese language which explore coreference
information to extract more informative relations without while obtaining the
information that would be lost if we discarded the extractions (6**) and (6***).
7</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>The present work described the participation of the FORMAS group in the
shared task for the collective elaboration of a coreference annotated corpus for
Portuguese texts at IberEval 2017. In this work, we discussed the notion of
coreference adopted by our group for the annotation process, the corpus we used
as well as our experience during the annotation process.</p>
      <p>We believe that, while the annotation showed moderate agreement from the
annotators, the resulting corpus can (and will) be an important resource to the
Portuguese language, allowing the development of interesting methods and
systems focusing on this language, considering the inherent di culty of solving this
problem and the scarcity of resources for the Portuguese language. Particularly,
we plan to evaluate the resulting corpus within the context of Open IE in the
near future, to measure how such a resource can foster the development of
methods that minimizes the extraction of uninformative relations, while being able
to extract all the information from a text.</p>
      <p>In regard to the shared task, we consider that the notion of coreference
adopted expected in the shared task was not clearly de ned, and, as such, we
felt the necessity to explicit the notions and choices adopted by our groups. As
a result of the underspeci cation of the task, it is not clear to us how the many
corpora generated by di erent groups participating in the shared task can be
united to form a corpus for coreference resolution in the Portuguese language. We
believe that, to merge all these corpora, it will be necessary to identify possible
inconsistencies in the annotation processes, apart from the pure quantitative
evaluation of inter-annotator agreement.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Banko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cafarella</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soderland</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Broadhead</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Open information extraction from the web</article-title>
          .
          <source>In: IJCAI</source>
          . vol.
          <volume>7</volume>
          , pp.
          <volume>2670</volume>
          {
          <issue>2676</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Batista</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forte</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martins</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Extraccao de relacoes sema^nticas de textos em portugu^es explorando a dbpedia e a wikipedia. linguamatica 5(1</article-title>
          ),
          <volume>41</volume>
          {
          <fpage>57</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cardoso</surname>
          </string-name>
          , N.:
          <article-title>Rembrandt-reconhecimento de entidades mencionadas baseado em relacoes e analise detalhada do texto</article-title>
          .
          <source>In: Encontro do Segundo HAREM</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fader</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christensen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soderland</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mausam</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Open information extraction: The second generation</article-title>
          .
          <source>In: IJCAI</source>
          . vol.
          <volume>11</volume>
          , pp.
          <volume>3</volume>
          {
          <issue>10</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Fonseca</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sesti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira</surname>
          </string-name>
          , R.:
          <article-title>Guia para anotaca~o de correfer^encia</article-title>
          . http://ontolp.inf.pucrs.br/corref/ibereval2017/CorrefVisual.zip (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gundel</surname>
            ,
            <given-names>J.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hedberg</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zacharski</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>Cognitive status and the form of referring expressions in discourse</article-title>
          . Language pp.
          <volume>274</volume>
          {
          <issue>307</issue>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hirschman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chinchor</surname>
          </string-name>
          , N.:
          <article-title>Muc-7 coreference task de nition</article-title>
          .
          <source>In: MUC-7 Proceedings. Science Applications International Corporation</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hirschman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vilain</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Automating coreference: The role of annotated training data</article-title>
          .
          <source>In: Proceedings of the AAAI Spring Symposium on Applying Machine Learning to Discourse Processing</source>
          . pp.
          <volume>118</volume>
          {
          <issue>121</issue>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hirst</surname>
          </string-name>
          , G.:
          <article-title>Discourse-oriented anaphora resolution in natural language understanding: A review</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>7</volume>
          (
          <issue>2</issue>
          ),
          <volume>85</volume>
          {
          <fpage>98</fpage>
          (
          <year>1981</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bollegala</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsuo</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ishizuka</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Using graph based method to improve bootstrapping relation extraction</article-title>
          .
          <source>Computational linguistics and intelligent text</source>
          processing pp.
          <volume>127</volume>
          {
          <issue>138</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mitkov</surname>
            ,
            <given-names>R.: Anaphora</given-names>
          </string-name>
          <string-name>
            <surname>Resolution. Longman</surname>
          </string-name>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Pires</surname>
            ,
            <given-names>J.C.B.</given-names>
          </string-name>
          : Extraca~o e Mineraca~o de Informaca~o Independente de
          <article-title>Dom nios da Web na L ngua Portuguesa</article-title>
          .
          <source>Master's thesis</source>
          , Universidade Federal de Goias (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Poesio</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Linguistic and cognitive evidence about anaphora</article-title>
          .
          <source>In: Anaphora Resolution</source>
          , pp.
          <volume>23</volume>
          {
          <fpage>54</fpage>
          . Springer Berlin Heidelberg (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Poesio</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stuckardt</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Versley</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Anaphora resolution (</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Poesio</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>A corpus-based investigation of de nite description use</article-title>
          .
          <source>Computational linguistics 24(2)</source>
          ,
          <volume>183</volume>
          {
          <fpage>216</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Santos</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Porqu^e o pagico? razoes para uma avaliacao conjunta</article-title>
          .
          <source>Linguamatica</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ), 1{
          <issue>8</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Van Deemter</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kibble</surname>
          </string-name>
          , R.:
          <article-title>On coreferring: Coreference in muc and related annotation schemes</article-title>
          .
          <source>Computational linguistics 26(4)</source>
          ,
          <volume>629</volume>
          {
          <fpage>637</fpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Van Deemter</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kibble</surname>
          </string-name>
          , R.:
          <article-title>On coreferring: Coreference in muc and related annotation schemes</article-title>
          .
          <source>Computational linguistics 26(4)</source>
          ,
          <volume>629</volume>
          {
          <fpage>637</fpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Construca~o de um corpus anotado para classi caca~o de entidades nomeadas utilizando a Wikipedia e a DBpedia. Master's thesis, Pontif cia Universidade Catolica do Rio Grande do Sul (</article-title>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Xavier</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Lima</surname>
            ,
            <given-names>V.L.S.:</given-names>
          </string-name>
          <article-title>A semi-automatic method for domain ontology extraction from portuguese language wikipedia's categories</article-title>
          .
          <source>In: Brazilian Symposium on Arti cial Intelligence</source>
          . pp.
          <volume>11</volume>
          {
          <fpage>20</fpage>
          . Springer, Berlin, Heidelberg (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Xavier</surname>
          </string-name>
          , C.C.,
          <string-name>
            <surname>de Lima</surname>
            ,
            <given-names>V.L.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Souza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Open information extraction based on lexical-syntactic patterns</article-title>
          .
          <source>In: Intelligent Systems (BRACIS)</source>
          ,
          <source>2013 Brazilian Conference on</source>
          . pp.
          <volume>189</volume>
          {
          <fpage>194</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Xavier</surname>
          </string-name>
          , C.C.,
          <string-name>
            <surname>de Lima</surname>
            ,
            <given-names>V.L.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Souza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Open information extraction based on lexical semantics</article-title>
          .
          <source>Journal of the Brazilian Computer Society</source>
          <volume>21</volume>
          (
          <issue>1</issue>
          ),
          <volume>4</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>