<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Coreference Annotation of the CSTNews Corpus</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Instituto Federal de Sa~o Paulo</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidade Federal de Goias</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universidade Federal de Sa~o Carlos</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universidade de S~ao Paulo</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Universidade do Algarve</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>102</fpage>
      <lpage>112</lpage>
      <abstract>
        <p>We report in this paper the coreference annotation process of the CSTNews corpus as part of a collective task of the IberEval 2017 conference. The annotated corpus is composed of 140 news texts written in Brazilian Portuguese language and counts with several annotation layers, including annotations in the morphosyntax/syntax, semantics, and discourse levels. The annotation, focused on nominal references, was conducted in a semi-automatic way by ve teams, achieving satisfactory annotation agreement results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Coreference resolution is the task of nding linguistic expressions in a text that
refer to the same entity [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. As an illustration of coreference occurrence, we show
below a short text with some coreferent elements in bold. In this short text, the
referring expressions \a passenger plane", \the airplane" and \it" refer to the
same entity and form a \coreference chain". It is interesting to notice that the
coreference resolution task includes pronominal anaphora resolution, which is
part of the problem.
      </p>
      <p>
        At least 17 people died after the crash of a passenger plane in the
Democratic Republic of Congo. According to an ONU spokeswoman, the
airplane was trying to land in the Bukavu airport in the midst of a
storm. It failed to reach the runway and fell in a forest 15 kilometers
away from the airport.
Coreference resolution and its subtasks have been investigated for a long time
in the Natural Language Processing (NLP) area. There are approaches based on
heuristics [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ], machine learning [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], and discourse theories, as Centering [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
and Veins Theory [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The task has also been the focus of a shared task in CoNLL
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Besides its history, the task is still a challenge, as it brings together the
di culties of automatically dealing with semantics and discourse. Such linguistic
analysis levels usually require deep linguistic processing capabilities from the
machines, which, in turn, demand automatic text interpretation techniques.
      </p>
      <p>Coreference is a very important information for NLP systems. For instance,
it is essential for summarization systems to produce coherent and cohesive
summaries, allowing the systems to properly \glue" together text passages; it is
useful to track, on the web, entities of interest; it may be necessary for
information extraction about entities of interest and for performing the related inference;
it may help simplifying texts, allowing changing pronominal anaphoras by
nominal antecedents, as certain anaphoric mentions cause di culties for people with
cognitive disabilities; and so on. Therefore, research e orts in such task are very
relevant. Although some coreference annotated corpora and automatic
coreference annotation softwares are available for English (the interested reader may
refer to the shared task in CoNLL), similar initiatives for other languages are
still rare, including Portuguese, which is the focus of this paper.</p>
      <p>
        We report here the coreference annotation process of a corpus of news texts
written in Brazilian Portuguese language - the CSTNews corpus [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The
annotation, focused on nominal references, was conducted in a semi-automatic way by
5 annotation teams as part of a collective task of the IberEval 2017 conference,
achieving satisfactory annotation agreement results (considering the di culty of
the task).
      </p>
      <p>In what follows (Section 2), we present some basic concepts regarding
coreference. In Section 3, we introduce the corpus we annotated. Section 4 reports
the annotation process and the achieved results, describing some challenges of
the task. Some nal remarks are presented in Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Coreferences</title>
      <p>Detecting coreferent elements and composing coreference chains is a task that
demands sophisticated linguistic knowledge. Coreference mainly happens at the
semantics/discourse interface, as it requires the identi cation of the meaning of
the elements (i.e., to what they refer to) in the same sentence or across sentences
in a text. There are also initiatives for detecting coreference relations across
di erent texts, which is relevant for multi-document processing purposes.</p>
      <p>
        As presented in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the entities in a text are evoked by \referring expressions".
The element that refers to a previous one in the text (e.g., \the airplane" in the
sample text in the Introduction section) is the \anaphoric element", while the
element that is referred to (\a passenger plane") is called the \antecedent" (or
\the referent", according to some authors). Such elements are said to corefer,
that is, they refer to the same extra-linguistic entity. When we track and store
all the references to an entity, we perform \coreference resolution", and the set
of references is a \coreference chain".
      </p>
      <p>
        Referring expressions may happen in several forms in a text, usually
syntactically realized as noun phrases. For instance, we may use pronouns (e.g., \it"),
common nouns (\airplane") and proper nouns (\William Shakespeare"),
optionally presenting determiners, pre- and/or post-modi ers (e.g., \the beautiful girl")
and having high size variance (e.g., \the beautiful and charming girl that was
looking at me"). In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the authors discuss such variations and the di culties
that they bring. The authors comment that the reference may happen directly,
when the same noun is used to refer to another one, as in the text passage \The
letter was signed yesterday. In the letter, the scientists argue that...", or in
an indirect way, when di erent terms are used, as in \The letter was signed
yesterday. In the text, the scientists argue that...". In such cases, the reference
to the original expression may be recovered by accessing linguistic and world
knowledge, as synonymy, hypernymy/hyponymy, and meronymy/holonymy
relations, several types of pronouns, acronyms and abbreviations, and verb
nominalizations, among several others. We suggest consulting the work of [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for the
interested reader.
      </p>
      <p>
        To the best of our knowledge, the only manually annotated corpus in
Brazilian Portuguese that is speci cally focused on coreference chains is the Summ-it
corpus [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], which is generally used for training and testing systems. There are
other corpora annotated with named entities, which might be used for such end,
but that were not speci cally built to address the task of coreference resolution.
As argued in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], \the Portuguese coreference resolution area is at an early
stage of development", but some initiatives exist. The CORP system1 [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ],
for instance, has been used by the research community.
      </p>
      <p>In such scenario, producing coreference annotated corpora and developing
and/or improving coreference resolution systems for Portuguese are key issues
to be pursued. The annotation e ort reported in this paper is a step towards
new advances in the NLP area for Portuguese. In what follows, we introduce the
corpus that was annotated.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The CSTNews Corpus</title>
      <p>
        The CSTNews corpus [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] was originally developed for multi-document
summarization purposes during the SUCINTO project2. The corpus includes 140 news
texts written in Brazilian Portuguese, from some main online news agencies in
Brazil, as Folha de Sa~o Paulo, Estad~ao, Jornal do Brasil, Gazeta do Povo, and O
Globo. The texts are grouped in 50 clusters (each cluster has 2 our 3 texts), and
the texts of a cluster are on the same topic. The texts are about Economy,
Politics, Sports, Science, Daily News, World News, and Financial subjects. Several
summaries are associated to each cluster, including single and multi-document
1 http://ontolp.inf.pucrs.br/corref/
2 http://www.icmc.usp.br/~taspardo/sucinto/
summaries, extractive and abstractive summaries, and automatically and
manually produced summaries, as the corpus was originally intended for application
in the summarization area.
      </p>
      <p>
        In the years following its creation, the corpus received several linguistic
annotation layers (mainly of discourse nature), according to the uses it had. The
corpus currently includes:
{ manual multi-document discourse annotation, according to the Cross-document
Structure Theory (CST) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which was the rst annotation in the corpus
and gave origin to its name;
{ manual single document discourse annotation, according to the Rhetorical
      </p>
      <p>
        Structure Theory (RST) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ];
{ manual subtopic segmentation, following the proposal of [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ];
{ manual informative aspect identi cation, according to the guidelines
proposed at the Guided Summarization task3 of the TAC conference [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ];
{ manual identi cation and normalization of temporal expressions, following
the proposal of [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ];
{ manual word sense disambiguation of verbs and (the most frequent) common
nouns, using Princeton WordNet as sense repository [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ];
{ manual text-summary alignment, indicating which sentences from the source
texts gave origin to the sentences of the summaries and through which
rewriting operations;
{ automatic morphosyntax and syntax annotation by the PALAVRAS parser
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>Except for the last one, which was automatic, all of these annotations were
manually carried out in systematic and controlled way. Satisfactory annotation
agreement results were obtained (when applicable), which allows to infer that
the data is reliable.</p>
      <p>
        With such annotations, the corpus has subsidized several research e orts,
including traditional multi-document summarization [22{24], the more recent
update summarization [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], summary coherence evaluation [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], the study of
human behavior for summary production [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], the development of discourse parsers
[
        <xref ref-type="bibr" rid="ref28 ref29">28, 29</xref>
        ] and application on information extraction [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], phrase generalization [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ],
and theoretical studies on some discourse aspects [32{34], among several others
that may be seen in the SUCINTO project website.
      </p>
      <p>The coreference annotation reported here constitutes a new annotation layer
in the CSTNews corpus. The annotation process is described in the following
section.
4</p>
    </sec>
    <sec id="sec-4">
      <title>The Annotation</title>
      <p>The coreference annotation of the CSTNews corpus was carried out as a
collective task, named \Collective Elaboration of a Coreference Annotated Corpus
3 https://tac.nist.gov//2010/Summarization/Guided-Summ.2010.guidelines.html
for Portuguese Texts"4, of the IberEval 2017 (Evaluation of Human Language
Technologies for Iberian Languages) conference5.</p>
      <p>In the collective task, annotation teams with at least 3 members should
submit their corpus for annotation. If selected to participate, the members of the
team would receive pre-processed texts to annotate. Part of the texts was from
the corpus submitted by the team (4 of them were used to compute the
annotation agreement of the team), and another part was from the corpora submitted
by other teams. Each member should then individually annotate the coreference
chains in his/her texts. Only nominal referring expressions were intended to be
annotated.</p>
      <p>
        As said above, the texts were presented in a pre-processed form. They were
automatically annotated by the CORP coreference resolution system [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ],
which performs (in some cases, with the aid of some other tools) part of speech
tagging, shallow syntactic parsing, and construction of coreference chains by
using syntactic heuristics (as string matching, copular constructions, and
juxtaposition of linguistic expressions in the text) and semantic knowledge (as synonymy
and hyponymy relations) obtained from the Onto.PT resource [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. The
collective annotation task consisted, therefore, in reviewing and eventually correcting
the automatic coreference annotation that was presented, which was done with
the aid of another tool, the CorrefVisual, which is a graphical interface that
allows to visualize the text and the available preprocessed coreference chains, and
to edit the chains.
      </p>
      <p>Figure 1 shows the CorrefVisual interface with an annotated text loaded.
The text is in Portuguese, as it is in the corpus. It is about an accident in an
airport, where an airplane crashed into a building. One may see the text at
the left of the panel, the coreference chains in the middle, and the remaining
referring expressions that did not form any chain, called \unique mentions", on
the right side. Above these unique mentions, there is an auxiliary panel to help
dealing with the chains. In the interface, each chain is associated to a di erent
color, and, every time that an element is selected, its occurrence in the text is
highlighted in the corresponding color.</p>
      <p>Correcting a chain consisted, therefore, in moving the referring expressions
among the windows in the tool, e.g., removing an expression from a chain and
adding it to another one, incorporating a unique mention to some existing chain,
creating new chains, etc. If necessary, the auxiliary panel would help in grouping
and moving the elements across the windows. The tool also allowed to edit the
referring expressions, by adding or removing words next to it (to the left or to
the right of the expression). This is an important step, as such expressions were
automatically detected and some errors probably occurred (in fact, during the
annotation, we have noticed that such errors were very frequent). Unfortunately,
the tool does not allow to use new expressions that were not previously detected
by the tool itself. The edition of the target expressions had also been explicitly
and highly discouraged by the organizers of the collective annotation task. Such
4 http://ontolp.inf.pucrs.br/corref/ibereval2017/
5 http://nlp.uned.es/IberEval-2017/
decision bene ts the annotation agreement results, but is also based on the idea
that we should make use of the available state of the art preprocessing tools,
with their current limitations and potentialities.</p>
      <p>
        Finally, one may also notice in the interface that a semantic category should
be assigned to each chain, indicating the type of the entities of the chain. In
the interface, the category appears above each chain, in capital letters. The
available semantic categories in the tool are based on the named entity typology
of the REPENTINO gazetteer [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ], namely: person, organization/place, event,
communication, product, document, abstraction, nature, another living being,
substance, and \other".
      </p>
      <p>In order to fully annotate the CSTNews corpus, 5 annotation teams were
assembled: one with four members and four with three members. Each team
was leaded by its more senior member, usually a researcher with some
experience in NLP and corpus annotation. All the members were native speakers of
Portuguese.</p>
      <p>
        Initially, all the teams read the material made available by the organizers of
the collective annotation task (regarding the task instructions and how to use the
CorrefVisual tool) and also studied a didactic reference paper on coreference for
Portuguese [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. After that, a training step was performed over a single text that
was provided by the task organizers. The teams had the chance to discuss some
annotation issues and, after that, they provided feedback to the task organizers
so that they could improve some functionalities.
      </p>
      <p>After the training with a single text, the task organizers considered that
the task was well understood and provided the teams with the actual texts to
annotate. In average, each annotator had to annotate from 11 to 13 texts, in a
period of a month and a half. In general, most of the teams managed to nish
the task in 2 to 3 weeks, with each member trying to annotate one text per day.</p>
      <p>Some issues were very challenging to deal with. Several members had severe
problems with the annotation tool, which apparently has inconsistent
functionalities in di erent operating systems. Given the di culties with the tool, some
annotators preferred to manually annotate the texts (in paper or in a di erent
electronic edition application) before replicating the process in the tool. More
serious than the inconsistencies of the tool were its limitations regarding the
inclusion of referring expressions and the guideline for avoiding editing the
expressions. Such problems are the main cause of some very poor annotations, with
wrong referring expressions in the chains (e.g., see in Figure 1 the expression a
aeronave que in the rst chain; in English, \the airplane that"), non-nominal
elements (as que; \that"), and expressions in the text that were not detected by
the tool and, therefore, could not be included in any chain.</p>
      <p>
        The task organizers provided the annotation agreement results to the teams.
Agreement was computed over all possible pairs of referring expressions that
formed coreference chains, and, for each pair of expressions, it was indicated
how many annotators said that the pair was in fact coreferring. Kappa measure
[
        <xref ref-type="bibr" rid="ref37">37</xref>
        ] was used for computing agreement. This measure is highly adopted in the
area because it corrects the results for expected chance agreement. The author
proposes that a minimum agreement value of 0.67 is important for temptative
conclusions to be draw from the data. However, the NLP area has learned that
di erent values may be expected in di erent tasks, as such values highly depend
on the di culty and subjectivity inherent in the task. In our case, we expected
lower values.
      </p>
      <p>Table 1 shows the kappa results for each team, including the general average.
As said before, for each team, 4 texts were used for computing kappa. One may
see that, excepting team 2, all the other teams reached an agreement of at least
0.50. In average, the teams presented an agreement of 0.54, which we consider
satisfactory given the annotation conditions (limited training during a short time
period, and di culties and limitations with the annotation tool).
Overall, including the other participants of the collective annotation task, the
minimum agreement value was 0.41, and the maximum was 0.64 (achieved by
our team 4).</p>
      <p>Some very interesting cases of disagreements may be found in the data, which
may be partially explained by the limited training, and partially by di erent
principles regarding the task and the high level of subjectivity in some cases.
For instance, to cite a few (related to the text shown in Figure 1):
{ while one annotator has built a coreference chain with all the elements that
refer to the airport (that is, Congonhas - which is the name of the airport
- and o aeroporto - \the airport", in English), another one has divided this
chain in two, one for the sense of airport and another one for the sense of the
location of the airport, being this di erence very di cult to realize (and this
may explain why the semantic categories \organization" and \place" have
been joined by the task organizers);
{ one annotator has considered that the expressions hipotese (\hypothesis")
and falha mec^anica (\mechanical failure") were coreferent (as the
\hypothesis" for the accident with the airplane was a \mechanical failure"), but
another one considered that these expressions are in di erent generalization
levels and built di erent chains for them;
{ some annotators have built chains for time expressions that refer to the same
event (in such cases, some inference was necessary to determine the correct
time that was referred to), while others did not, considering that time is
not a proper element to a coreference chain (the attentive reader probably
realized that the semantic categories in the annotation tool did not include
\time", which is a traditional class in named entity recognition tasks);
{ some annotators have included in the chains the occurrences of relative
pronouns (e.g., que; \that/which", in English), as they refer to some previous
entities and were detected by the annotation tool, but other annotators did
not consider such items because they are not strictly nominal expressions
(which were the focus of the annotation).</p>
      <p>Di erences as the previous ones may result in broad variations in the nal
annotation. For instance, for the text of Figure 1, the three annotators produced
from 32 to 46 coreference chains (with unique mentions varying from 49 to 86
referring expressions).</p>
      <p>The annotation task required a lot of attention and dedication from the
annotators, which had to read several times each text and its referring
expressions in order to nd all the chains. In many situations, background and domain
knowledge was necessary for correctly identifying the chains. Usually,
annotators consulted the web (mainly Wikipedia) to solve these cases. In average, the
annotators took above 1 hour to annotate each text.</p>
      <p>We present some nal remarks in the next section.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Final Remarks</title>
      <p>All the annotated data is in an XML format, which is a traditional way of
marking and making data available. It shall be available in the SUCINTO project
website, as it constitutes an additional linguistic annotation layer of the
CSTNews corpus.</p>
      <p>
        We expect that the produced coreference annotation fosters other research
initiatives on discourse processing tasks. For the short term, the new data may
help to improve summarization models, speci cally those involving coherence
and cohesion evaluation, for which the occurrence and distribution of referring
expressions are very important features (see, e.g., the entity-based model
proposed in [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ]).
      </p>
      <p>For future work, concerning the CSTNews corpus, the task of pronominal
anaphora resolution remains to be done, as it was not directly tackled in the
reported annotation e ort.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments References</title>
      <p>The authors are grateful to FAPESP, CAPES and CNPq for supporting this
work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          :
          <article-title>Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics and Speech Recognition</article-title>
          . Prentice
          <string-name>
            <surname>Hall</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hobbs</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          :
          <article-title>Resolving pronoun references</article-title>
          .
          <source>Lingua</source>
          <volume>44</volume>
          (
          <issue>4</issue>
          ) (
          <year>1978</year>
          )
          <volume>311</volume>
          {
          <fpage>338</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mitkov</surname>
            ,
            <given-names>R.: Anaphora</given-names>
          </string-name>
          <string-name>
            <surname>Resolution. Pearson Education</surname>
          </string-name>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Haponchyk</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A practical perspective on latent structured prediction for coreference resolution</article-title>
          .
          <source>In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics</source>
          . Volume
          <volume>2</volume>
          . (
          <year>2015</year>
          )
          <volume>143</volume>
          {
          <fpage>149</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Wiseman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rush</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shieber</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          :
          <article-title>Learning global features for coreference resolution</article-title>
          .
          <source>In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          . (
          <year>2016</year>
          )
          <volume>994</volume>
          {
          <fpage>1004</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Grosz</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weisten</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Centering: A framework for modeling the local coherence of discourse</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>21</volume>
          (
          <issue>2</issue>
          ) (
          <year>1995</year>
          )
          <volume>203</volume>
          {
          <fpage>225</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cristea</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ide</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romary</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Veins theory. an approach to global cohesion and coherence</article-title>
          .
          <source>In: Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics</source>
          . (
          <year>1998</year>
          )
          <volume>281</volume>
          {
          <fpage>285</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Pradhan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xue</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uryupina</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , Y., eds.: Joint Conference on EMNLP and
          <article-title>CoNLL { Shared Task, Association for Computational Linguistics (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Cardoso</surname>
            ,
            <given-names>P.C.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maziero</surname>
            ,
            <given-names>E.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castro</surname>
            <given-names>Jorge</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.L.R.</given-names>
            ,
            <surname>Seno</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.M.R.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Di</given-names>
            <surname>Felippo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Rino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.H.M.</given-names>
            ,
            <surname>Nunes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.G.V.</given-names>
            ,
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.A.S.:</surname>
          </string-name>
          <article-title>CSTNews { a discourseannotated corpus for single and multi-document summarization of news texts in brazilian portuguese</article-title>
          .
          <source>In: Proceedings of the 3rd RST Brazilian Meeting</source>
          . (
          <year>2011</year>
          )
          <volume>88</volume>
          {
          <fpage>105</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Vieira</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goncalves</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Souza</surname>
            ,
            <given-names>J.G.C.</given-names>
          </string-name>
          : Processamento computacional de anafora e correfer^encia.
          <source>Revista de Estudos da Linguagem</source>
          <volume>16</volume>
          (
          <issue>1</issue>
          ) (
          <year>2008</year>
          )
          <volume>263</volume>
          {
          <fpage>284</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Collovini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carbonel</surname>
            ,
            <given-names>T.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fuchs</surname>
            ,
            <given-names>J.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coelho</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rino</surname>
            ,
            <given-names>L.H.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira</surname>
          </string-name>
          , R.:
          <article-title>Summ-it: Um corpus anotado com informaco~es discursivas visando a sumarizaca~o automatica</article-title>
          .
          <source>In: Proceedings of V Workshop em Tecnologia da Informaca~o e da Linguagem Humana</source>
          . (
          <year>2007</year>
          )
          <volume>1605</volume>
          {
          <fpage>1614</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Fonseca</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Improving coreference resolution with semantic knowledge</article-title>
          .
          <source>In: Proceedings of the International Conference on Computational Processing of the Portuguese Language. (2016a)</source>
          <volume>213</volume>
          {
          <fpage>224</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Fonseca</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Corp: Coreference resolution for portuguese</article-title>
          .
          <source>In: Proceedings of the International Conference on Computational Processing of the Portuguese Language - Demonstration Session. (2016b)</source>
          <volume>9</volume>
          {
          <fpage>11</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Fonseca</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sesti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atonitsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira</surname>
          </string-name>
          , R.: CORP:
          <article-title>Uma abordagem baseada em regras e conhecimento sema^ntico para a resoluca~o de correfer^encias</article-title>
          .
          <source>LinguaMATICA</source>
          <volume>9</volume>
          (
          <issue>1</issue>
          ) (
          <year>2017</year>
          )
          <volume>3</volume>
          {
          <fpage>18</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.:</given-names>
          </string-name>
          <article-title>A common theory of information fusion from multiple text sources, step one: Cross-document structure</article-title>
          .
          <source>In: Proceedings of the 1st ACL SIGDIAL Workshop on Discourse and Dialogue</source>
          . (
          <year>2000</year>
          )
          <volume>74</volume>
          {
          <fpage>83</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>W.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          :
          <article-title>Rhetorical structure theory: A theory of text organization</article-title>
          .
          <source>Technical Report Technical Report ISI/RS-87-190</source>
          (
          <year>1987</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Texttiling:
          <article-title>Segmenting text into multi-paragraph subtopic passages</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>23</volume>
          (
          <issue>1</issue>
          ) (
          <year>1997</year>
          )
          <volume>33</volume>
          {
          <fpage>64</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Owczarzak</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dang</surname>
          </string-name>
          , H.T.:
          <article-title>Who wrote what where: Analyzing the content of human and automatic summaries</article-title>
          .
          <source>In: Proceedings of the Workshop on Automatic Summarization for Di erent Genres</source>
          ,
          <article-title>Media, and</article-title>
          <string-name>
            <surname>Languages.</surname>
          </string-name>
          (
          <year>2011</year>
          )
          <volume>25</volume>
          {
          <fpage>32</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Baptista</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagege</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mamede</surname>
          </string-name>
          , N.:
          <article-title>Identi caca~o, classi caca~o e normalizaca~o de expresso~es temporais do portugu^es: A experi^encia do segundo HAREM e o futuro</article-title>
          . In Mota,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Santos</surname>
          </string-name>
          , D., eds.:
          <article-title>Desa os na avaliaca~o conjunta do reconhecimento de entidades mencionadas: O Segundo HAREM</article-title>
          . (
          <year>2008</year>
          )
          <volume>33</volume>
          {
          <fpage>54</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>WordNet: An Electronic Lexical Database</article-title>
          . MIT Press (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Bick</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The Parsing System PALAVRAS: Automatic Grammatical Analysis of Portuguese in a Constraint Grammar Framework</article-title>
          . Aarhus University Press (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Silveira</surname>
            ,
            <given-names>S.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Branco</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Enhancing multi-document summaries with sentence simpli cation</article-title>
          .
          <source>In: Proceedings of the 14th International Conference on Arti cial Intelligence</source>
          .
          <article-title>(</article-title>
          <year>2012</year>
          )
          <volume>742</volume>
          {
          <fpage>748</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Ribaldo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cardoso</surname>
            ,
            <given-names>P.C.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pardo</surname>
            ,
            <given-names>T.A.S.:</given-names>
          </string-name>
          <article-title>Exploring the subtopic-based relationship map strategy for multi-document summarization</article-title>
          .
          <source>Journal of Theoretical and Applied Computing</source>
          <volume>23</volume>
          (
          <issue>1</issue>
          ) (
          <year>2016</year>
          )
          <volume>183</volume>
          {
          <fpage>211</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Cardoso</surname>
            ,
            <given-names>P.C.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pardo</surname>
            ,
            <given-names>T.A.S.:</given-names>
          </string-name>
          <article-title>Multi-document summarization using semantic discourse models</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>56</volume>
          (
          <year>2016</year>
          )
          <volume>57</volume>
          {
          <fpage>64</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Nobrega</surname>
            ,
            <given-names>F.A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pardo</surname>
            ,
            <given-names>T.A.S.:</given-names>
          </string-name>
          <article-title>Update summarization for portuguese</article-title>
          .
          <source>In: Proceedings of the 6th Brazilian Conference on Intelligent Systems</source>
          (To appear). (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Dias</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pardo</surname>
            ,
            <given-names>T.A.S.:</given-names>
          </string-name>
          <article-title>A discursive grid approach to model local coherence in multi-document summaries</article-title>
          .
          <source>In: Proceedings of the 16th Annual SIGdial Meeting on Discourse and Dialogue</source>
          . (
          <year>2015</year>
          )
          <volume>60</volume>
          {
          <fpage>67</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Camargo</surname>
          </string-name>
          , R.T.,
          <string-name>
            <surname>Di</surname>
            <given-names>Felippo</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.A.S.:</surname>
          </string-name>
          <article-title>On strategies of human multidocument summarization</article-title>
          .
          <source>In: Proceedings of the 10th Brazilian Symposium in Information and Human Language Technology</source>
          . (
          <year>2015</year>
          )
          <volume>141</volume>
          {
          <fpage>150</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Maziero</surname>
            ,
            <given-names>E.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirst</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pardo</surname>
            ,
            <given-names>T.A.S.</given-names>
          </string-name>
          :
          <article-title>Semi-supervised never-ending learning in rhetorical relation identi cation</article-title>
          .
          <source>In: Proceedings of the Recent Advances in Natural Language Processing</source>
          . (
          <year>2015</year>
          )
          <volume>436</volume>
          {
          <fpage>442</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Braud</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coavoux</surname>
            , M.,
            <given-names>S gaard</given-names>
          </string-name>
          , A.:
          <article-title>Cross-lingual RST discourse parsing</article-title>
          .
          <source>In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics</source>
          . Volume
          <volume>1</volume>
          . (
          <year>2017</year>
          )
          <volume>292</volume>
          {
          <fpage>304</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Ponti</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korhonen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Event-related features in feedforward neural networks contribute to identifying causal relations in discourse</article-title>
          .
          <source>In: Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics</source>
          . (
          <year>2017</year>
          )
          <volume>25</volume>
          {
          <fpage>30</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <given-names>Di</given-names>
            <surname>Felippo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Nenkova</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Phrase generalization: a corpus study in multidocument abstracts and original news alignments</article-title>
          .
          <source>In: Proceedings of the 10th Linguistic Annotation Workshop</source>
          . (
          <year>2016</year>
          )
          <volume>151</volume>
          {
          <fpage>159</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Cardoso</surname>
            ,
            <given-names>P.C.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pardo</surname>
            ,
            <given-names>T.A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taboada</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>On the contribution of discourse to topic segmentation</article-title>
          .
          <source>In: Proceedings of the 14th Annual Meeting of the Special Interest Group on Discourse and Dialogue</source>
          . (
          <year>2013</year>
          )
          <volume>92</volume>
          {
          <fpage>96</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Maziero</surname>
            ,
            <given-names>E.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castro</surname>
            <given-names>Jorge</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.L.R.</given-names>
            ,
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.A.S.:</surname>
          </string-name>
          <article-title>Revisiting cross-document structure theory for multi-document discourse parsing</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>50</volume>
          (
          <issue>2</issue>
          ) (
          <year>2014</year>
          )
          <volume>297</volume>
          {
          <fpage>314</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Souza</surname>
            ,
            <given-names>J.W.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Di</surname>
            <given-names>Felippo</given-names>
          </string-name>
          , A.:
          <article-title>O corpus cstnews e sua complementaridade temporal</article-title>
          .
          <source>In: PROPOR Workshop on Tools and Resources for Automatically Processing Portuguese and Spanish</source>
          . (
          <year>2014</year>
          )
          <volume>105</volume>
          {
          <fpage>109</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>H.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomes</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Eco and onto.pt: a exible approach for creating a portuguese wordnet automatically</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>48</volume>
          (
          <issue>2</issue>
          ) (
          <year>2014</year>
          )
          <volume>373</volume>
          {
          <fpage>393</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Sarmento</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinto</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabral</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>REPENTINO - a wide-scope gazetteer for entity recognition in portuguese</article-title>
          .
          <source>In: Proceedings of the 7th International Workshop on Computational Processing of the Portuguese Language</source>
          . (
          <year>2006</year>
          )
          <volume>31</volume>
          {
          <fpage>40</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <surname>Carletta</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Assessing agreement on classi cation tasks: the kappa statistic</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>22</volume>
          (
          <issue>2</issue>
          ) (
          <year>1996</year>
          )
          <volume>249</volume>
          {
          <fpage>254</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Barzilay</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapata</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Modeling local coherence: An entity-based approach</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>34</volume>
          (
          <issue>1</issue>
          ) (
          <year>2008</year>
          )
          <volume>1</volume>
          {
          <fpage>34</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>