<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Bilingual Summary Corpus for Information Extraction and other Natural Language Processing Applications∗</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Horacio Saggion</string-name>
          <email>horacio.saggion@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandra Szasz</string-name>
          <email>sandra.szasz@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Pompeu Fabra Departament de Tecnologies de la Informaci ́o i les Comunicacions Grupo TALN C/Tanger 122 - Barcelona - 08018</institution>
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>We are grateful to Programa Ram ́on y Cajal from Ministerio de Ciencia e Innovaci ́on</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <fpage>28</fpage>
      <lpage>34</lpage>
      <abstract>
        <p>Cross-lingual information extraction, the task of extracting information from multiple-multilingual sources, can benefit from the availability of a corpus of equivalent documents in various languages. We present a dataset of pairs of summaries in Spanish and English in various application domains and demonstrate its use in information extraction experiments. The dataset has been manually annotated with semantic information.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Cross-lingual information extraction, the
task of extracting information from
multiplemultilingual sources, is a problem which has
received considerably less attention than
extraction from mono-lingual sources. In this
paper, we are concerned with the creation
of a dataset for the development and
evaluation of cross-lingual information extraction
systems. Our corpus is a set of pairs of
summaries in Spanish and English in various
domains. An example of the dataset is shown
below:
17 julio 2006 Isla de Java: un
maremoto de magnitud 7,7 Richter de
magnitud provoca un ’tsunami’ que
caus´o la muerte de 596 personas.</p>
      <p>On 17 July at 03:19:25 p.m. local
time an earthquake measuring 7.7 on
the Richter scale struck offshore
immediately south of West Java at a
depth of 10 km. The areas affected by
the earthquake and resultant tsunami
included the districts of Taskimalaya,
Ciamis, Sukabumi and Garut in West
Java province, Cilacap, Kebumen and
Banyumas in Central Java and the
Gunung Kidul and Bantul districts in the
province of Yogyakarta. No. Deaths
500.</p>
      <p>These elements in the dataset are
nontranslated equivalent summaries which have
been found on the Web. They report on
the same event, in this case an earthquake,
but because they are not translations of one
another, they contain different information,
for example the Spanish summary reports
596 people dead while the English summary
reports 500 people dead. The English
summary is more verbose and contains
information about the time of the event and
various locations affected by the tremor thus
being the two elements complementary. The
dataset can be used for training information
extraction systems, studying
template-totext bilingual generation, and automatic
knowledge modelling.</p>
      <p>This paper gives an overview of the
dataset and initial experiments showing its
potential application. The rest of this paper
is structured as follows: Section 2 we explain
related work and then, in Section 3 we
describe the data set created. After that, in
Section 4 we illustrate how we have used the
corpus and in Section 5 we present our
conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        There are various multilingual datasets in
the machine translation field such as the
Europarl Multilingual Corpus (Koehn, 2005) or
the United Nations Parallel Corpus
        <xref ref-type="bibr" rid="ref4">(Eisele
y Chen, 2010)</xref>
        . Related to the work
presented here are those datasets prepared for
text summarization or information
extraction research. Among them we have
identified the SummBank corpus
        <xref ref-type="bibr" rid="ref13">(Saggion et al.,
2002)</xref>
        created for the study of multi-lingual
summarization in Chinese and English. The
documents in this corpus are translations
of one another and contain announcements
of a local administration. The corpus has
been used in text summarization and
information retrieval experiments
        <xref ref-type="bibr" rid="ref8">(Radev et al.,
2003)</xref>
        . Because of the content and
annotation provided with the dataset, this corpus is
probably less suitable for information
extraction. The CAST corpus
        <xref ref-type="bibr" rid="ref5">(Or˘asan, Mitkov,
y Hasler, 2003)</xref>
        contains newswire texts and
popular science articles in English where
annotations are added to indicate: (i)
essential sentences, (ii) unessential fragments in
sentences, and (iii) links between sentences
when one sentence is needed to understand
another. Because of the particular
annotation schema used, the corpus has potential
applications for sentence compression. The
SumTime-Meteo Corpus
        <xref ref-type="bibr" rid="ref9">(Reiter y Sripada,
2002)</xref>
        provides weather summaries in English
from numerical data and is potentially useful
in data to text generation applications and
information extraction. The Ziff-Davis
corpus contains technical documents in English
and their human created summaries and has
been used in text summarization experiments
(Knight y Marcu, 2000). The dataset of the
Message Understanding Conferences (ARPA,
1993) is probably the best known set for the
development of information extraction
systems.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Data Set Creation and</title>
    </sec>
    <sec id="sec-4">
      <title>Annotation</title>
      <p>The dataset under development is a
comparable corpus of Spanish and English
summaries for four different domains:
aviation accidents, rail accidents, earthquakes,
and terrorist acts; this later subset is still
under development. Further domains will
be incorporated in the future for researchers
interested in evaluating the robustness and
adaptation capabilities of different natural
language processing techniques. In order
to collect the summaries, a keyword search
strategy was used to search for documents on
the Internet using Google Search. Keywords
per domain were defined and used to select
a set of Web pages in Spanish, for example
the keywords “lista de terremotos” could
be used to search for documents in the
earthquake domain. The pages returned
by the search engine were examined to
verify if they actually contained an event
summary and in that case a document was
created for the summary (it is not unusual
to find multiple summaries in a single Web
page). The documents were given names
indicating the type of event and the date
of the event/incident. A set of around 50
summaries per domain in Spanish were
collected in this manner. After this, for
each event summary originally in Spanish
the Internet was searched for an equivalent
English summary (not a translation) using
keywords in English, this time manually
derived from the Spanish summary. For
example if an earthquake event mentioned
a particular date and intensity, then those
elements were used as keywords. Following
this procedure we found equivalent English
summaries for most of the Spanish ones.</p>
      <p>For each domain (event or incident) a
set of semantic components (i.e., slots) were
identified based on intuition and on the
actual data observed in a set of summaries for
the domain. The slots/components making</p>
      <sec id="sec-4-1">
        <title>Information</title>
        <p>Airline
Cause
DateOfAccident
Destination
FlightNumber
NumberOfVictims
Origin
Passenger
Place
Survivors
Tripulation
TypeOfAccident
TypeOfAircraft
Year</p>
      </sec>
      <sec id="sec-4-2">
        <title>Information</title>
        <p>Cause
DateOfAccident
Destination
NumberOfVictims
Origin
Place
Survivors
TypeOfAccident
TypeOfTrain
up the templates which model the domain
are shown in Table 1.</p>
        <p>Corpus examples (pairs of summaries in
the two languages) for the three domains are
shown in Table 2. In order to manually
annotate the summaries with semantic
information, we have used the GATE annotation
framework (Maynard et al., 2002). To
facilitate the annotation process an annotation
schema was used so that in the GATE
Graphical User Interface the target text span to
be annotated can be selected, and annotated
with one valid category from the annotation
schema. The summaries are annotated by
one person, however a second person checks
the annotations for any inconsistency. Note
that because we are dealing with short texts,
the annotation process is easier than that of
annotating a full event report.</p>
        <p>The number of event components found in
the set of summaries is reported in Tables 3,
4 and 5.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Uses of the Corpus</title>
      <p>
        We have started using the corpus in
monolingual as well as in cross-lingual information
extraction. Information extraction is the
mapping of natural language texts (e.g. news
articles, web pages, e-mails) into predefined
structured representations or templates
(Grishman, 1997) such as those we defined in
Table 1. Various techniques have been used
in the development of information extraction
systems including rule-based approaches
relying on robust partial syntactic analysis
        <xref ref-type="bibr" rid="ref2">(Appelt et al., 1993)</xref>
        , Hidden Markov Models
(Leek, 1997; Freigtag y McCallum, 1999),
and a combination of supervised machine
learning
        <xref ref-type="bibr" rid="ref3">(Ciravegna, 2001)</xref>
        and weakly
supervised machine learning
        <xref ref-type="bibr" rid="ref10 ref16">(Yangarber,
2003; Riloff, 1996)</xref>
        . In recent years there
has been an increasing interest in the
application of information extraction for the
“Semantic Web” using ontologies as
knowledge representation formalisms
        <xref ref-type="bibr" rid="ref12 ref6">(Maynard
et al., 2007; Saggion et al., 2007)</xref>
        as well as
on multilingual and cross-lingual
information extraction
        <xref ref-type="bibr" rid="ref12 ref4 ref6 ref7">(Poibeau y Saggion, 2007;
Poibeau, Saggion, y Yangarber, 2008)</xref>
        . It has
been shown that extraction from multiple
Incident
Aviation Accident
Railway Accident
Earthquake
      </p>
      <p>Semantic Schema
Airline; Cause; DateOfAccident; Destination;
FlightNumber; Origin; Passenger; Place; Survivors; Tripulation;
TypeOfAccident; TypeOfAircraft; Victims; Year
Cause; DateOfAccident; Destination; Origin; Passenger;
Survivors; TrainLine; Tripulation; TypeOfAccident;
TypeOfTrain; Victims; Year
City; Country; DateOfEarthquake; Depth; Epicentre;
Fatalities; Homeless; Injured; Magnitude;
OtherPlacesAffected; Province; Region; Survivors; TimeOfEarthquake;
TotalVictims</p>
      <p>Aviation Accident
2009 30 de junio: el vuelo 626 de Yemenia choc´o en cercan´ıas
a Comoras, en el Oc´eano Indico.
2009 June 30 Yemenia Flight 626, an Airbus A310-300 flying
from Sana’a, Yemen to Moroni, Comoros, crashes into the
Indian Ocean with 153 people aboard; one 12-year-old is found
clinging to the wreckage.</p>
      <p>Railway Accident
12 enero 1997 8 muertos y 25 heridos en el descarrilamiento
del tren r´apido Mil´an-Roma en las proximidades de Piacenza
(Italia).</p>
      <p>January 12, 1997 A Pendolino train derails just before a train
station at Piacenza, Italy, killing 8 people and injuring 29
others.</p>
      <p>Earthquake
27 mayo 2006 Isla de Java (Indonesia): un terremoto de
magnitud 6,2 Richter causa al menos 6.234 muertos, 20.000
heridos y 340.000 desplazados.</p>
      <p>
        May 27, 2006 A powerful earthquake struck Indonesia’s
central province of Java early Saturday morning at 0554 Hrs
local time (26 May 2254 Hrs GMT), flattening buildings and
killing over 4900 people.
multilingual sources can lead to improved
semantic indexing
        <xref ref-type="bibr" rid="ref11">(Saggion et al., 2003)</xref>
        when
compared to monolingual or single source
extraction. It has also been shown that
cross-lingual extraction
        <xref ref-type="bibr" rid="ref4 ref6">(Hakkani-Tu¨r, Ji, y
Grishman, 2007)</xref>
        can be used as a filtering
step to improve retrieval in a target language.
4.1
      </p>
      <p>Experiments
Our cross-lingual information extraction
experiments involve the use of a system trained
in a source language to extract
information from translations from another language.</p>
      <p>
        However, to test how useful the dataset is, we
have started with monolingual experiments
per domain and language (e.g., six systems in
total). The systems are a pipeline of text
processing tools followed by a process of token
classification based on Support Vector
Machines (Li et al., 2002). The machine learning
component was adjusted through testing and
evaluation cycles. The text analysis
components are as follows:
• For English: we used default
processors from the GATE system: tokenizer,
parts-of-speech tagger, rule-based
morphological analysis, dictionary lookup,
and named entity recognition and
classification;
Event
Train Accident Spanish
Train Accident English
Aviation Accident Spanish
Aviation Accident English
Earthquake Spanish
Earthquake English
• For Spanish: we used the TreeTagger
software
        <xref ref-type="bibr" rid="ref15">(Schmid, 1995)</xref>
        and our own
trainable named entity recognizer.
      </p>
      <p>Basic linguistic features were used to
train Spanish and English extraction
systems. Both the Spanish and English
systems use for each token to be classified a
context window of five positions containing
the following token features: orthography
(e.g., word capitalization), word root,
partsof-speech, named entity type, and dictionary
(gazetteer lookup) information.</p>
      <p>Because each dataset is relatively small,
we have performed 10-fold cross-validation
experiments reporting here aggregated
precision, recall, and f-score figures. Table 6
presents the results. The English extraction
system performs better than the Spanish
system in the train and aviation accident
domains, while the Spanish system performs
better than the English one in the earthquake
domain. This could be due to the fewer
human annotations in the English earthquakes
compared to the Spanish counterpart. It is
worth noting that the English summaries are
more verbose in this domain making
extraction more difficult. Although the obtained
results are modest, they have to be assessed
taken into account the limited syntactic and
semantic information available from the text
processors. In order to test how the
systems cope with noisy data we have
translated the Spanish summaries into English and
the English summaries into Spanish using
Google Translator and have applied the
information extraction systems to each
translation. In these experiments, for each
translation T in a domain D, the extraction
system is trained with all documents except the
document which is equivalent to T and the
resulting system is applied to summary T.</p>
      <p>
        Evaluation metrics are also computed and
aggregated over all documents. In these
experiments we have obtained in most
domains and languages f-scores over 0.60 which
although not directly comparable with the
mono-lingual results are certainly
encouraging, full details on these experiments can be
found in
        <xref ref-type="bibr" rid="ref14">(Saggion y Szasz, 2011)</xref>
        .
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this paper we have presented an overview
of a dataset with potential interest for
crosslingual natural language processing
applications. To the best of our knowledge this is one
of the few datasets in this field for the pair
Spanish/English. We have shown
information extraction and cross-lingual extraction
as potential applications of the dataset. Our
current work involves the expansion of the
dataset to cover additional domains such as
terrorism and sports. In future work we will
address automatic domain modelling from
summaries and information extraction
induction. We also plan to use the cross-lingual
extraction results to improve mono-lingual
mono-document extraction.
Nation Documents. En Nicoletta
Calzolari (Conference Chair) Khalid Choukri
Bente Maegaard Joseph Mariani Jan
Odijk Stelios Piperidis Mike Rosner,
y Daniel Tapias, editores, Proceedings
of the Seventh conference on
International Language Resources and Evaluation
(LREC’10), Valletta, Malta, may.
European Language Resources Association
(ELRA).</p>
      <p>Freigtag, D. y A. K. McCallum. 1999.</p>
      <p>Information Extraction with HMMs and
Shrinkage. En Proceesings of Workshop
on Machine Learnig for Information
Extraction, p´aginas 31–36.</p>
      <p>Grishman, R. 1997. Information
extraction: Techniques and challenges. En
Maria Teresa Pazienza, editor,
Information Extraction: A Multidisciplinary
Approach to an Emerging Information
Technology, International Summer School
(SCIE-97), volumen 1299 de Lecture
Notes in Computer Science, p´aginas 10–
27, Frascati, Italy, Jul. Springer Verlag.
Hakkani-Tu¨r, D., Heng Ji, y R.
Grishman. 2007. Using Information
Extraction to Improve Cross-lingual Document
Retrieval. En Proceedings of the 1st Intl.
Workshop on Multi-source Multi-lingual
Information Extraction and
Summarization Workshop.</p>
      <p>Knight, K. y M. Marcu. 2000.
Statisticsbased summarization - step one: Sentence
compression. En AAAI/IAAI, p´aginas
703–710, Austin, Texas.</p>
      <p>Koehn, Philipp. 2005. Europarl: A
Parallel Corpus for Statistical Machine
Translation. En Conference Proceedings:
the tenth Machine Translation Summit,
p´aginas 79–86, Phuket, Thailand. AAMT,
AAMT.</p>
      <p>Leek, T.R. 1997. Information
Extraction Using Hidden markov Models.
Informe t´ecnico, University of California,
San Diego, USA.</p>
      <p>Li, Y., H. Zaragoza, R. Herbrich, J.
ShaweTaylor, y J. Kandola. 2002. The
Perceptron Algorithm with Uneven Margins.
En Proceedings of the 9th International
Conference on Machine Learning
(ICML2002), p´aginas 379–386.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Advanced</given-names>
            <surname>Research Projects Agency</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Proceedings of the Fifth Message Understanding Conference (MUC-5)</article-title>
          . Morgan Kaufmann, California.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Appelt</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.R.</given-names>
            <surname>Hobbs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bear</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Israel</surname>
          </string-name>
          , M. Kameyama, y
          <string-name>
            <given-names>M.</given-names>
            <surname>Tyson</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Description of the JV-FASTUS system as used for MUC-5</article-title>
          .
          <source>En Proceedings of the Fourth Message Understanding Conference MUC-5</source>
          , p´aginas 221-
          <fpage>235</fpage>
          . Morgan Kaufmann, California.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Ciravegna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Adaptive information extraction from text by rule induction and generalisation</article-title>
          .
          <source>En Proceedings of the 17th International Joint Conference on Artificial Intelligence (IJCAI</source>
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Eisele</surname>
          </string-name>
          , Andreas y Yu Chen.
          <year>2010</year>
          .
          <article-title>MultiUN: A Multilingual Corpus from United Maynard</article-title>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yankova</surname>
          </string-name>
          , K. Bontcheva, y
          <string-name>
            <given-names>W.</given-names>
            <surname>Peters</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Natural Language Technology for Information Integration in Business Intelligence</article-title>
          . En W. Abramowicz, editor,
          <source>10th International Conference on Business Information Systems</source>
          , Poland,
          <fpage>25</fpage>
          -
          <lpage>27</lpage>
          April.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Or</surname>
            ˘asan,
            <given-names>C.</given-names>
          </string-name>
          , R. Mitkov, y
          <string-name>
            <given-names>L.</given-names>
            <surname>Hasler</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>CAST: a Computer-Aided Summarisation Tool</article-title>
          .
          <source>En Proceedings of EACL2003</source>
          , p´
          <source>aginas 135 - 138</source>
          , Budapest, Hungary, April.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Poibeau</surname>
          </string-name>
          , T. y H. Saggion, editores.
          <year>2007</year>
          .
          <article-title>1st International Workhop on Multi-Source, Multi-Lingual Information Extraction and Summarization</article-title>
          . RANLP, September.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Poibeau</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , H. Saggion, y R. Yangarber, editores.
          <year>2008</year>
          .
          <article-title>2nd International Workhop on Multi-Source, Multi-Lingual Information Extraction and Summarization</article-title>
          . COLING, September.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Radev</surname>
          </string-name>
          , Dragomir Radev, Wai Lam, Arda C Elebi,
          <string-name>
            <surname>Simone</surname>
            <given-names>Teufel</given-names>
          </string-name>
          , John Blitzer, Danyu Liu, Horacio Saggion, Hong Qi, Elliott Drabek, y Johns Hopkins U.
          <year>2003</year>
          .
          <article-title>Evaluation challenges in large-scale document summarization</article-title>
          .
          <source>En In: Proceedings of the 41st Annual Meeting on Association for Computational Linguistics</source>
          , p´aginas 375-
          <fpage>382</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Reiter</surname>
            , E. y
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sripada</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Squibs and discussions: human variation and lexical choice</article-title>
          .
          <source>Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Riloff</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>1996</year>
          .
          <article-title>Automatically generating extraction patterns from untagged text</article-title>
          .
          <source>Proceedings of the Thirteenth Annual Conference on Artificial Intelligence</source>
          , p´aginas 1044-
          <fpage>1049</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Saggion</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maynard</surname>
          </string-name>
          , O. Hamza, y
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wilks</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Multimedia Indexing through Multisource and Multilingual Information Extraction; the MUMIS project</article-title>
          .
          <source>Data and Knowledge Engineering</source>
          ,
          <volume>48</volume>
          :
          <fpage>247</fpage>
          -
          <lpage>264</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Saggion</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Funk</surname>
          </string-name>
          , D. Maynard, y
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Ontology-based information extraction for business applications</article-title>
          .
          <source>En Proceedings of the 6th International Semantic Web Conference (ISWC</source>
          <year>2007</year>
          ), Busan, Korea, November.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Saggion</surname>
            , H.,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Teufel</surname>
            , L. Wai, y
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Strassel</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Developing Infrastructure for the Evaluation of Single and Multi-document Summarization Systems in a Cross-lingual Environment</article-title>
          .
          <source>En 3rd International Conference on Language Resources and Evaluation (LREC</source>
          <year>2002</year>
          ), p´aginas 747-754,
          <string-name>
            <surname>Las</surname>
            <given-names>Palmas</given-names>
          </string-name>
          , Gran Canaria, Spain.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Saggion</surname>
            , H. y
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Szasz</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Multi-domain cross-lingual information extraction from clean and noisy texts</article-title>
          .
          <source>En Proceedings of the Brazilian Symposium on Information and Human Language Technology, Cuiab´a, Brazil</source>
          ,
          <fpage>24</fpage>
          -
          <lpage>26</lpage>
          October. SBC.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <year>1995</year>
          .
          <article-title>Improvements in partof-speech tagging with an application to german</article-title>
          .
          <source>En In Proceedings of the ACL SIGDAT-Workshop</source>
          , p´aginas 47-
          <fpage>50</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Yangarber</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>2003</year>
          .
          <article-title>Counter-Training in Discovery of Semantic Patterns</article-title>
          .
          <source>En Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics (ACL'03).</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>