<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Measuring Comparability of Multilingual Corpora Extracted from Wikipedia ∗</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pablo Gamallo Otero</string-name>
          <email>pablo.gamallo@usc.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Issac Gonz´alez L´opez</string-name>
          <email>isaacjgonzalez@cilenis.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro de Investigaci ́on en Tecnolox ́ıas, da Informaci ́on (CITIUS), Universidade de Santiago de Compostela</institution>
          ,
          <addr-line>Galiza</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Cilenis S.L., Language Engineering Solutions</institution>
          ,
          <addr-line>Santiago de Compostela, Galiza</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <fpage>8</fpage>
      <lpage>13</lpage>
      <abstract>
        <p>Comparable corpora can be used for many linguistic tasks such as bilingual lexicon extraction. By improving the quality of comparable corpora, we improve the quality of the extraction. This article describes some strategies to build comparable corpora from Wikipedia and proposes a measure of comparability. Experiments were performed on Portuguese, Spanish, and English Wikipedia.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Wikipedia is a free, multilingual, and
collaborative encyclopedia containing entries
(called “articles”) for almost 300 languages
(281 in July 2011). English is the more
representative one with about 3 million
articles. However, Wikipedia is not a parallel
corpus as their articles are not translations
from one language into another. Many works
have been published in the last years
focused on its use and exploitation for
multilingual tasks in natural language processing:
extraction of bilingual dictionaries
        <xref ref-type="bibr" rid="ref2 ref3 ref7 ref8">(Yu y
Tsujii, 2009; Tyers y Pieanaar, 2008)</xref>
        , alignment
and machine translation
        <xref ref-type="bibr" rid="ref6">(Adafre y de Rijke,
2006; Tom´as, Bataller, y Casacuberta, 2001)</xref>
        ,
multilingual information retrieval
        <xref ref-type="bibr" rid="ref3 ref7">(Pottast,
Stein, y Anderka, 2008)</xref>
        . There also exists
      </p>
      <p>
        In addition, multilingual articles of
Wikipedia have been used as a source to build
comparable corpora (Gamallo y Gonz´alez,
2010). The EAGLES - Expert Advisory
Group on Language Engineering Standards
Guidelines (see http://www.ilc.pi.cnr.
it/EAGLES96/browse.html) defines a
“comparable corpus” as one which selects
similar texts in more than one language or
variety. One of the main advantages of
comparable corpora is their versatility to be used
in many linguistic tasks
        <xref ref-type="bibr" rid="ref1">(Maia, 2003)</xref>
        , like
bilingual lexicon extraction
        <xref ref-type="bibr" rid="ref3 ref3 ref5 ref7 ref7">(Gamallo y
Pichel, 2008; Saralegui, Vicente, y Gurrutxaga,
2008)</xref>
        , information retrieval, and knowledge
engineering. Besides, they can also be used
as training corpus to improve statistic
machine learning systems, in particular when
parallel corpora are scarce for a given pair of
languages. Another advantage concerns their
availability. In contrast with parallel corpora,
which require (not always available)
translated texts, comparable corpora are easily
retrieved from the web. Among the different
web sources of comparable corpora,
Wikipedia is likely the largest repository of
similar texts in many languages. We only require
the appropriate computational tools to make
them comparable.
      </p>
      <p>
        By taking into account multilingual
potentialities of Wikipedia, our main objective
is to define a method to measure the
similarity (or degree of comparability) of
different comparable corpora built from
Wikipedia. For this purpose, first we describe some
strategies to extract monolingual corpora in
Portuguese, Spanish, and English from
Wikipedia, by making use of some categories
(“Archaeology”, “Biology”, “Physics”, etc.)
to make them comparable according to a
specific topic. These strategies were
described in detail in (Gamallo y Gonz´alez, 2010).
Then, we propose a measure of
comparability to verify whether the corpora are lowly
or highly comparable. For many extraction
tasks, such as bilingual lexicon extraction,
using highly comparable corpora often leads
to better results. There are some works
proposing comparability measures between
monolingual corpora
        <xref ref-type="bibr" rid="ref4">(Li y Gaussier, 2010;
Saralegui y Alegria, 2007)</xref>
        , based on the use of
existing bilingual dictionaries. However,
instead of exploiting dictionaries to compute the
comparability degree, we take advantage of
the translation equivalents inserted in
Wikipedia by means of interlanguage links.
      </p>
      <p>This paper is organized as follows. Section
2 introduces two strategies to build
comparable corpora from Wikipedia. Next, in Section
3, we propose some comparability measures.
Then, Section 4 describe some experiments
performed in order to measure the
comparability between different corpora built using
the strategies defined in Sec. 2 . The last
section discusses future tasks that will be
implemented in order to extend and improve our
tools.
2.</p>
      <p>Two strategies to Build
Wikipedia-Based Comparable
Corpora</p>
      <p>The input of our strategies is
CorpusPedia1, a friendly and easy-to-use XML
structure, generated from Wikipedia dump files.
In CorpusPedia, all the internal links found
in the text are put in a vocabulary list
identified with the tag links. In the same way, all
the categories (or topics) used to classify each
article are inserted in the tag category. In
addition, there is a tag called translations which
codifies a list of interlanguage links (i.e., links
to the same articles in other languages) found
in each article. Categories and translations
are very useful features to build comparable
corpora. Given these features, we developed
two strategies aimed to extract corpora with
different degrees of comparability.</p>
    </sec>
    <sec id="sec-2">
      <title>Not-Aligned Corpus This strategy ex</title>
      <p>tracts those articles in two languages
having in common the same topic,
where the topic is represented by a
category and its translation (for instance,
the English-Spanish pair
“ArchaeologyArqueolog´ıa”). It results in a not-aligned
comparable corpus, consisting of texts
in two languages. We called it
“notaligned” because the version of an article
in one language may have not its
corresponding version in the other language.
Aligned Corpus The goal is to extract
pairs of bilingual articles related by
interlanguage links if, at least, one of both
contains a required category. It results
in a comparable corpus that is aligned
article by article.</p>
      <p>In Section 4, we will measure the degree
of comparability of corpora built by means
of these two strategies. Before that, we will
define how to measure comparability between
Wikipedia-based corpora.</p>
      <p>Comparability Measures</p>
      <p>For a comparable corpus C of Wikipedia
articles, constituted for instance by a
Portuguese part Cp and a Spanish part Cs, a
comparability coefficient can be defined on the basis
1The software to build CorpusPedia, as well as
CorpusPedia files for English, French, Spanish,
Portuguese, and Galician, are freely available at http:
//gramatica.usc.es/pln/
of finding, for each Portuguese term tp in the
vocabulary Cpv of Cp, its interlanguage link (or
translation) in the vocabulary Csv of Cs. The
vocabulary of a Wikipedia corpus is the set of
“internal links” found in that corpus. So, the
two corpus parts, Cp and Cs, tend to have a
high degree of comparability if we find many
internal links in Cpv that can be translated (by
means of interlanguage links) into many
internal links in Csv. Let T ransbin(tp, Csv) be a
binary function which returns 1 if the
translation of the Portuguese term tp is found in
the Spanish vocabulary Csv. The binary Dice
coefficient, Dicebin, between two parts of a
comparable corpus C is then defined as:
Dicebin(Cp, Cs) =
2 Ptp∈Cpv T ransbin(tp, Csv)</p>
      <p>|Cpv| + |Csv|</p>
      <p>We consider that it is not necessary to
define the counterpart of the translation
function, since the number of ambiguous terms
is very low in Wikipedia, and most cases of
ambiguity are solved with the so-called
“disambiguated pages”.</p>
      <p>To avoid a bias towards common internal
links, that is, towards those links occurring
in most articles, we define a specific version
of tf idf weight for each term. In particular,
tf idf (tp) is the frequency of term tp in the
Portuguese part of the comparable corpus,
multiplied by its inverse article frequency in
the whole Portuguese Wikipedia. By taking
into account the tf idf of terms, we can
define a weighted measure of comparability. Let
T ranstf idf (tp, Csv) be a function which
returns the smallest value (min) of two tf idf
scores, both tf idf (tp) and tf idf (ts), where
ts is the Spanish translation of tp in the
Spanish part Cs. The weighted Dice coefficient,
Dicetf idf , between two parts of a
comparable corpus C is then defined as follows:
2 Ptp∈Cpv T ranstf idf (tp, Csv)
Dicetf idf (Cp, Cs) = Ptp∈Cpv tf idf(tp) + Pts∈Csv tf idf(ts)</p>
      <p>The experiments described in the next
section will be performed with the two
comparability measures defined here.</p>
      <p>Experiments and Results</p>
      <p>Taking CorpusPedia as input source, we
performed several experiments to build
different comparable corpora for three
language pairs, namely Portuguese-Spanish,</p>
      <p>Portuguese-English, and Spanish-English.
These corpora were built using the two
strategies described in Section 2 and five domain
specific seed terms (in the three languages)
considered as representative of five domain
topics: “Archaeology”, “Linguistics”,
“Physics”, “Biology”, and “Sport”.</p>
      <p>Table 1 shows the (binary and tf idf) Dice
scores obtained from measuring the
comparability degree of 30 different comparable
corpora. For each corpus, the table also shows
the size (in Mb) of its two parts. In
particular, the first column introduces the two
languages of the corpus (pt = Portuguese, sp =
Spanish, en = English) and the type of
strategy (aligned or not aligned) used to build
it. In the second and third columns, we show
the two Dice scores. The forth column shows
the size of the two parts of the corpus, and
the last column contains the two seed terms
employed to generate the corpus. In Table 2,
we show the Dice scores as well as the size of
nine pairs of monolingual corpora randomly
generated from Wikipedia.</p>
      <p>We can observe first that there are
significant differences in terms of comparability
between the Dice scores in Table 1 and those
obtained from the randomly generated
monolingual pairs in Table 2. It follows that
corpora built by means of our strategies (not
aligned and aligned) are actually comparable.
Then, we should note that in the
comparable corpora of Table 1, the Dice scores based
on tf idf are about 70 % higher than those
based on the binary function. By contrast, in
randomly generated corpora (Table 2), there
are no significant differences between Dicebin
and Dicetd idf . It means that our tf idf
makes the Dice similarity score higher if the two
evaluated corpus parts are actually
comparable.</p>
      <p>As it was expected, not-aligned corpora
tend to be larger than the aligned ones.
However, if we just compare the smallest parts
of each corpus, the differences are not very
important: the smallest parts of not-aligned
corpora are only 15 % larger than those of
aligned corpora. This is in accordance with
the fact that aligned corpora are more
balanced in terms of size, since no part is much
larger than the other one. As far the corpus
size is concerned, let us note that, in
average, English parts are clearly larger than the
Spanish ones, which are slightly larger than
the Portuguese ones. In general, English
arCorpora
pt-sp (not aligned)
pt-en (not aligned)
sp-en (not aligned)
pt-sp (aligned)
pt-en (aligned)
sp-en (aligned)
pt-sp (not aligned)
pt-en (not aligned)
sp-en (not aligned)
pt-sp (aligned)
pt-en (aligned)
sp-en (aligned)
pt-sp (not aligned)
pt-en (not aligned)
sp-en (not aligned)
pt-sp (aligned)
pt-en (aligned)
sp-en (aligned)
pt-sp (not aligned)
pt-en (not aligned)
sp-en (not aligned)
pt-sp (aligned)
pt-en (aligned)
sp-en (aligned)
pt-sp (not aligned)
pt-en (not aligned)
sp-en (not aligned)
pt-sp (aligned)
pt-en (aligned)
sp-en (aligned)
pt-sp (not aligned)
pt-en (not aligned)
sp-en (not aligned)
pt-sp (aligned)
pt-en (aligned)
sp-en (aligned)
.068
.041
.090
.179
.127
.181
.078
.054
.074
.140
.128
.150
.200
.123
.270
.237
.178
.220
.130
.102
.068
.197
.186
.213
.083
.026
.047
.175
.189
.206
.111
.069
.109
.185
.161
.194</p>
    </sec>
    <sec id="sec-3">
      <title>Dice</title>
      <p>(bin)</p>
      <p>Dice
(tf-idf )
.086
.067
.140
.199
.140
.226
.129
.136
.170
.214
.194
.257
.374
.287
.403
.390
.348
.387
.227
.193
.129
.328
.308
.294
.148
.085
.136
.266
.334
.290
.192
.153
.195
.279
.264
.290</p>
      <p>Size
(in Mb)
0.6Mb/3.4Mb
0.6Mb/8.4Mb
0.4Mb/8.4Mb
0.4Mb/0.2Mb
0.4Mb/1.1Mb
2.0Mb/2.9Mb
0.8Mb/1.7Mb
0.8Mb/5.1Mb
1.7Mb/5.1Mb
0.6Mb/0.8Mb
0.5Mb/1.2Mb
0.9Mb/1.7Mb
4.4Mb/4.8Mb
4.4Mb/12Mb
4.8Mb/12Mb
3.6Mb/4.7Mb
3.8Mb/11Mb
3.4Mb/7.6Mb
2.4Mb/1.5Mb
2.4Mb/9.4Mb
1.5Mb/9.4Mb
1.6Mb/2.8Mb
1.8Mb/4.5Mb
0.9Mb/1.3Mb
11Mb/35Mb
11Mb/333Mb
35Mb/333Mb
9.7Mb/15Mb
11Mb/20Mb
20Mb/29Mb
3.8Mb/9.3Mb
3.8Mb/73Mb
9.3Mb/73Mb
3.2Mb/4.7Mb
3.5Mb/7.6Mb
6.2Mb/8.5Mb</p>
    </sec>
    <sec id="sec-4">
      <title>Seed terms</title>
      <sec id="sec-4-1">
        <title>Arqueologia, Arqueolog´ıa</title>
        <p>Arqueologia, Archaeology
Arqueolog´ıa, Archaeology
Arqueologia, Arqueolog´ıa
Arqueologia, Archaeology
Arqueolog´ıa, Archaeology</p>
      </sec>
      <sec id="sec-4-2">
        <title>Lingu´ıstica, Lingu¨´ıstica</title>
        <p>Lingu´ıstica, Linguistics
Lingu¨´ıstica, Linguistics
Lingu´ıstica, Lingu¨´ıstica
Lingu´ıstica, Linguistics
Lingu¨´ıstica, Linguistics</p>
      </sec>
      <sec id="sec-4-3">
        <title>F´ısica, F´ısica</title>
        <p>F´ısica, Physics
F´ısica, Physics
F´ısica, F´ısica
F´ısica, Physics
F´ısica, Physics</p>
      </sec>
      <sec id="sec-4-4">
        <title>Biologia, Biolog´ıa</title>
        <p>Biologia, Biology
Biolog´ıa, Biology
Biologia, Biolog´ıa
Biologia, Biology
Biolog´ıa, Biology</p>
      </sec>
      <sec id="sec-4-5">
        <title>Desporto, Deporte</title>
        <p>Desporto, Sport
Deporte, Sport
Desporto, Deporte
Desporto, Sport
Deporte, Sport</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Overall</title>
    </sec>
    <sec id="sec-6">
      <title>Overall</title>
    </sec>
    <sec id="sec-7">
      <title>Overall</title>
    </sec>
    <sec id="sec-8">
      <title>Overall</title>
    </sec>
    <sec id="sec-9">
      <title>Overall</title>
    </sec>
    <sec id="sec-10">
      <title>Overall</title>
      <p>Cuadro 1: Dice similarity between several comparable corpora in Portuguese, Spanish, and
English.</p>
      <p>Corpora
pt-sp1 (random)
pt-en1 (random)
sp-en1 (random)
pt-sp2 (random)
pt-en2 (random)
sp-en2 (random)
pt-sp3 (random)
pt-en3 (random)
sp-en3 (random)</p>
    </sec>
    <sec id="sec-11">
      <title>Dice</title>
      <p>(bin)</p>
      <p>Dice
(tf-idf )
.012
.003
.003
.016
.017
.017
.008
.001
.005
.012
.003
.003
.014
.014
.015
.006
.001
.005</p>
      <p>Size
(in Mb)
2.2Mb/0.9Mb
2.2Mb/0.4Mb
0.9Mb/0.4Mb
1.5Mb/3.0Mb
1.5Mb/42Mb
3.0Mb/42Mb
0.2Mb/0.5Mb
0.2Mb/1.4Mb
0.5Mb/1.4Mb
Cuadro 2: Dice similarity between randomly generated pairs of monolingual corpora.
ticles tend to have more words than Spanish
and Portuguese articles. As it was suggested
by one of the reviewers of the article, one of
the reasons for the difference in size in the
case of aligned corpora is that Spanish and
Portuguese entries seem to be summaries of
the English ones. So, to increase
comparability between an aligned pair of articles, the
longer article could be shortened by
removing those parts which are not present in the
other language, obtaining, this way, a more
comparable pair of articles.</p>
      <p>Finally, as it was expected, aligned
corpora are significantly more comparable (i.e.,
higher Dice coefficient) than not-aligned
corpora. In average, Dicetd idf increases 80 % the
comparability of aligned-corpora with regard
to not-aligned ones. So, considering that
aligned corpora only decreases 15 % in size in
relation to not-aligned corpora, we can
conclude that the aligned strategy seems to be
more appropriate to build comparable corpora
from Wikipedia.</p>
      <p>Conclusions and Future Work</p>
      <p>The emergence of multilingual resources,
such a Wikipedia, makes it possible to
design new methods and strategies to compile
corpus from the web, methods that are
more efficient and powerful than the
traditional ones. In particular, the semi-structured
information underlying Wikipedia turns out
to be very useful to build comparable
corpora. In this article, we proposed two strategies
to build comparable corpora from Wikipedia
and a way to measure their degree of
comparability. The experiments led us to
conclude that corpora aligned article by article are
more comparable than not aligned corpora.</p>
      <p>Besides, they consist of two balanced corpus
parts in terms of size. Finally, they are not
much smaller than not aligned corpora.</p>
      <p>
        In future work, we will be focused on how
to improve the strategies to build
comparable corpora by extending coverage (more
articles) without losing comparability. For this
purpose, we will test and evaluate
techniques to expand categories using a list of
similar terms identified as hyponyms or
cohyponyms of the source category. In order to
find hyponyms and co-hyponyms of a term, it
will be required to build an ontology of
categories using the semi-structured information
of Wikipedia
        <xref ref-type="bibr" rid="ref2 ref8">(Chernov et al., 2006; Ponzetto
y Navigli, 2009; de Melo y Weikum, 2010)</xref>
        . On
the other hand, we will evaluate
comparability in an indirect way. In particular, we will
use the generated corpora on tasks requiring
comparable corpora as input (e.g., bilingual
lexicon extraction). The better the extracted
lexicon, the more comparable the input
corpus should be. Finally, we believe that our
method for aligning pairs of articles could be
useful for related tasks, such as Wikipedia
infoboxes alignment in different languagues
        <xref ref-type="bibr" rid="ref2 ref8">(Adar, Skinner, y Weld, 2009)</xref>
        .
      </p>
      <p>Bibliograf´ıa
Adafre, S.F. y M. de Rijke. 2006. Finding
similar sentences across multiple languages
in wikipedia. En 11th Conference of the
European Chapter of the Association for</p>
      <p>Computational Linguistics, p´aginas 62–69.</p>
      <p>Adar, Eytan, Michael Skinner, y Daniel S.</p>
      <p>Weld. 2009. Information arbitrage across
multi-lingual wikipedia. En Second ACM
International Conference on Web Search
and Data Mining , WSDM.</p>
      <p>Chernov, Sergey, Tereza Iofciu, Wolfgang</p>
      <p>Nejdl, y Xuan Zhou. 2006. Extracting
semantic relationships between wikipedia
categories. En SemWiki2006 - From Wiki
to Semantics, Budva, Montenegro.
de Melo, Gerard y Gerhard Weikum. 2010.</p>
      <p>Menta: inducing multilingual taxonomies
from wikipedia. En Proceedings of the
19th ACM international conference on
Information and knowledge management,</p>
      <p>CIKM ’10, p´aginas 1099–1108.</p>
      <p>Filatova, Elena. 2009. Directions for
Exploiting Asymmetries in Multilingual
Wikipedia. En CLEAWS3, p´aginas 30–37,
Colorado.</p>
      <p>Gamallo, Pablo y Isaac Gonz´alez. 2010.
Wikipedia as a multilingual source of
comparable corpora. En LREC 2010 Workshop
on Building and Using Comparable
Corpora, p´aginas 19–26, Valeta, Malta.</p>
      <p>Gamallo, Pablo y Jos´e Ramom Pichel.</p>
      <p>2008. Learning Spanish-Galician
Translation Equivalents Using a Comparable
Corpus and a Bilingual Dictionary. LNCS,
4919:413–423.</p>
      <p>Li, Bo y Eric Gaussier. 2010. Improving
corpus comparability for bilingual lexicon
extraction from comparable corpora. En
20th International Conference on
Computational Linguistics (COLING 2010,
p´aginas 644–652.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Maia</surname>
          </string-name>
          , Belinda.
          <year>2003</year>
          .
          <article-title>What Are Comparable Corpora</article-title>
          .
          <source>En Workshop on Multilingual Corpora: Linguistic Requirements and Technical Perspectives</source>
          , p´aginas 27-34, Lancaster, UK.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Ponzetto</surname>
          </string-name>
          , Simone Paolo y Roberto Navigli.
          <year>2009</year>
          .
          <article-title>Large-scale taxonomy mapping for restructuring and integrating wikipedia</article-title>
          .
          <source>En Proceedings of the 21st international jont conference on Artifical intelligence</source>
          , p´
          <year>aginas 2083</year>
          -
          <year>2088</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Pottast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          , y
          <string-name>
            <given-names>M.</given-names>
            <surname>Anderka</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>A wikipedia-based multilingual retrieval model</article-title>
          .
          <source>En Advances in Information Retrieval</source>
          , p´aginas 522-
          <fpage>530</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Saralegui</surname>
            ,
            <given-names>X. y I.</given-names>
          </string-name>
          <string-name>
            <surname>Alegria</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Similitud entre documentos multil´ıngu¨es de car´acter cient´ıfico-t´ecnico en un entorno Web</article-title>
          .
          <source>En Procesamiento del Lenguaje Natural</source>
          , p´
          <fpage>agina</fpage>
          39.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Saralegui</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , I. San Vicente, y
          <string-name>
            <given-names>A.</given-names>
            <surname>Gurrutxaga</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Automatic generation of bilingual lexicons from comparable corpora in a popular science domain</article-title>
          .
          <source>En LREC 2008 Workshop on Building and Using Comparable Corpora.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>Tom´as</article-title>
          , J., J. Bataller, y
          <string-name>
            <given-names>F.</given-names>
            <surname>Casacuberta</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Mining Wikipedia as a Parallel and Comparable Corpus</article-title>
          .
          <source>En Language Forum, volumen 1</source>
          , p´
          <fpage>agina</fpage>
          34.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Tyers</surname>
            ,
            <given-names>M.F. y J.A.</given-names>
          </string-name>
          <string-name>
            <surname>Pieanaar</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Extracting Bilingual Word Pairs from Wikipedia</article-title>
          .
          <source>En LREC</source>
          <year>2008</year>
          , SALTMIL Workshop, Marrakesh, Marocco.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Yu</surname>
          </string-name>
          , Kun y Junichi Tsujii.
          <year>2009</year>
          .
          <article-title>Bilingual dictionary extraction from wikipedia</article-title>
          .
          <source>En Machine Translation Summit XII</source>
          , Ottawa, Canada.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>