<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Korean DBpedia and an Approach for Complementing the Korean Wikipedia based on DBpedia</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eun-kyung Kim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthias Weidl</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Key-Sun Choi</string-name>
          <email>kschoi@world.kaist.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Soren Auer</string-name>
          <email>auer@informatik.uni-leipzig.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Semantic Web Research Center, CS Department, KAIST, Korea</institution>
          ,
          <addr-line>305-701</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universitat Leipzig, Department of Computer Science</institution>
          ,
          <addr-line>Johannisgasse 26, D-04103 Leipzig</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the rst part of this paper we report about experiences when applying the DBpedia extraction framework to the Korean Wikipedia. We improved the extraction of non-Latin characters and extended the framework with pluggable internationalization components in order to facilitate the extraction of localized information. With these improvements we almost doubled the amount of extracted triples. We also will present the results of the extraction for Korean. In the second part, we present a conceptual study aimed at understanding the impact of international resource synchronization in DBpedia. In the absence of any information synchronization, each country would construct its own datasets and manage it from its users. Moreover the cooperation across the various countries is adversely a ected.</p>
      </abstract>
      <kwd-group>
        <kwd>Synchronization</kwd>
        <kwd>Wikipedia</kwd>
        <kwd>DBpedia</kwd>
        <kwd>Multi-lingual</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Wikipedia is the largest encyclopedia of mankind and is written collaboratively
by people all around the world. Everybody can access this knowledge as well as
add and edit articles. Right now Wikipedia is available in 260 languages and the
quality of the articles reached a high level [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, Wikipedia only o ers
full-text search for this textual information. For that reason, di erent projects
have been started to convert this information into structured knowledge, which
can be used by Semantic Web technologies to ask sophisticated queries against
Wikipedia. One of these projects is DBpedia [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which stores structured
information in RDF. DBpedia reached a high-quality of the extracted information
and o ers datasets in 91 di erent languages. However, DBpedia lacks su cient
support for non-English languages. For example, DBpedia only extracts data
from non-English articles that have an interlanguage link3, to an English article.
      </p>
    </sec>
    <sec id="sec-2">
      <title>3 http://en.wikipedia.org/wiki/Help:Interlanguage links</title>
      <p>Therefore, data which could be obtained from other articles is not included and
hence cannot be queried. Another problem is the support for non-Latin
characters which among other things results in problems during the extraction process.
Wikipedia language editions with a relatively small number of articles (compared
to the English version) could bene t from an automatic translation and
complementation based on DBpedia. The Korean Wikipedia, for example, was founded
in October 2002 and reached ten thousand articles in June 20054 . Since
February 2010, it has over 130,000 articles and is the 21st largest Wikipedia5. Despite
of this growth, compared to the English version with 3.2 million articles it is still
small.</p>
      <p>The goal of this paper is two-fold: (1) to improve the DBpedia extraction from
non-Latin language editions and (2) to automatically translate information from
the English DBpedia in order to complement the Korean Wikipedia.</p>
      <p>
        The rst aim is to improve the quality of the extraction in particular for the
Korean language and to make it easier for other users to add support for their
native languages. For this reason the DBpedia framework will be extended with a
plug-in system. To query the Korean DBpedia dataset a Virtuoso Server[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] with
a SPARQL endpoint will be installed. The second aim is to translate infoboxes
from the English Wikipedia into Korean and insert it into the Korean Wikipedia
and consequently the Korean DBpedia as well. In recent years, there has been
signi cant research in the area of coordinated control of multi-languages.
Although English has been accepted as a global standard to exchange information
between di erent countries, companies and people, the majority of users are
attracted by projects and web sites if the information is available in their native
language as well. Another important fact is that various Wikipedia editions from
di erent languages can o er more precise information related to a large number
of native speakers of the language, such as countries, cities, people and culture.
For example, information about @¬Ü (transcribing into Kim Jaegyu), a music
conductor of the Republic of Korea, is only available in Korean and Chinese at
the moment.
      </p>
      <p>The paper is structured as follows: In Section 2, we give an overview about
the work on the Korean DBpedia. Section 3 explains the complementation of
Korean Wikipedia using DBpedia. In Section 4, we review related work. Finally,
we discuss future work and conclude in Section 5.
2</p>
      <sec id="sec-2-1">
        <title>Building the Korean DBpedia</title>
        <p>The Korean DBpedia uses the same framework to extract datasets from Wikipedia
as the English version. However, the framework does not have su cient support
for non-English languages, especially for the non-Latin alphabet based languages.</p>
        <p>For testing and development purposes, a dump of the Korean Wikipedia was
loaded into a local MySQL database. The rst step was to use the current
DBpedia extraction framework in order to obtain RDF triples from the database.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 http://stats.wikimedia.org/EN/ChartsWikipediaKO.htm</title>
    </sec>
    <sec id="sec-4">
      <title>5 http://meta.wikimedia.org/wiki/List of Wikipedias</title>
      <p>
        At the beginning the focus was on infoboxes, because infobox templates o er
already semi-structured information. But instead of just extracting articles that
have a corresponding article in the English Wikipedia, like the datasets
provided by DBpedia, all articles have been processed. More information about the
DBpedia framework and the extraction process can be found in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        After the extraction process and the evaluation of the RDF triples,
encoding problems have been discovered and xed. In fact most of these problems
will occur not only in Korean, but for all languages with non-Latin
characters. Wikipedia and DBpedia use the UTF-8 and URL encoding. URI's in
DBpedia have the form http://ko.dbpedia.org/resource/Name, where Name
is taken from the URL of the source Wikipedia article, which has the form
http://ko.wikipedia.org/wiki/Name. This approach has certain advantages.
Further information can be found in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. For example an URI for 4 
(transcribing into Gottingen), as a property in a RDF Triple, would look as follows:
http://ko.dbpedia.org/property/%EA%B4%B4%ED%8C%85%EA%B2%90
      </p>
      <p>This property URI contains the \%" character and thus cannot be serialized
as RDF/XML. For this reason another way has to be found to represent
properties with \%" encoding. The solution in use by DBpedia is to replace \%"
with \ percent ". This resulted in very long and confusing properties which also
produced errors during the extraction process. This has not been a big issue for
the English DBpedia, since it contains very few of those characters. For other
languages this solution is unsuitable. To solve it, di erent solutions have been
discussed. The rst possibility is to just drop the triples that contain such
characters. Of course this is not an applicable solution for languages that mainly
consist of characters which have to be encoded. The second solution was to use
a shorter encoding but with this approach the Wikipedia encoding cannot be
maintained. Another possibility is to use the \%" character and add an
underscore at the end of the string. With this modi cation, the Wikipedia encoding
could be maintained and the RDF/XML can be serialized. At the moment we use
this solution during the extraction process. The use of IRI6's instead of URI's
is another possibility which we will discuss in Section 5. An option has been
added to the framework con guration to control which kind of encoding should
be used.</p>
      <p>Because languages di er in grammar as well, it is obvious that every language
uses its own kind of formats and templates. For example, dates in the English
and in the Korean Wikipedia look as follows:</p>
      <p>For that reason every language has to de ne its own extraction methods.
To realize this, a plug-in system has been added to the DBpedia extraction
framework (see Fig. 1).</p>
    </sec>
    <sec id="sec-5">
      <title>6 http://www.w3.org/International/articles/idn-and-iri/</title>
      <p>The plug-in system consists of two parts: the default part and an optional
part. The default part contains extraction methods for the English Wikipedia
and functions for datatype recognition, for example, currencies and
measurements. This part will always be applied rst, independent from which language
is actually extracted.</p>
      <p>The second part is optional. It will be used automatically if the current
language is not English. The extractor will load the plug-in le for the corresponding
language if it exists. If the extractor did not nd a match in the default part,
it will use the optional part to check the current string for corresponding
templates. The same approach is used for sub-templates which are contained in the
current template.</p>
      <p>After these problems have been resolved and the plug-in system has been
added to the DBpedia extraction framework, the dataset derived from the
Korean Wikipedia infoboxes consists of more than 103,000 resource descriptions
with more than 937,000 RDF triples in total. The old framework only extracted
55,105 resource descriptions with around 485,000 RDF triples. The amount of
triples and templates was almost doubled. The extended framework also
extracted templates which have not been extracted by the old framework at all.
A comparison between some example templates extracted by the old framework
and the extended version can be found in Fig. 2.</p>
      <p>We already started to extend the support for other extractors. Until now the
extractors mentioned in Table 1 are supported for Korean.
3</p>
      <sec id="sec-5-1">
        <title>Complementation of the Korean</title>
      </sec>
      <sec id="sec-5-2">
        <title>DBpedia</title>
      </sec>
      <sec id="sec-5-3">
        <title>Wikipedia using</title>
        <p>The infobox is manually created by authors that create or edit an article. As a
result, many articles have no infoboxes and other articles contain infoboxes which
are not complete. Moreover, even the interlanguage linked articles do not use the
same infobox template or contain di erent amount of information. This is what
we call the imbalance of information. This problem raises an important issue
about multi-lingual access on the Web. In the Korean Wikipedia multi-lingual
access is prevented due to the lack of interlanguage links.</p>
        <p>For example (see Fig. 3), the \Blue House"7 article in English contains the
infobox template Korean Name. However, the \Blue House" could be regarded as
either a Building or a Structure. The \White House" article, which is very similar
to the former, uses the Historic Building infobox template. Furthermore, the
interlanguage linked \­@ " (\Blue House" in Korean) page does not contain
any infobox.</p>
        <p>There are several types of imbalances of infobox information between two
di erent languages. According to the presence of infobox, we classi ed
interlanguages linked pairs of articles into three groups: Short Infobox (S), Distant
Infobox (D), and Missing Infobox (M):
7 The \Blue House" is the executive o ce and o cial residence of the South Korean
head of state, the President of the Republic of Korea.
{ The S-group contains pairs of articles which use the same infobox template
but have a di erent amount of information. For example, an English-written
article and a non-English-written article, which have an interlanguage link
and use the same infobox template, but have a di erent amount of template
attributes.
{ The D-group contains pairs of articles which use di erent infobox
templates. D-group emerges due to the di erent degrees of each Wikipedia
communities' activities. In communities where many editors lively participate,
template categories and formats are more well-de ned and more ne-grained.
For example, philosopher, politician, o ceholder and military person
templates in English are matched just person template in Korean. It appears
not only a Korean Wikipedia but also non-English Wikipedias.
{ The M-group contains pairs of articles where an infobox exists on only one
side.</p>
        <p>As a rst step, we concentrate on S-group and D-group. We tried to enrich
the infobox using dictionary-based term translation. In this work, DBpedia's
English triples in infoboxes are translated into Korean triples. We used bilingual
dictionary which is originally created for English-to-Korean translations from
Wikipedia interlanguage links. Then we added translation patterns for
multiterms using bilingual alignment of collocations. Multi-terms are set of single
terms such as Computer Science.</p>
        <p>
          DBpedia[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is a community which harvests the information from infoboxes.
We translated English DBpedia into Korean. We also developed the Korean
infobox extraction module in Python. This module identi es records contained
in infoboxes, and then parse out the needed elds. A comparison of datasets is
as follows:
{ English Triples in DBpedia: 43,974,018
{ Korean Dataset (Existing Triples/Translated Triples): 354,867/12,915,169
        </p>
        <p>We can get translated Korean triples over 30 times larger than existing
Korean triples. However, a large amount of translated triples has no prede ned
templates in Korean. There may be a need to form a template schema to
organize the ne-grained template structure.</p>
        <p>Thus we have built a template ontology, OntoCloud8, from DBpedia and
Wikipedia, which was released on September 2009, to e ciently build the
template structure. The construction of OntoClolud consists of the following steps:
(1) extracting templates of DBpedia as concepts in an ontology, for example,
the Template:Infobox Person (2) extracting attributes of these templates, for
example, name of Person. These attributes are mapped to properties in
ontology. (3) constructing the concept hierarchy by set inclusion of attributes, for
example, Book is a subclass of Book series.</p>
        <p>{ Book series = fname, title orig, translator, image, image caption, author,
illustrator, cover artist, country, language, genre, publisher, media type, pub date,
english pub date, preceded by, followed byg.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>8 http://swrc.kaist.ac.kr/ontocloud</title>
      <p>For the ontology building, similar types of templates are mapped to a concept.
For example, the Template:infobox baseball player and Template:infobox asian
baseball player describe baseball player. Moreover di erent format of properties
with same meaning should be re ned, for example, `birth place', `birthplace and
age' and `place birth' are mapped to `birthPlace'. OntoCloud v0.2 includes 1,927
classes, 74 object properties and 101 data properties.</p>
      <p>We provided the rst implementation of the DBpedia/Wikipedia multi-lingual
enrichment research.
4</p>
      <sec id="sec-6-1">
        <title>Related Work</title>
        <p>
          DBpedia[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] focuses on extracting information from Wikipedia and make it usable
for the Semantic Web. There are several other projects which have the same goal.
        </p>
        <p>
          The rst project is Yago[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Yago extracts information from Wikipedia and
WordNet. It concentrates on the category system and the Infoboxes of Wikipedia
and combines this information with the taxonomy of WordNet.
        </p>
        <p>
          Another approach is Semantic MediaWiki [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. It is an extension for
MediaWiki, the system used for Wikipedia. This extension allows you to add
structured data into Wikis by using a speci c syntax.
        </p>
        <p>The third project is Freebase, an online database of structured data. Users
can edit this database in a similar way as Wikipedia can be edited.</p>
        <p>
          In the area of cross-language data fusion, another project has been launched
[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The goal is to extract Infobox data from multiple Wikipedia editions and
fusing the extracted data among editions. To increase the quality of articles,
missing data in one edition will be complemented by data from other editions.
If a value exists more than once, the property which is most likely correct will
be selected.
        </p>
        <p>
          The DBpedia ontology has been created manually based on the most
commonly used Infoboxes within Wikipedia. Kylin Ontology Generator[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] is an
autonomous system for re ning such an ontology. To achieve this, the system
combines Wikipedia Infoboxes with WordNet using statistical-relational
learning.
        </p>
        <p>
          The last project is CoreOnto. It is the research project about IT ontology
infrastructure and service technology development 9. There are several
components and solutions for semi-automated ontology construction. One of them is
the CAT2ISA[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] which is a toolkit to extract isa/instanceOf relation from
category structure. It supports not only lexical patterns, but it also analyze other
category links related to the given category link to determine whether the given
category link is isa/instanceOf relation or not.
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>9 CoreOnto http://ontocore.org</title>
      <p>5</p>
      <sec id="sec-7-1">
        <title>Future Work and Conclusion</title>
        <p>Fig. 4. English-Korean Enrichment in DBpedia/Wikipedia</p>
        <p>As future work, it is planned to support more extractors for the Korean
language and improve the quality of the extracted datasets. The support for
Yago and WordNet could be covered by DBpedia-OntoCloud. The OntoCloud
has been linked to WordNet where OntoCloud is an ontology transformed from
the templates of infobox (English). To make the dataset accessible for everybody
a server will be set up at the SWRC10 at KAIST11.</p>
        <p>Because encoding for Korean results in unreadable strings for human beings
the idea has been raised to use IRI's instead of URI's. It is still uncertain if all
tools of the tool chain can handle IRI's. Nevertheless it is already possible to
extract the data with IRI's if desired. These triples contain characters that are
not valid in XML. The number of triples with such characters is only 48 and
can be ignored for now. We also plan to set up a Virtuoso Server to query the
Korean DBpedia over SPARQL.</p>
        <p>The results from the translated Infoboxes should be evaluated precisely and
improved afterwards. After the veri cation of the triples, the Korean DBpedia
10 http://swrc.kaist.ac.kr
11 Korea Advanced Institute of Science and Technology: http://www.kaist.ac.kr
can be updated. This will be helpful to guarantee that the same information can
be recognized in di erent languages.</p>
        <p>The concept shown in Fig. 4 describes the information enrichment process
within the Wikipedia and DBpedia. As described earlier, we rst synchronized
two di erent language versions of DBpedia using translation. After that we can
create infoboxes using translated values. In the nal study of this project, we will
generate new sentences for Wikipedia using newly added data from DBpedia.
These sentences will be published in the appropriate Wikipedia articles. It can
help authors to edit articles and to create infoboxes when a new article is created.
The system can support the authors by suggesting the right template.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Giles</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <article-title>Internet encyclopedias go head to head</article-title>
          .
          <source>Nature</source>
          ,
          <volume>438</volume>
          (
          <year>2005</year>
          ),
          <fpage>900</fpage>
          -
          <lpage>901</lpage>
          ,
          <year>2005</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          , Jens Lehmann, Georgi Kobilarov, Soren Auer, Christian Becker, Richard Cyganiak, Sebastian Hellmann,
          <article-title>DBpedia - A crystallization point for the Web of Data</article-title>
          .
          <source>Journal of Web Semantics</source>
          ,
          <volume>7</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
          ,
          <year>2009</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Orri</given-names>
            <surname>Erling</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ivan</given-names>
            <surname>Mikhailov</surname>
          </string-name>
          .
          <article-title>RDF support in the Virtuoso DBMS</article-title>
          . Volume P-
          <volume>113</volume>
          <source>of GI-Edition - Lecture Notes in Informatics (LNI)</source>
          ,
          <source>ISSN 1617-5468</source>
          , Bonner Kollen Verlag,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Soren Auer and
          <string-name>
            <given-names>Jens</given-names>
            <surname>Lehmann</surname>
          </string-name>
          .
          <article-title>W hat have innsbruck and leipzig in common? Extracting semantics from wiki content</article-title>
          .
          <source>In Enricho Fanconi</source>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Kifer</surname>
          </string-name>
          , and Wolfgang May, editors,
          <source>ESWC</source>
          , volume
          <volume>4519</volume>
          <source>of LNCS</source>
          , pages
          <fpage>503</fpage>
          -
          <lpage>517</lpage>
          .
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Soren Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak and
          <string-name>
            <given-names>Zachary</given-names>
            <surname>Ives</surname>
          </string-name>
          .
          <article-title>DBpedia: A Nucleus for a Web of Open Data</article-title>
          .
          <source>LNCS, ISSN 1611- 3349</source>
          , Springer Berlin/Heidelberg,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Hellmann</surname>
          </string-name>
          , Claus Stadler, Jens Lehmann, Soren Auer.
          <article-title>DBpedia Live Extraction</article-title>
          . Universitat Leipzig (
          <year>2008</year>
          ), http://www.informatik.unileipzig.de/ auer/publication/dbpedia-live-extraction.pdf
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fabian</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Suchanek</surname>
          </string-name>
          , Gjergji Kasneci, Gerhard Weikum.
          <article-title>YAGO: A Large Ontology from Wikipedia and WordNet</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          , Volume
          <volume>6</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>3</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pages</surname>
            <given-names>203</given-names>
          </string-name>
          <source>-217</source>
          ,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Markus</given-names>
            <surname>Kro</surname>
          </string-name>
          <article-title>tzsch, Denny Vrandecic, and Max Volkel. Wikipedia and the Semantic Web { The Missing Links</article-title>
          . In Jakob Voss and Andrew Lih, editors,
          <source>Proceedings of Wikimania</source>
          ,
          <year>2005</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. Max Volkel, Markus Krotzsch, Denny Vrandecic, Heiko Haller, and
          <string-name>
            <given-names>Rudi</given-names>
            <surname>Studer</surname>
          </string-name>
          .
          <source>Semantic Wikipedia. WWW</source>
          <year>2006</year>
          , pages
          <fpage>585</fpage>
          -
          <lpage>594</lpage>
          , ACM,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Eugenio</surname>
            <given-names>Tacchini</given-names>
          </string-name>
          , Andreas Schultz, Christian Bizer:
          <article-title>Experiments with Wikipedia Cross-Language Data Fusion</article-title>
          . 5th Workshop on Scripting and
          <article-title>Development for the Semantic Web (SFSW2009</article-title>
          ), Crete,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Fei</surname>
            <given-names>Wu</given-names>
          </string-name>
          , Daniel S. Weld:
          <article-title>Automatically Re ning the Wikipedia Inforbox Ontology</article-title>
          .
          <source>International World Wide Web Conference</source>
          , Beijing,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <article-title>DongHyun Choi and Key-Sun Choi: Incremental Summarization using Taxonomy, KCAP 2009</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>