<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Building Linked Open Data of the Life Science Dictionary</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yasunori Yamamoto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shoko Kawamoto</string-name>
          <email>shoko@dbcls.rois.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Database Center for Life Science</institution>
          ,
          <addr-line>Bunkyo, Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>There is a growing need for efficient and integrated access to databases provided by diverse institutions. Using a linked data design pattern allows the diverse data on the Internet to be linked effectively and accessed efficiently by computers. In addition, providing a dictionary to translate words into another language in Resource Description Framework (RDF) is useful to cross a language barrier such as English and Japanese when we want to access datasets in multiple languages. Here, we built a Linked Open Dataset of the Life Science Dictionary (LSD) with links to DBpedia. LSD consists of various lexical resources including English-Japanese / Japanese-English dictionaries and a thesaurus using the MeSH vocabulary. The latest version of LSD contains 110 thousand English and 120 thousand Japanese terms. Since we believe that LSD is a useful language resource in the life science domain to process Japanese and English text data seamlessly, linking LSD to DBpedia enables us to find related knowledge more easily and therefore contributes to the life science research community.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>Multi-lingual linked data</kwd>
        <kwd>Dictionary</kwd>
        <kwd>Linked Open Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        To link heterogeneous databases and provide users with access to them in an
integrated manner, publishing datasets following the linked data design pattern [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] has
increasing appeal to database developers and users. It enables us to access raw data
using the World Wide Web approach such as Uniform Resource Identifier (URI) and
Hypertext Transfer Protocol (HTTP). To follow that pattern, we express every datum
using Resource Description Framework (RDF).
      </p>
      <p>
        In addition, there are lots of non-English resources on the Internet, and the number
of non-English RDF datasets is increasing [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For example, the National Diet Library
(NDL) of Japan provides an RDF version of Web NDL Authorities [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this
situation, there are growing needs of cross language RDF resources that can play a role to
link monolingual RDF data sets of different languages [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. One example is DBpedia
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which has made the contents of Wikipedia available in RDF. Wikipedia [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is an
open, collaboratively developed encyclopedia project, and DBpedia is the largest hub
on the Linked Open Data (LOD). However, it is not necessarily reliable as a
translation dictionary in a specific domain. For example, Wikipedia has 149 pages in the
category of "World Health Organization essential medicines" in English, but has only
56 in Japanese [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Life Science Dictionary (LSD) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] consists of various lexical resources including
English-Japanese / Japanese-English dictionaries and a thesaurus using the Medical
Subject Headings (MeSH) vocabulary, the NLM controlled vocabulary thesaurus used
for indexing articles for PubMed. It also contains co-occurring data that show how
often a pair of terms appears in a MEDLINE entry. LSD has been edited and
maintained by the LSD project since 1993. Members of this project are experts in the
domain. In this situation, we built an RDF version of LSD and made a set of links from
LSD words to their corresponding DBpedia titles. We aim at using this dataset as a
complement to DBpedia in the life science domain.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods and Results</title>
      <p>We used the latest version (Mar. 2011) of LSD that contains 110k English and 120k
Japanese terms. Besides, we used the English titles of DBpedia version 3.7 in
NTriples (labels_en.nt.bz2). There are 8,826,375 titles.</p>
      <p>
        We made links from LSD to DBpedia titles by using a series of string match
methods from the simplest (exact match) to more sophisticated ones such as cosine
similarity-based match. For each DBpedia title, our linking process looks for its
corresponding LSD word by applying the following methods in that order: exact match,
Fingerprint Key Collision (FKC) method, bi-gram FKC, tri-gram FKC [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and cosine
similarity-based match [
        <xref ref-type="bibr" rid="ref10 ref9">9,10</xref>
        ]. Once a match is found, the linking process stops working
on that DBpedia title and takes a next one. This means that if the process finds an
exactly literal word in LSD, it does not use any more sophisticated methods. As for
the application of the cosine similarity-based match, we set its threshold relatively
lower (70%) to find prospective terms broadly. On the other hand, to filter out
undesirable lexically approximate matches such as interleukin-1 and interleukin-2, we
made several pattern matching-based filtering rules.
      </p>
      <p>
        As a result, we obtained 81,065 links and represented them using the
skos:exactMatch predicate [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In addition, of 155,888 English terms in LSD that
have their Japanese translations, about a half of them (79,345) have been linked.
Although the number of the DBpedia entries that have links to their corresponding
Japanese WikiPedia entries is 390,994, only 9,816 of them have been linked to the
LSD English terms. This indicates that the coverage of DBpedia as a translation
dictionary in the life science domain is very limited currently.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Discussions and Conclusions</title>
      <p>We built a linked open dataset of LSD with links to DBpedia. This is a first trial and
we found that some exactly matched words are not semantically identical such as
single letter words. For those words, DBpedia often has entries for disambiguation,
and we should take them into consideration. We have published the LSD-LOD along
with its ontology and hope them to be used widely to utilize multilingual resources in</p>
      <p>Acknowledgements. We thank Dr. Shuji Kaneko for permitting us to release LSD
under CC BY-ND 2.0. This work is funded by the Integrated Database Project,
Ministry of Education, Culture, Sports, Science and Technology of Japan and National
Bioscience Database Center (NBDC) of Japan Science and Technology Agency
(JST).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Linked</given-names>
            <surname>Data - Connect Distributed</surname>
          </string-name>
          <article-title>Data across the Web</article-title>
          , http://linkeddata.org/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Gracia</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montiel-Ponsoda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Challenges for the multilingual Web of Data</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          .
          <volume>11</volume>
          ,
          <fpage>63</fpage>
          -
          <lpage>71</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Web</surname>
            <given-names>NDL</given-names>
          </string-name>
          Authorities, http://id.ndl.go.jp/auth/ndla/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>DBpedia-a crystallization point for the web of data</article-title>
          .
          <source>Journal of Web Semantics</source>
          ,
          <volume>7</volume>
          (
          <issue>3</issue>
          ),
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>5. Wikipedia, http://wikipedia.org/</mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>6. http://en.wikipedia.org/wiki/Category:World_Health_Organization_essential_medicines</mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kawamoto</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohtake</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fujita</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takekoshi</surname>
            , Ugawa,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takeuchi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaneko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Life Science Dictionary:
          <article-title>statistical and collocational analyses of life science English</article-title>
          .
          <source>20th IUBMB International Congress of Biochemistry and Molecular Biology and 11th FAOBMB Congress</source>
          ,
          <string-name>
            <surname>Kyoto</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. Clustering In Depth, http://code.google.com/p/google-refine/wiki/ClusteringInDepth</mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>W. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravikumar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fienberg</surname>
            ,
            <given-names>S. E.</given-names>
          </string-name>
          :
          <article-title>A comparison of string distance metrics for name-matching tasks</article-title>
          .
          <source>IJCAI-2003 Workshop on Information Integration on the Web (IIWeb-03)</source>
          ,
          <fpage>73</fpage>
          -
          <lpage>78</lpage>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Okazaki</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsujii</surname>
          </string-name>
          , J.:
          <article-title>Simple and Efficient Algorithm for Approximate Dictionary Matching</article-title>
          .
          <source>Proceedings of the 23rd International Conference on Computational Linguistics (Coling</source>
          <year>2010</year>
          ), Beijing, China,
          <fpage>851</fpage>
          -
          <lpage>859</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Miles</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bechhofer</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>SKOS-Simple Knowledge Organization System Reference, W3C Recommendation (=</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>