<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unsupervised Information Extraction using BabelNet and DBpedia</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Amir H. Jadidinejad</string-name>
          <email>amir@jadidi.info</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Islamic Azad University</institution>
          ,
          <addr-line>Qazvin Branch, Qazvin</addr-line>
          ,
          <country country="IR">Iran</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <volume>1019</volume>
      <abstract>
        <p>Using linked data in real world applications is a hot topic in the field of Information Retrieval. In this paper we leveraged two valuable knowledge bases in the task of information extraction. BabelNet is used to automatically recognize and disambiguate concepts in a piece of unstructured text. After extracting all possible concepts, DBpedia is leveraged to reason about the type of each concept using SPARQL.</p>
      </abstract>
      <kwd-group>
        <kwd>Concept Extraction</kwd>
        <kwd>Linked Data</kwd>
        <kwd>BabelNet</kwd>
        <kwd>DBpedia</kwd>
        <kwd>SPARQL</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>BABELNET</title>
      <p>
        BabelNet[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a multilingual lexicalized semantic network and ontology. It was
automatically created by linking the largest multilingual Web encyclopedia – i.e.
Wikipedia1 – to the most popular computational lexicon of the English language – i.e.
WordNet[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It contains an API for programmatic access of 5.5 million concepts and a
multilingual knowledge-rich Word Sense Disambiguation (WSD) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. With the aid of this
API, we can extract all possible concepts in a piece of text. These concepts are linked
to DBpedia, one of the more famous parts of the Linked Data project.
      </p>
    </sec>
    <sec id="sec-2">
      <title>DBPEDIA</title>
      <p>
        DBpedia[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is a project aiming to extract structured content from the information
created as part of the Wikipedia project. This structured information is made available
on Semantic Web formats. DBpedia allows users to query relationships and properties
associated with Wikipedia concepts. In this paper we used SPARQL to query DBpedia.
It's possible to reason about the type of each concept (PER, LOC, ORG, MISC) with
the aid of a classic deductive reasoning using classes and subclasses. For example,
"Settlement" is defined as a subclass of "Place" (although maybe not directly). That means
that all Things that are "Settlements" are also "Places". "Tehran" is a "Settlement", so
it is also a "Place". Using the following query:
1 http://www.wikipedia.org
      </p>
      <p>It's possible to reason about the type of every “?thing” such as:
http://dbpedia.org/resource/Tehran. A similar query is used for LOCATION and
ORGANIZATION.
3</p>
    </sec>
    <sec id="sec-3">
      <title>IMPLEMENTATION DETAILS</title>
      <p>Our proposed solution shows in Figure 1. The input text is passed to "Text2Concept"
module. This module is used "BabelNet" and "Knowledge-rich WSD" algorithm to
recognize a list of concepts. Finally, "Text Reasoner" module reason about the type of each
concept with the aid of DBpedia using a simple deductive reasoning.</p>
      <p>BabelNet</p>
      <p>Knowledge-rich</p>
      <p>WSD
Input Text</p>
      <p>Text2Concept
C1:T1, C2:T2,..., Ck:Tk
C1, C2,..., Ck
Text Reasoner</p>
      <p>DBpedia
#MSM2013 Concept Extraction Challenge Making Sense of Microposts III</p>
      <p>Making Sense of Microposts III</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S. P.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>BabelNet: The Automatic Construction, Evaluation and Application of a Wide-Coverage Multilingual Semantic Network</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>193</volume>
          ,
          <fpage>217</fpage>
          -
          <lpage>250</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>George</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <string-name>
            <surname>Miller</surname>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>WordNet: A Lexical Database for English</article-title>
          .
          <source>Communications of the ACM</source>
          .
          <volume>38</volume>
          ,
          <issue>11</issue>
          ,
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S. P.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Multilingual WSD with Just a Few Lines of Code: the BabelNet API</article-title>
          .
          <source>In Proc. of the 50th Annual Meeting of the Association for Computational Linguistics (ACL</source>
          <year>2012</year>
          ), Jeju, Korea,
          <fpage>67</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>DBpedia - A Crystallization Point for the Web of Data</article-title>
          .
          <source>Journal of Web Semantics: Science, Services and Agents on the World Wide Web</source>
          ,
          <volume>7</volume>
          ,
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>