<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Rule-Based Relation Extraction System using DBpedia and Syntactic Parsing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kamel Nebhi</string-name>
          <email>kamel.nebhi@unige.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LATL, Department of linguistics University of Geneva Switzerland</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we present a rule-based relation extraction approach which uses DBpedia and linguistic information provided by the syntactic parser Fips. Our goal is twofold: (i) the morpho-syntactic patterns are de ned using the syntactic parser Fips to identify relations between named entities (ii) the RDF triples extracted from DBpedia are used to improve RE task by creating gazetteer relations.</p>
      </abstract>
      <kwd-group>
        <kwd>Kamel Nebhi</kwd>
        <kwd>relation extraction</kwd>
        <kwd>information extraction</kwd>
        <kwd>linked open data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Relation Extraction (RE), de ned as the task of recognizing semantic relations
between pairs of terms in text, has received renewed interest in the \Web of
Data" era, when many billions of RDF triples are actually published on the
Linked Open Data (LOD) cloud1.</p>
      <p>While supervised approaches for RE tasks require much human e ort,
unsupervised approaches need improvement to obtain best results.</p>
      <p>In this paper, we propose a rule-based RE method for French exploiting LOD
such as DBpedia dataset. Furthermore, the system uses linguistic informations
provided by the syntactic parser Fips to de ne rules.</p>
      <p>The main contributions of this paper are twofold: (i) the morpho-syntactic
patterns are de ned using the syntactic parser Fips to identify relations between
named entities (ii) the RDF triples extracted from DBpedia are used to improve
RE task by creating a relation gazetteer.</p>
      <p>This article is structured as follows: section 2 describes some cognate work
on relation extraction; section 3 explains how DBpedia is exploited by our RE
approach; section 4 provides details on the proposed approach; section 5 contains
our experimental results. We conclude and give some perspectives in section 6.
1 The interactive view of the Linked Open Data cloud sets is available here : http:
//lod-cloud.net/
The researches on RE task are divided on three main approaches: supervised
approach, distant supervised approach and unsupervised approach.</p>
      <p>
        Traditional supervised RE has mostly employed kernel-based approaches [
        <xref ref-type="bibr" rid="ref14 ref20">20,
14</xref>
        ]. But recent work on Open Information Extraction [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ] have exploited
lexical and syntactic features to solve the problems of incoherent and uninformative
extractions. Nevertheless, labeling training data require substantial human
effort, leading to signi cant recent interest in distant supervision.
      </p>
      <p>
        Distant supervision was introduced in bioinformatics by [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Since then, the
approach has gained in popularity [
        <xref ref-type="bibr" rid="ref11 ref12 ref4">4, 12, 11</xref>
        ]. The main idea of this approach is
to create its own training data by heuristically matching the contents of relation
repositories to corresponding text. However, we observe that this method leads
to noisy patterns and poor extraction performance when the method is not
directly applied to the text we are working with. To solve this problem, [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ]
use multiple instance learning and multi-instance multi-label learning.
      </p>
      <p>
        The traditional methods in unsupervised RE collect co-occurrences of word
pairs with strings between them, and nally calculate term co-occurrence or
generate surface patterns [
        <xref ref-type="bibr" rid="ref2 ref8">2, 8</xref>
        ]. In addition to surface patterns, [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] generates
dependency patterns to obtain semantic information for concept pairs.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>DBpedia</title>
      <p>
        Linked Open Data refers to data published with a number of best practices based
on W3C standards for publishing and connecting structured data on the Web
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In the past few years, we have assisted to a growth in the LOD publishing
on the Web. Since 2007, many knowledge bases such as DBpedia, Freebase or
YAGO have been integrated in the LOD cloud. In this context, exploiting a very
large-scale information resource can enhance information extraction process [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        For our experiments, we used relations extracted from DBpedia French
databank2. DBpedia [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] de nes LOD URIs for millions of concepts by extracting
structured information from Wikipedia. The French DBpedia dataset contains
data about 100 000 persons, 100 000 locations and 27 020 organizations. Out of
them, we extracted 287 976 relation instances.
      </p>
      <p>Table 1 shows details of relations extracted from the French DBpedia dataset.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Architecture</title>
      <p>
        Our IE system is built on GATE [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to annotate entities in text and to detect
relations between them. Figure 1 describes the architecture of our RE system.
      </p>
      <p>
        To start, articles are submitted to an ontology-based Named Entity
Recognition (NER) pipeline [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] including a basic linguistic pre-processing, a NE
2 http://fr.dbpedia.org/
Relation Type
per:birthPlace
per:birthDate
per:spouse
per:residence
location:country
location:mayor
location:region
org:adminCenter
org:leaderName
org:foundedBy
org:foundingYear
org:foundationPlace
      </p>
      <p>Total</p>
      <p>Size
gazetteer and rules written in JAPE which is a nite state transducer. Semantic
annotation is performed with respect to the DBpedia ontology classes.</p>
      <p>The articles are then parallelly processed by the syntactic parser Fips. We
chose Fips because it produces syntactic structures with (binary) relations
between constituents and because it is robust and accurate enough for our task.
NE received from the NER pipeline are submitted to the relation gazetteer. For
each sentence and the entity pair on it, the relation gazetteer will identify known
relations. Then, we use JAPE patterns based on information, such as functional
relations procuding by Fips and/or POS tags, to extract other binary relations.
4.1</p>
      <sec id="sec-3-1">
        <title>The Fips Parser</title>
        <p>
          Fips [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] is a deep symbolic multilingual parser based on generative grammar
concepts. The parser uses a bottom up parsing algorithm with parallel treatment
of alternatives, as well as heuristics to rank alternatives.
        </p>
        <p>In Fips, each syntactic constituent is represented as a simpli ed X-bar
structure of the form [XP L X R] with no intermediate level, where X is a variable
ranging over the set of lexical categories3. L and R stand for (possibly empty)
lists of, respectively, left and right subconstituents. The lexical level contains
detailed morphosyntactic and semantic information available from the
manuallybuilt lexicons. In the structures returned by the parser, extraposed elements
3 The lexical categories are N(oun), Adj(ective), V(erb), P(reposition), Adv(erb),
C(onjonction), Inter(jection), to which we add the two functional categories T(ense)
and F(unctional).
(interrogative phrases, relative pronouns, clitics, etc.) are coindexed with empty
constituents in canonical positions (i.e., typical argument or adjunct positions).
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Lexical and Syntactic Patterns</title>
        <p>To characterize binary relationships, we use two di erent types of relation
extraction patterns: lexical patterns, built from words and word class information,
and dependency patterns with syntactic information.</p>
        <p>Kirkouk , dans le nord de
l'</p>
        <p>Irak
&lt;city&gt; , PRP DET nord PRP DET &lt;country&gt;</p>
        <p>Nestle a ete creee en 1866 a Vevey par Henry Nestle.
organization \en"
\a"</p>
        <p>\par"
nsubj
date
pobj
location person
pobj</p>
        <p>
          pobj
The data set we use for our experiments is the Quaero broadcast news extended
NE corpus [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. We have written our rules using a training set of 188 articles.
Then, to evaluate the performance of the system we applied the processing
resources on a test corpora of 150 articles4. We manually constructed relation
mentions, according to the DBpedia ontology, only between entity mentions in
the same sentence. Then, we compare the system with the gold standard. We
present the results according to the di erent methods in order to examine the
contribution of the lexical and syntactic patterns.
        </p>
        <p>RE Model</p>
        <p>DBpedia</p>
        <p>DBpedia+Lex Patterns
DBpedia+Lex&amp;Syn Patterns</p>
        <p>Pre%</p>
        <p>Rec%
32.2
49.0
75.5
20.7
30.0
62.1</p>
        <p>F1%
25.2
37.2
68.1
In this paper, we propose a novel approach to RE by exploiting syntactic
information and DBpedia dataset. We show that this approach provides several
advantages and improves RE performance. It was designed mainly to enrich the
annotation produced by an ontology-based information extraction system, but
can be used for other domains, such as improvement of DBpedia.
4 For the evaluation, we use relations between Person, Organization, Location and</p>
        <p>Date.</p>
        <p>In future work, we plan to incorporate an anaphora resolution system to
improve RE task. We also try to apply the method for other kind of relations.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          et al.:
          <article-title>DBpedia { A Crystallization Point for the Web of Data</article-title>
          .
          <source>Journal of Web Semantics: Science, Services and Agents on the WWW</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Banko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al.:
          <article-title>Open information extraction from the Web</article-title>
          .
          <source>In Proceedings of IJCAI</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Linked Data - The Story So Far</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems</source>
          ,
          <volume>5</volume>
          :1{
          <fpage>22</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bunescu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mooney</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Learning to extract relations from the web using minimal supervision</article-title>
          .
          <source>In Proceedings of ACL</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ciravegna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gentile</surname>
            ,
            <given-names>A. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>LODIE: Linked Open Data for Web-scale Information Extraction</article-title>
          . SWAIE:
          <fpage>11</fpage>
          -
          <lpage>22</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Craven</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumlien</surname>
          </string-name>
          , J.:
          <article-title>Constructing biological knowledge bases by extracting information from text sources</article-title>
          .
          <source>In Proceedings of ICISMB</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cunningham</surname>
          </string-name>
          , H. et al.:
          <source>Text Processing with GATE (Version</source>
          <volume>6</volume>
          ). University of She eld,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Davidov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rappoport</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Classi cation of Semantic Relationships between Nominals Using Pattern Clusters</article-title>
          .
          <source>In Proceedings of ACL</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Fader</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soderland</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Etzioni</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Identifying relations for open information extraction</article-title>
          .
          <source>In Proceedings of EMNLP</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Mausam</surname>
            , Schmitz,
            <given-names>M.</given-names>
          </string-name>
          et al.:
          <article-title>Open language learning for information extraction</article-title>
          .
          <source>In Proceedings of EMNLP-CoNLL</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Min</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          et al.:
          <article-title>Distant Supervision for Relation Extraction with an Incomplete Knowledge Base</article-title>
          ,
          <source>In Proceedings of NAACL-HLT</source>
          ,
          <year>2013</year>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mintz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al.:
          <article-title>Distant supervision for relation extraction without labeled data</article-title>
          .
          <source>In Proceedings of ACL-IJCNLP</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Nebhi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Ontology-Based Information Extraction for French Newspaper Articles</article-title>
          .
          <source>In Proceedings of KI</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riccardi</surname>
          </string-name>
          , G.:
          <article-title>Convolution kernels on constituent, dependency and sequential structures for relation extraction</article-title>
          .
          <source>In Proceedings of EMNLP</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Rosset</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          et al.:
          <article-title>Structured Named Entities in two distinct press corpora: Contemporary Broadcast News and Old Newspapers</article-title>
          .
          <source>In Proceedings of LAW VI</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Riedel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Modeling relations and their mentions without labeled text</article-title>
          .
          <source>In Proceedings of ECML</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Surdeanu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al.
          <article-title>: Multi-instance Multi-label Learning for Relation Extraction</article-title>
          .
          <source>In Proceedings of EMNLP-CoNLL</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Wehrli</surname>
          </string-name>
          , E.:
          <article-title>Fips, a deep linguistic multilingual parser</article-title>
          :
          <source>In Proceedings of ACL 2007 Workshop on deep Linguistic Processing</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          et al.:
          <article-title>Unsupervised relation extraction by mining Wikipedia texts using information from the web</article-title>
          .
          <source>In Proceedings of ACL and AFNLP</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Zelenko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aone</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Richardella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Kernel Methods for Relation Extraction</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          , (
          <volume>3</volume>
          ):
          <fpage>1083</fpage>
          -
          <lpage>1106</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>