<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Relation Extraction Using TBL with Distant Supervision</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maengsik Choi</string-name>
          <email>nlpmschoi@kangwon.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harksoo Kim</string-name>
          <email>nlpdrkim@kangwon.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Program of Computer and Communications Engineering, College of IT, Kangwon National University</institution>
          ,
          <country>Republic of Korea</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Supervised machine learning methods have been widely used in relation extraction that finds the relation between two named entities in a sentence. However, their disadvantages are that constructing training data is a cost and time consuming job, and the machine learning system is dependent on the domain of the training data. To overcome these disadvantages, we construct a weakly labeled data set using distant supervision and propose a relation extraction system using a transformation-based learning (TBL) method. This model showed a high F1-measure (86.57%) for the test data collected using distant supervision but a low F1-measure (81.93%) for gold label, due to errors in the training data collected by the distant supervision method.</p>
      </abstract>
      <kwd-group>
        <kwd>Relation extraction</kwd>
        <kwd>Transformation-based learning</kwd>
        <kwd>Distant supervision</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In natural language documents, there are a huge number of relations between named
entities. Automatic extraction of these relations from documents would be highly
beneficial in the data analysis fields such as question answering and social network
analysis. Previous studies on relation extraction [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1,2,3,4</xref>
        ] have been mainly conducted
through supervised learning methods using Automatic Content Extraction Corpus
(ACE Corpus) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. However, recent studies have investigated some methods based on
simple rules rather than on complex algorithms because the supervised learning
method using weakly labeled data generated by distant supervision has made it possible to
use a large amount of data [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6,7,8</xref>
        ]. The distant supervision reduces the construction
cost of training data by automatically generating a large amount of training data. In
this paper, we propose a relation extraction system to generate weakly labeled data
using DBpedia ontology [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and classify the relations between two named entities by
using transformation-based learning (TBL) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>Relation Extraction Method Based on TBL</title>
      <p>As shown in Figure 1, the proposed system consists of three parts: (1) a distant
supervision part that generates weakly labeled data from a relation knowledge base and
articles, (2) a training part that generates a TBL model by using the weakly labeled
data as training data, and (3) an applying part that extracts relations from new articles
using the TBL model.
For distant supervision, we use DBpedia ontology as a knowledge base. The DBpedia
ontology is populated using a rule-based semi-automatic approach that relies on
Wikipedia infoboxes, a set of subject-attribute-value triples that represents a summary of
some unifying aspect that the Wikipedia articles share. As shown in Figure 2, the
proposed system extracts sentences from Wikipedia articles by using the triple
information of Wikipedia infobox.
For example, the second sentence in Figure 2 is extracted from the weakly labeled
data having a “Born” relation between “Madonna” and “Bay City, Michigan” and a
“Residence” relation between “Madonna” and “New York City.”
2.2</p>
      <sec id="sec-2-1">
        <title>Relation Extraction Using TBL</title>
        <p>
          To extract features from the weakly labeled data, the proposed system performs
natural language processing using Apache OpenNLP [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], as follows.
1. Sentences are separated using SentenceDetectorME.
2. Parts of speech (POS) are annotated using Tokenizer and POSTaggerME.
3. Parsing results are converted to dependency trees, and head words are extracted
from the dependency trees.
        </p>
        <p>For TBL, the proposed system uses two kinds of templates; a morpheme-level
template and a syntax-level template. As shown in Figure 3, the morpheme-level
template is constituted by the combination of words and POSs that are present in the left
and right three words of the target named entities based on the POS tagging results.
In Figure 3, the numbers described as “-n” and “+n” mean a left n-th word and right
n-th word, respectively, from the target named entities. Then, “NEC” means the class
name of the target named entity. As shown in Figure 4, the syntax-level template is
constituted by the combination of two target named entities, common head word of
the two named entities, head word (parent node in a dependency tree) of the two
named entities, and dependent word (child node in a dependency tree) of the two
named entities based on dependency parsing results.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation</title>
      <sec id="sec-3-1">
        <title>Data Sets and Experimental Settings</title>
        <p>We used DBpedia ontology 3.9 in order to generate weakly labeled data. Then, we
selected the ‘person’ class as a relation extraction domain. Next, we selected five
attributes highly occurred in the infobox templates of the ‘person’ class. Finally, we
constructed a weakly labeled data set by using distant supervision. Table 1 shows the
distribution of the weakly labeled data set.
For the experiment, 90% of the weakly labeled data were used as training data, and
the remaining 10% were used as test data (hereinafter “weak-label test data”). In
addition, to measure the reliability of the weak-label test data, a gold-label test data set,
which was manually annotated with correct answers after randomly selecting 822
items, was constructed. During manual annotation, 395 noise sentences out of 822
sentences were removed.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Experiment Results</title>
        <p>As shown in Table 2, the proposed system showed high performance with regard to
the weak-label test data but low performance with regard to the gold-label test data.
This is because of noise sentences (i.e., sentences that do not describe the relation
between two named entities) that are included in the training data collected through
the distant supervision method. For example, “Christopher Plummer” and “Academy
Award” have an “Award” relation, and based on this, the extracted sentence (“The
film also features an extensive supporting cast including Amanda Peet, Tim Blake
Nelson, Alexander Siddig, Amr Waked and Christopher Plummer, as well as
Academy Award winners Chris Cooper, William Hurt”) describes the “Award” relation
between “Chris Cooper and William Hurt” and “Academy Award.” The results of
analysis of the gold-label test data showed that there were 395 noise sentences out of
822 sentences.
The proposed system showed high performance for the weak-label test data. The fact
reveals that the proposed system can show better performance for the gold-label test
data only if good quality training data is ensured. In order to solve this problem, a
method that can reduce the noise of the training data collected through distant
supervision is required.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>We proposed a relation extraction system using TBL based on distant supervision. In
the proposed system, transformation rules were extracted based on a morpheme-level
template that reflects the linguistic characteristics around named entities. They were
also extracted based on a syntax-level template that reflects the syntactic dependency
between two named entities. The experiment results indicated that higher performance
could be expected only when the quality of the training data, which were extracted
based on the distant supervision method, is improved. In the future, we will study on a
method to increase the quality of training data collected by distant supervision.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This research was supported by Basic Science Research Program through the National
Research Foundation of Korea(NRF) funded by the Ministry of Education, Science
and Technology(2013R1A1A4A01005074), and was supported by ATC(Advanced
Technology Center) Program “Development of Conversational Q&amp;A Search
Framework Based On Linked Data: Project No. 10048448”. It was also supported by 2014
Research Grant from Kangwon National University(No. C1010876-01-01).
6</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Culotta</surname>
          </string-name>
          , Aron, and Jeffrey Sorensen:
          <article-title>Dependency tree kernels for relation extraction</article-title>
          .
          <source>In: Proceedings of the 42nd Annual Meeting on Association for Computational Linguistics</source>
          .
          <article-title>As-sociation for Computational Linguistics</article-title>
          , (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bunescu</surname>
          </string-name>
          , Razvan C., and
          <string-name>
            <surname>Raymond J. Mooney</surname>
          </string-name>
          :
          <article-title>A shortest path dependency kernel for relation extraction</article-title>
          .
          <source>In: Proceedings of HLT/EMNLP</source>
          , pp.
          <fpage>724</fpage>
          -
          <lpage>731</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Zhang, Min, Jie Zhang, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Su</surname>
          </string-name>
          .
          <article-title>Exploring syntactic features for relation extraction using a convolution tree kernel</article-title>
          .
          <source>In: Proceedings of HLT-NAACL</source>
          , pp.
          <fpage>288</fpage>
          -
          <lpage>295</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. H.</given-names>
            , &amp;
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Q.</surname>
          </string-name>
          :
          <article-title>Tree kernel-based relation extraction with context-sensitive structured parse tree information</article-title>
          .
          <source>In: Proceedings of EMNLP-CoNLL</source>
          pp.
          <fpage>728</fpage>
          -
          <lpage>736</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>NIST</surname>
          </string-name>
          <year>2007</year>
          .
          <article-title>The NIST ACE evaluation website</article-title>
          . http://www.nist.gov/speech/tests/ace
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Mintz</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mike</surname>
          </string-name>
          , et al:
          <article-title>Distant supervision for relation extraction without labeled data</article-title>
          .
          <source>In: Proceedings of the 47th Annual Meeting of the Association for Computational Linguistics</source>
          , pp.
          <fpage>1003</fpage>
          -
          <lpage>1011</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Tseng</surname>
          </string-name>
          ,
          <string-name>
            <surname>Yuen-Hsien</surname>
          </string-name>
          , et al:
          <article-title>Chinese Open Relation Extraction for Knowledge Acquisition</article-title>
          .
          <source>EACL</source>
          <year>2014</year>
          , pp.
          <fpage>12</fpage>
          -
          <lpage>16</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , Yanping, Qinghua Zheng, and Wei Zhang.:
          <article-title>Omni-word Feature and Soft Constraint for Chinese Relation Extraction</article-title>
          .
          <source>In: Proceedings of the 52nd Annual Meeting of the Associ-ation for Computational Linguistics</source>
          , pages
          <fpage>572</fpage>
          -
          <lpage>581</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>9. http://wiki.dbpedia.org/Downloads39</mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ngai</surname>
          </string-name>
          , Grace, and Radu Florian:
          <article-title>Transformation-based learning in the fast lane</article-title>
          .
          <source>In: Proceed-ings of NAACL</source>
          , pp.
          <fpage>40</fpage>
          -
          <lpage>47</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>11. https://opennlp.apache.org/</mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Aprosio</surname>
          </string-name>
          , Alessio Palmero, Claudio Giuliano, and Alberto Lavelli.
          <article-title>: Extending the Coverage of DBpedia Properties using Distant Supervision over Wikipedia</article-title>
          .
          <source>In: NLP-DBPEDIA@ ISWC</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Choi</surname>
          </string-name>
          , Maengsik, and Harksoo Kim.:
          <article-title>Social relation extraction from texts using a supportvector-machine-based dependency trigram kernel</article-title>
          .
          <source>In: Information Processing &amp; Management</source>
          <volume>49</volume>
          .1 pp.
          <fpage>303</fpage>
          -
          <lpage>311</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>