<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Seq2RDF: An end-to-end application for deriving Triples from Natural Language Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yue Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tongtao Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhicheng Liang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heng Ji</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Deborah L. McGuinness</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Rensselaer Polytechnic Institute</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present an end-to-end approach that takes unstructured textual input and generates structured output compliant with a given vocabulary. We treat the triples within a given knowledge graph as an independent graph language and propose an encoder-decoder framework with an attention mechanism that leverages knowledge graph embeddings. Our model learns the mapping from natural language text to triple representation in the form of subject-predicate-object using the selected knowledge graph vocabulary. Experiments on three di erent data sets show that we achieve competitive F1-Measures over the baselines using our simple yet e ective approach. A demo video is included.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Converting free text into usable structured knowledge for downstream
applications usually requires expert human curators, or relies on the ability of machines
to accurately parse natural language based on the meanings in the knowledge
graph (KG) vocabulary. Despite many advances in text extraction and
semantic technologies, there is yet to be a simple system that generates RDF triples
from free text given a chosen KG vocabulary in just one step, which we consider
an end-to-end system. We aim to automate the process of translating a natural
language sentence into a structured triple representation de ned in the form of
subject-predicate-object, s-p-o for short, and build an end-to-end model
based on an encoder-decoder architecture that learns the semantic parsing
process from text to triple without tedious feature engineering and intermediate
steps. We evaluate our approach on three di erent datasets and achieve
competitive F1-measures outperforming our proposed baselines, respectively. The
system, data set and demo are publicly available12.
Inspired by the sequence-to-sequence model[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] in recent Neural Machine
Translation, we attempt to use this model to bridge the gap between natural
language and triple representation. We consider a natural language sentence X =
[x1; : : : ; xjXj] as a source sequence, and we aim to map X to an RDF triple
Y = [y1; y2; y3] with regard to s-p-o as a target sequence that is aligned with
1 https://github.com/YueLiu/NeuralTripleTranslation
2 https://youtu.be/ssiQEDF-HHE
a given KG vocabulary set or schema. Given DBpedia for example, we take a
large amount of existing triples from DBpedia as ground truth facts for training.
Our model learns how to form a compliant triple with appropriate terms in the
existing vocabulary. Furthermore, the architecture of the decoder enables the
model to capture the di erences, dependencies and constraints when selecting
s-p-o respectively, which makes the model a natural t for this learning task.
      </p>
      <p>Lake George is
at
the southeast base</p>
      <p>of the Adirondack Mountains
Bi-directional LSTM</p>
      <p>Concatenate
Encoder</p>
      <p>&lt;Start_of_Triple&gt;
dbr:Lake_George_(New_York)</p>
      <p>dbr:George_Lake
dbr:Lake_George_(Florida)
dbr:Lake_George_(New_South_Wales)
Decoder
dbo:country
dbo:birthplace
dbo:location
dbo:isPartOf</p>
      <p>
        dbr:Adirondacks
dbr:Adirondack_Mountains
yago:Mountain109359803
dbr:Whiteface_Mountain
As shown in Figure 1, the model consists of an encoder taking in a natural
language sentence as sequence input and a decoder generating the target RDF
triple. The model pursues the maximized conditional probability
p(Y jX) =
3
Y p(yjy&lt;td ; X);
td=1
(1)
Both encoder and decoder are recurrent neural networks3 with Long Short Term
Memory (LSTM) cells. We apply the attention mechanism that forces the model
to learn to focus on speci c parts of the input sequence when decoding, instead
3 We use tf.contrib.seq2seq.sequence loss which is a weighted cross-entropy loss
for a sequence of logits. We concatenate the last hidden output of forward and
backward LSTM networks, the concatenated vector comes with xed dimensions
of relying only on the last hidden state of the encoder. Furthermore, in order
to capture the semantics of the entities and relations within our training data,
we apply domain speci c resources[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to obtain the word embeddings and the
TransE model[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to obtain KG embeddings for entities and relations in the KG.
We use these pre-trained Word embeddings and KG embeddings for entities and
relations to initialize the encoder and decoder embedding matrix, respectively,
and results show that this approach improves the overall performance.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Experiments</title>
      <p>
        Data Sets We ran experiments on two public datasets NYT4[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], ADE5 with
selected vocabularies and a Wiki-DBpedia dataset that is produced by distant
supervision6. For data obtained by distant supervision, the test set is manually
labeled to ensure its quality. Each data set is an annotated corpus with
corresponding triples in the form of either s-p-o or entity mentions and relation
types at the sentence level. Details are available on our GitHub page.
      </p>
      <p>Text Berlin is the capital city of Germany.</p>
      <p>
        Triple dbr:Germany dbo:capital dbr:Berlin
Evaluation Metrics We consider pipeline-based approaches that combine
Entity Linking (EL) and Relation Classi cation (RC) as state of the art. We
propose several baselines with combined outputs from state-of-the-art EL7 and RC
for evaluation. We use F1-measure to evaluate triple generation (an output is
considered correct only if s-p-o are all correct) in comparison with the baselines.
Baselines We implement multiple baselines including a classical supervised
learning using simple Lexical features, a state-of-the-art recurrent neural
network (RNN) approach with LSTM [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and one with a Gate Recurrent Unit
(GRU) variant. Then we evaluate the performance on triple generation with
results combining EL and RC. The hyper-parameters in our model are tuned with
10-fold cross-validation on the training set according to the best F1-scores. We
applied the same settings to the baselines. The details regarding the parameters
and settings are available on our GitHub page for replication purposes.
4
      </p>
    </sec>
    <sec id="sec-3">
      <title>Result Analysis</title>
      <p>We achieve the best F1 Measure of 84.3 on the triple generation from Table 2.
Note that the baseline approaches that we implemented are pipeline-based, and
thus they are very likely to propagate errors to downstream components.
However, our model merges the two di erent tasks of EL and RC into one during the
4 New York Times articles: https://github.com/shanzhenren/CoType
5 Adverse drug events: https://sites.google.com/site/adecorpus
6 http://deepdive.stanford.edu/distant_supervision
7 Stanford, Domain speci c NER
decoding, which composes a major advantage over pipeline-based approaches
that usually apply separate models on EL and RC. The most common errors
are caused by Out of vocabulary and Noise from overlapping relations
in text. As we do not cover all rare entity names or consider multiple triple
situations, these errors are valid in some sense.
F1-Measure F1-Measure F1-Measure
36.8
58.7
59.8
64.2
71.4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>We present an end-end system for translating a natural language sentence to
its triple representation. Our system performs competitively on three di erent
datasets and our assumption on enhancing the model with pre-trained KG
embeddings improves performance across the board. It is easy to replicate our work
and use our system following the demonstration. In the future, we plan to
redesign the decoder and enable the generation of multiple triples per sentence.
Acknowledgement This work was partially supported by the NIEHS Award
0255-0236-4609 / 1U2CES026555-01.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bordes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Usunier</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Duran</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yakhnenko</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Translating embeddings for modeling multi-relational data</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>2787</volume>
          {
          <issue>2795</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mathews</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Exploiting task-oriented resources to learn word embeddings for clinical abbreviation expansion</article-title>
          .
          <source>Proceedings of BioNLP 15</source>
          pp.
          <volume>92</volume>
          {
          <issue>97</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Miwa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bansal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>End-to-end relation extraction using lstms on sequences and tree structures</article-title>
          .
          <source>arXiv preprint arXiv:1601.00770</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voss</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abdelzaher</surname>
          </string-name>
          , T.F.,
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          .: Cotype:
          <article-title>Joint extraction of typed entities and relations with knowledge bases</article-title>
          .
          <source>In: Proceedings of the 26th International Conference on World Wide Web</source>
          . pp.
          <volume>1015</volume>
          {
          <issue>1024</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          :
          <article-title>Sequence to sequence learning with neural networks</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>3104</volume>
          {
          <issue>3112</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>