<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Di cult Relations: Extracting Novel Facts from Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ismini Lourentzou</string-name>
          <email>ismini.lourentzou@ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna Lisa Gentile</string-name>
          <email>annalisa.gentile@ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Gruhl</string-name>
          <email>dgruhl@us.ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jane Fortner</string-name>
          <email>jfortner@us.ibm.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Freemon</string-name>
          <email>mfreemon@us.ibm.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kendra Grande</string-name>
          <email>kgrande@us.ibm.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IBM Research Almaden</institution>
          ,
          <addr-line>CA</addr-line>
          ,
          <country country="US">US</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IBM Watson Health</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Creating, populating, updating and maintaining a knowledge resource requires intense human e ort. Automatic Information Extraction techniques play a crucial role for this task, but many ongoing production systems still require a large component of human annotation. In this work we investigate how to better take advantage of human annotations by performing active learning on multiple IE tasks concurrently, speci cally Relation Extraction and Named Entity Recognition. Our proposed approach adaptively requests annotations for one task or the other depending on the current overall performance of the combined extraction. We show promising results on a small use case extracting relations expressing Adverse Drug Reactions from unannotated sentences.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The task of curating Knowledge Bases (KB) has received substantial attention
in the recent years. To extract facts from text or other unstructured or
semistructured content sources, many methods have been proposed [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], mostly
borrowing ideas from research areas such as Natural Language Processing, Machine
Learning, Statistics etc. If we consider the sole task of KB population, the human
component appears in all phases at various degrees, from manual e orts from
dedicated teams, like WordNet or Cyc, to collaboratively created resources as
in the case of Wikipedia, Wikidata, etc. to more automated e orts where facts
can be (semi-)automatically extracted from various sources and the human can
only be involved at the validation step.
      </p>
      <p>When dealing with highly curated KBs, especially in the medical domain it is
of paramount importance that every novel entity, property or relation added to
the KB is nearly 100% accurate. For this reason resources such as pharmaceutical
KBs, biology KBs (such as genomics data), Information Extraction (IE) systems
for clinical trial data etc. have the requirement that every addition is vetted by at
least one human. We therefore consider the case of a semi-automatic extraction
of novel facts from text, where although the human is ultimately responsible
for the addition of new facts to the KB, they can be e ectively supported by
Information Extraction methods.</p>
      <p>We propose an active learning approach that extracts entities and their
relation from text in parallel. We develop a system that consists of an active learning
pipeline and two neural models for sequence labeling and classi cation to perform
both Named Entity Recognition (NER) and Relation Extraction (RE) tasks. We
bootstrap the model by using small amounts of available previously vetted data
- in a cold start scenario we collect some annotated examples by \observing" the
user in her task. When new text is analyzed, and if a full relation is extracted,
we simply pass it for validation. In cases where no entities or relations can be
identi ed we prompt speci c annotation tasks to the end user to collect targeted
annotations.</p>
      <p>The main contribution of this work is a co-training procedure for NER and
RE that seamlessly employs the Active Learning paradigm in real-world
knowledge curation tasks. We test the approach for the task of extracting Adverse Drug
Reactions (ADRs) from medically relevant text, which involves identifying Drugs
and Symptoms as well as whether a causal relation among the two is expressed
in the text. We show the bene t of performing both NER and RE concurrently
by taking full advantage of the continuous interaction with the human (active
learning paradigm). The method does not rely on any manually engineered
features, nor other Natural Language Processing tools but simply leverages word
and character embeddings, therefore can be easily ported to di erent languages
or text styles requiring only a few initial examples.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related</title>
    </sec>
    <sec id="sec-3">
      <title>Work</title>
      <p>
        The literature on harvesting information from text for the purpose of knowledge
creation is extremely vast [
        <xref ref-type="bibr" rid="ref11 ref13">11, 13</xref>
        ], as well as the literature on speci c methods
to solve NER and RE as individual tasks [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], where RE systems mostly rely on
the assumption that entities have been pre-tagged. Early approaches for joint
NER and RE treat these as two separate tasks in a pipeline and exploit their
interactions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], while integrated approaches have also been explored with a
diverse set of methods [
        <xref ref-type="bibr" rid="ref12 ref7">7, 12</xref>
        ]. More recent approaches exploit neural joint models
[
        <xref ref-type="bibr" rid="ref10 ref5">10, 5</xref>
        ] and have also been applied in biomedical literature for extracting ADRs
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The main drawback of current joint models are that they are either (i)
timeconsuming or (ii) rely on complex structures and (iii) on the availability of large
annotated datasets, an assumption that almost never holds for \di cult" (long
tail) entities and relations. While active learning - which has the advantage of
producing usable models at early stage of training- has been explored for NER
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and RE [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], it has not been investigated, to the best of our knowledge, for joint
extraction. We aim to ll this gap by designing an active learning experiment
for joint NER and RE to extract ADRs from text.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Learning entities and relations with limited annotations</title>
      <p>We treat NER and RE as separate but interconnected tasks: each of them is
solved by a speci c neural model and we design a pipeline to ameliorate the
human annotation process. This choice has several bene ts, including the ability
of concurrently having multiple annotators with di erent levels of expertise on
di erent tasks. A joint model would limit us to a single annotator per example
for both tasks and would require a large pool of training data due to increased
complexity. By separating the two tasks we can leverage the interaction between
the two modules so as to limit the number of required annotations.</p>
      <p>
        The NER model is an LSTM-CNN-CRF combination similar to [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The RE
model is mostly based on our previous work [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] which learns one relation at a
time. Given a sentence s our goal is to identify whether s: (i) contains entities
of interest and (ii) expresses a certain relation r among them. We begin with
a pool of unlabeled sentences and we ask our user to label k examples with
information about entities3. We train our NER on the small batch and use the
(imperfect) model to extract entities from the remaining examples and consider
sentences containing target entities as positive examples of the target relation.
When asked for further annotation for such sentences the human only needs
to correct mistakes rather than produce the full annotation for the relation. We
collect k of such \approved" annotated relations, train the RE and re-iterate this
process. Intuitively, NER does not simply spot all entities of the target class,
but only generates entity candidates which are likely to express a relation. The
RE module assess if the candidates actually express a relation, which provides
feedback to the NER component. This interleaved training paradigm enforces an
agreement between the two modules. Minimizing the disagreement of two
models is a typical method used in co-training [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. We test the proposed method
to extract ADRs from askapatient.com and involve our medical knowledge
curators as humans-in-the-loop. The dataset consists of 1100 (positive and negative)
examples of causal relationships between drugs and adverse drug events.
      </p>
      <p>To select examples for NER annotation, we rank unlabeled instances based on
their likelihood of being good entity candidates, i.e. likelihood of the generated
sequence of tags: 1 max P NER (y1; : : : ; ynjx), which we compute using the
y1;:::;yn
Viterbi algorithm. For passing instances to the RE module, we choose between
the candidates generated by the NER component or backtrack to uncertainty
sampling when the NER cannot nd enough instances. In Fig.1 we show the
results of our method (joint NER+RE) compared to two baselines that
apply the NER and RE modules sequentially. OptimalNER+RE simulates an
optimal NER module by utilizing the gold standard labels of the dataset. Our
method achieves comparable performance as starting with an optimal NER.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this work we propose an co-training active learning pipeline for NER and RE
to extract entities and their relations from text. Our system continuously trains
the models while assisting the knowledge curators in their task of maintaining a
3 Batch size can adjust based on the task and the annotator, here we set it to 100
examples.
medical Knowledge Resource up-to-date. We show promising results on a small
use case to extract Adverse Drug Reactions from unstructured text.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aggarwal</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Mining text data</article-title>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lasko</surname>
            ,
            <given-names>T.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mei</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Denny</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>A study of active learning methods for named entity recognition in clinical text</article-title>
          .
          <source>Journal of biomedical informatics 58</source>
          ,
          <volume>11</volume>
          {
          <fpage>18</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grishman</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>An e cient active learning framework for new relation types</article-title>
          .
          <source>In: IJCNLP</source>
          . pp.
          <volume>692</volume>
          {
          <issue>698</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grishman</surname>
          </string-name>
          , R.:
          <article-title>Improving name tagging by reference resolution and relation detection</article-title>
          .
          <source>In: ACL</source>
          . pp.
          <volume>411</volume>
          {
          <fpage>418</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Katiyar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cardie</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Going out on a limb: Joint extraction of entity mentions and relations without dependency trees</article-title>
          .
          <source>In: ACL</source>
          . vol.
          <volume>1</volume>
          , pp.
          <volume>917</volume>
          {
          <issue>928</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Joint models for extracting adverse drug events from biomedical text</article-title>
          .
          <source>In: IJCAI</source>
          . pp.
          <volume>2838</volume>
          {
          <issue>2844</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
          </string-name>
          , H.:
          <article-title>Incremental joint extraction of entity mentions and relations</article-title>
          .
          <source>In: ACL</source>
          . vol.
          <volume>1</volume>
          , pp.
          <volume>402</volume>
          {
          <issue>412</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lourentzou</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alba</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coden</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gentile</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gruhl</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Mining relations from unstructured content</article-title>
          .
          <source>In: PAKDD</source>
          . pp.
          <volume>363</volume>
          {
          <issue>375</issue>
          (
          <year>2018</year>
          ), https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -93037-4 29
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
          </string-name>
          , E.:
          <article-title>End-to-end sequence labeling via bi-directional lstm-cnns-crf</article-title>
          .
          <source>In: ACL</source>
          . vol.
          <volume>1</volume>
          , pp.
          <volume>1064</volume>
          {
          <issue>1074</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Miwa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bansal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>End-to-end relation extraction using lstms on sequences and tree structures</article-title>
          .
          <source>arXiv preprint arXiv:1601.00770</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Automatic Knowledge Graph Re nement: A Survey of Approaches and Evaluation Methods</article-title>
          .
          <source>SWJ 0</source>
          ,
          <issue>1</issue>
          {
          <issue>0</issue>
          (
          <year>2015</year>
          ). https://doi.org/10.3233/SW-160218
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yih</surname>
          </string-name>
          , W.t.:
          <article-title>Global inference for entity and relation identi cation via a linear programming formulation</article-title>
          . Introduction to statistical relational learning pp.
          <volume>553</volume>
          {
          <issue>580</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Weikum</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <article-title>Ho art</article-title>
          , J.,
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Ten Years of Knowledge Harvesting: Lessons and Challenges</article-title>
          .
          <source>Data Engineering</source>
          <volume>5</volume>
          ,
          <issue>41</issue>
          {
          <fpage>50</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Z.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Semi-supervised learning by disagreement</article-title>
          .
          <source>Knowledge and Information Systems</source>
          <volume>24</volume>
          (
          <issue>3</issue>
          ),
          <volume>415</volume>
          {
          <fpage>439</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>