<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Relationship Selection Task?</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>OpenStax, Rice University</institution>
          ,
          <addr-line>Houston TX 77005</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Stanford University</institution>
          ,
          <addr-line>450 Serra Mall, Stanford, CA 94305</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>An intelligent textbook is a traditional textbook enhanced with a knowledge graph (KG) making it a source of enhanced learning and instruction. The nodes of a KG are key terms in a textbook and the edges are the relationships between the terms. Relationship selection is the process of selecting the most appropriate relationship type between two di erent terms to be incorporated into a KG. We demonstrate a tool that allows learners to select relationships between terms embedded in a textbook sentence. We created this tool as a key component of a scalable infrastructure for KG construction through crowdsourcing of relationships between automatically extracted terms from a textbook. This task has the potential to be exibly adapted to di erent textbooks and content domains. It is also suitable for encouraging relational processing and, we believe that it has instructional value. Therefore, our future work is focused on the pedagogical evaluation of the relationship selection task with students reading from a textbook.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge Graph</kwd>
        <kwd>Intelligent Textbooks</kwd>
        <kwd>Relationship Selection</kwd>
        <kwd>Concept Mapping</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Intelligent Textbooks (ITBs) using Arti cial Intelligence (AI) and knowledge
graphs (KG) allow students to dynamically interact with the textbook content,
increasing their ability to understand concepts, raising engagement, and thereby,
improving academic performance. Initial trials of ITBs that utilize KGs have
been found to improve student grade outcomes by a full letter grade over the
control group that was using a conventional textbook [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        However, the process of constructing KGs that power such ITBs are time
consuming and resource intensive. For example, the Inquire Biology ITB that was
created by author Chaudhri for the popular introductory textbook, Campbell's
Biology [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] required 5 person years from biology subject matter experts for
knowledge engineering a KG for the rst 10 chapters. Scaling this e ort to all
? Supported by the National Science Foundation (NSF).
      </p>
      <p>We thank Abhay Agarwal for deploying the task on AWS and making signi cant
additions to the code documentation.</p>
      <p>Copyright © 2021 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>
        To create a scalable AI-crowdsourcing hybrid infrastructure for KG
construction, we focused on three elements of the overall process. First, to capture the
critical terms from textbook content, we used an adapted version of the BERT
language deep learning model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Second, to identify the relationships between
term-pairs, we created a novel relationship selection task (RST) for
crowdsourcing the relationships. Third, we developed de-noising methods to e ectively fuse
crowdsourced responses while accounting for task di culty and participant
competency. For the purposes of this demo, we highlight the development of the RST.
      </p>
      <p>Relationship selection refers to the process of selecting the relationships
between key terms in a textbook. To illustrate a relationship, consider the sentence
eukaryotic cells contain a nucleus. In this example, the key terms are \eukaryotic
cells" and \nucleus", and the relationship linking these two terms is \contains".
Failure to identify important relationships a ects the quality of a KG and
ultimately limits the performance of the ITB.</p>
      <p>
        The example relationships presented in Fig. 1 can be encoded
computationally with tuples of the form (entity, relationship, entity), which is the standard
structure used in knowledge graphs. The relationships needed for ITBs include
taxonomy-based relationships (e.g., \prokaryotic cells" and \eukaryotic cells" are
both subset of the class of \cells"), meronymic relationships (e.g., "nucleus" is
inside "eukaryotic cell"), event structure relationships (e.g., "Anaphase" is a sub
step of "mitosis") and causal relationships (e.g., "di usion" enables "respiratory
gas exchange") [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ].
      </p>
      <p>
        There has been signi cant interest in the ML community in automatically
identifying such relationships from natural language text. However, automatic
extraction of relationships from text requires massive amounts of training data
and rarely yields the high accuracy needed for an ITB. Therefore, our strategy
was to develop a crowdsourcing task that could not only serve the purpose of
creating the necessary relationship data, but also provide pedagogical bene ts to
student participants who are actively learning the material. To this end, we
leveraged the popular educational task of concept mapping [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]). Concept mapping
is an educational activity wherein a student takes individual concepts,
represented as nodes, and de nes the labeled edges between them. The task is ideal
for present purposes, as the end product aligns with the KGs we ultimately hope
to develop. Moreover, the process of creating the concept maps is believed to be
bene cial for student learning [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Speci cally, concept maps have been shown
to be e ective at promoting relational processing [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] during learning. Hence,
we created a modi ed concept mapping activity, speci cally designed to foster
relational processing.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The development of the Relationship Selection Task</title>
      <p>We describe the key steps in the development of RST: identifying relationship
vocabulary, identifying terms, and developing a crowdsourcing tool.</p>
      <sec id="sec-2-1">
        <title>2.1 Identifying an appropriate set of relationships</title>
        <p>
          Our rst step in this process was to choose an appropriate set of relationships
also known as relationship vocabulary. Our relationship vocabulary is based on
an upper ontology called Component Library or CLIB [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. As CLIB was used
extensively for constructing KB Bio 101 [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], we analyzed the most frequently used
relationships. Such relationships included relations for describing the structure
and function of entities, structure of processes, and causal relationships between
processes. We were also informed by the empirical experience of the e
ectiveness of these relationships in practice as well as more recent work on linguistic
analysis of relations [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. For example, linguistic analysis suggests that some
relationships from the CLIB are confusing, such as agent, object, and base. We
replaced these confusing relationships with a general, but clearly understood
relationship participant. The relationships we currently support include taxonomic
relationships for classes and instances, structural relationships such as has part
and material, spatial relationships such as is inside and is above, functional
relationships such as has function and facilitates, event structure relationships,
such as subevent and next event, and causal relationships such as enables and
prevents. We also allow the possibility that no direct relationship may exist, as
well as opportunities for crowd workers to de ne new relationships.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Automatically extracting terms and creating tasks</title>
        <p>We used automated term extraction to identify all the terms in the textbook. As
the precision and recall of the automated method is not perfect, the biologists
on the team validated the terms. We then parsed the textbook section into
individual sentences and automatically identi ed all term pairs that existed in each
sentence. A sentence that contained N terms would have n2 possible pairings,
with each pairing considered to be a single task. After generating all possible
tasks, we presented them to crowd workers using the tool that we describe next.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Developing the Crowdsourcing Tool</title>
        <p>We developed a tool to guide the user in choosing the correct relationship
between a pair of terms in the context of a sentence. The intended user of this
tool is a student who does not have any formal training in knowledge
engineering. During the development of the tool, we iteratively validated our designs
through rapid prototyping with such users. The user is rst asked to read a
section from the textbook and then to undergo a short training on the
relationships. We designed the training using simple common-sense examples that
new users would nd easy to understand. For example, we explain the is inside
relationship using a visual in which a cat is shown hiding inside a box (Fig. 2).
We developed similar illustrations for all the di erent
relationships supported by the tool. After the training,
the user completes a series of tasks through an
interactive dialog to identify relationships between various
textbook terms. All possible relationships between a
pair of concepts can be extremely large. Some simple
insights make our task tractable, namely: most term Fig. 2. Illustration of the
pairs are not related (i.e., the nal graph is sparse), inside relationship.
the terms that are connected are also likely to co-occur closely in the text, and
that we can group the relationships into families so the user rst chooses a
relationship family before choosing the actual relationship (Fig. 3).</p>
        <p>Fig. 3. Respondents on the Relationship Selection Task rst select the correct family
of relationship, followed by the actual relationship.</p>
        <p>As a concrete example, consider this two-step selection in the dialog shown in
Figure 3 where the user is asked to relate the terms \cytoplasm" and \nucleus".
The user rst chooses the appropriate relationship family for the terms, including
taxonomic, spatial, and component-based relationships. The user further can
select that the terms have no relationship between them, that they are unsure of
the relationship, or that they would like to de ne a new relationship to relate the
terms. In this example, the correct relationship family is a spatial relationship
and clicking on this option takes them to a second set of options to specify
which spatial family relationship is correct. In this dialog, they have an option
to ip the order of the terms to ensure that the chosen relationship applies in the
correct direction. Once they ip the order of terms, they can correctly indicate
that the nucleus is inside cytoplasm.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Future directions</title>
      <p>Using the RST, we successfully crowdsourced relationships between terms in
sections of college level biology and psychology textbooks respectively, and are
currently investigating its pedagogical e cacy.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Porter</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A library of generic concepts for composing knowledge bases</article-title>
          .
          <source>In: Proceedings of the 1st international conference on Knowledge capture</source>
          . pp.
          <volume>14</volume>
          {
          <issue>21</issue>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Chaudhri</surname>
            , V.K., Cheng,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Overtholtzer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roschelle</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spaulding</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greaves</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gunning</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Inquire biology: A textbook that answers questions</article-title>
          .
          <source>AI</source>
          Magazine
          <volume>34</volume>
          (
          <issue>3</issue>
          ),
          <volume>55</volume>
          {
          <fpage>72</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chaudhri</surname>
            ,
            <given-names>V.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dinesh</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inclezan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , EDU, M.:
          <article-title>Three lessons for creating a knowledge base to enable explanation, reasoning and dialog</article-title>
          .
          <source>In: Proceedings of the Second Annual Conference on Advances in Cognitive Systems ACS</source>
          . vol.
          <volume>187</volume>
          , p.
          <fpage>203</fpage>
          .
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chaudhri</surname>
            ,
            <given-names>V.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wessel</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heymans</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Kb bio 101: A challenge for tptp rst-order reasoners</article-title>
          .
          <source>In: CADE-24 Workshop on Knowledge Intensive Automated Reasoning. Citeseer</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gisborne</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Donaldson</surname>
          </string-name>
          , J.:
          <article-title>Thematic roles and events</article-title>
          .
          <source>In: The Oxford Handbook of Event Structure</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Grimaldi</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poston</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karpicke</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          :
          <article-title>How does creating a concept map a ect item-speci c encoding?</article-title>
          <source>Journal of Experimental Psychology: Learning, Memory, and Cognition</source>
          <volume>41</volume>
          (
          <issue>4</issue>
          ),
          <volume>1049</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Nesbit</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adesope</surname>
            ,
            <given-names>O.O.</given-names>
          </string-name>
          :
          <article-title>Learning with concept and knowledge maps: A metaanalysis</article-title>
          .
          <source>Review of educational research 76(3)</source>
          ,
          <volume>413</volume>
          {
          <fpage>448</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Novak</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          , Can~as,
          <string-name>
            <surname>A.J.:</surname>
          </string-name>
          <article-title>The origins of the concept mapping tool and the continuing evolution of the tool</article-title>
          .
          <source>Information visualization 5</source>
          (
          <issue>3</issue>
          ),
          <volume>175</volume>
          {
          <fpage>184</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Reece</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meyers</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Urry</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cain</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wasserman</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Minorsky</surname>
          </string-name>
          , P.V.:
          <source>Campbell Biology Australian and New Zealand Edition</source>
          , vol.
          <volume>10</volume>
          .
          <article-title>Pearson Higher Education AU (</article-title>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>