<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>How to Summarize Big Knowledge Subjects</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ling Zheng</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yehoshua Perl</string-name>
          <email>perl@njit.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>James Geller</string-name>
          <email>james.geller@njit.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>NJIT</institution>
          ,
          <addr-line>Newark, NJ</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-One manifestation of the “Big Knowledge'' challenge is providing automated tools for summarization of ontology content to facilitate user comprehension. An aggregation approach for the automatic identification and display of major subjects covered by an ontology's content is presented. The results show that our methodology is viable in capturing the “big picture” of ontology content.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>II. METHODS</title>
      <p>It is customary to use summaries to obtain an orientation
into Big Data. For example, in a large drug repository, there
may be x Antibiotics, y Beta blockers, etc., but how can we
summarize ABK to obtain an orientation into its content? In
this poster, we concentrate on the orientation aspect of
identifying important subjects in a large ontology.</p>
      <p>
        We have previously developed the theory of Abstraction
Networks [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for summarizing ABK. For SNOMED CT
hierarchies, we have developed a kind of Abstraction
Network called partial area taxonomy (taxonomy for short)
[4]. Each node of a taxonomy represents a unit (group) of
concepts that is named by its root concept. This reflects the
structure of the terminology well. However, many groups are
small (measured by the number of concepts), and one cannot
see the forest for the trees, i.e., one cannot perceive the
summary due to too many small groups. Hence, a more
compact summary capturing mainly the large units is needed.
      </p>
      <p>However, if we remove all units below a given size b,
then a large portion of the knowledge is not accounted for.
To remedy this problem, we can aggregate descendant units
with fewer than b concepts (“small units”) into the closest
ancestor unit with at least b concepts (a “large unit”). By
varying the integer parameter b, we can control the
granularity of the summary, i.e., how large is the smallest
unit in the summary and how many large units are in the
summary. Previously [5], we defined this as an aggregate
taxonomy.</p>
      <p>However, there is another problem due to the structure of
a partial area taxonomy and its dependency on the concepts
where new relationships are introduced. We discovered that
some important subjects disappear (do not appear in the
summary at all) due to the small sizes of their units, in spite
of having many small related descendant units of the same
subject area. For example, the unit of the subject Specimen
from nervous system has only 12 concepts, but many more
descendant concepts belong to this subject area. The reason
is that the unit itself, the root of which has many descendant
concepts, is small due to the fact that some children or
grandchildren have new relationships and thus are not
included in the unit of the root, but introduce their own units.</p>
      <p>To overcome this difficulty, we define an aggregated
weight for each unit. Then the aggregated weight equals the
sum of the size x of the unit itself and the sizes of all its
descendant units smaller than x. In this way, the decision
which “small units” to eliminate from the summary can now
be based on the aggregated weight of the subject root. For
example, the unit Specimen from nervous system has 12
concepts and it does not appear in the aggregate taxonomy
when b&gt;12. However, its aggregated weight is 42, because it
has 22 descendant units with fewer than 12 concepts,
summarizing 30 descendant concepts. Considering the
aggregated weight, the unit Specimen from nervous system
will appear in the aggregate taxonomy as long as b&lt;=42.</p>
      <p>We tested this idea for the Specimen hierarchy of
SNOMED CT. A domain expert MD with extensive
experience in ontologies (G.E.) identified 21 major subjects
for Specimen as a gold standard list. They were mapped to
the closest Specimen concepts. The partial area containing
each such concept is listed in Table 1, followed by its size
and aggregated weight (weight for short). The aggregated
weight is used to decide whether the partial area and its
subject appear in the aggregate taxonomy if the aggregated
weight &gt;= b for various values of b. Thus, small units will be
eliminated from the aggregate taxonomy based on their
aggregated weights rather than their own sizes. In this way,
the unit Specimen from nervous system will not be eliminated
from the summary representation.</p>
    </sec>
    <sec id="sec-2">
      <title>III. RESULTS</title>
      <p>A subject is identified by our methodology if the
corresponding partial area appears in the aggregate taxonomy
for parameter b. The results and the corresponding recall R,
precision P and F values are listed at the bottom of Table 1
for various b values. The optimal aggregate taxonomy (Fig. 1)
for identifying subjects is obtained for the maximum F
value=0.48 for b=25. Twelve out of 21 subjects (R=0.57) are
identified among the 29 partial areas (P=0.41) of the
aggregate taxonomy. The 12 subject partial areas identified
are highlighted in yellow in Fig. 1. We note that both recall
and precision are important. The recall gives the portion of
the original subject list identified, while the precision gives
the ratio of the aggregate taxonomy units which qualify as
important subjects. We optimize F, which combines both of
them symmetrically. Fig. 1 was shown to G.E. who after
inspecting it determined that 13 more units (highlighted in
pink) qualify as important subjects. The original list should
have 21+13=34 subjects, 25 of which were identified. Hence,
for this enhancement R is 0.74 (=25/34), P is 0.86 (=25/29)
and F is 0.79, which improve the original results. Thus, the
methodology was shown successful in automatically
identifying most of the subjects in a sample of ABK, an
imperative BK challenge, by summarizing the important
subjects in a SNOMED CT hierarchy.</p>
      <p>Concept Partial-area
Blood spc Blood spc
Body substance smp Body substance smp
Fluid smp Fluid smp
Bone marrow spc Bone marrow spc
Spc from bone Musculoskeletal smp
Spc from nervous system Spc from nervous system
Dermatological smp Dermatological smp
Device spc Device spc
Spc from digestive system Spc from digestive system
Endocrine smp Endocrine smp
Male genital smp Spc from trunk
Genitourinary smp Spc from trunk
Hair spc Dermatological smp
Musculoskeletal smp Musculoskeletal smp
Spc from skin Dermatological smp
Soft tissue smp Soft tissue smp
Cardiovascular smp Cardiovascular smp
Spc from eye Spc from head and neck structure
Joint smp Musculoskeletal smp
Lesion smp Lesion smp
Stool spc Body substance smp</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Geller</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perl</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.
          <article-title>The Big Knowledge to Use (BK2U) Challenge</article-title>
          . Workshop on Data Science,
          <article-title>Learning and Applications to Biomedical and Health Sciences (DSLA-BHS)</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>V. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaufman</surname>
          </string-name>
          , D. R.,
          <string-name>
            <surname>Arocha</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Conceptual change in the biomedical and health sciences domain</article-title>
          .
          <source>Advances in instructional psychology</source>
          .
          <year>2000</year>
          ;
          <volume>5</volume>
          :
          <fpage>329</fpage>
          -
          <lpage>392</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] [4] [5]
          <string-name>
            <surname>Halper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perl</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , et al.
          <article-title>Abstraction networks for terminologies: Supporting management of "big knowledge".</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>Artif Intell Med</source>
          .
          <year>2015</year>
          ;
          <volume>64</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Min</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , et al.
          <article-title>Structural methodologies for auditing SNOMED</article-title>
          .
          <source>J Biomed Inform</source>
          .
          <year>2007</year>
          ;
          <volume>40</volume>
          (
          <issue>5</issue>
          ):
          <fpage>561</fpage>
          -
          <lpage>581</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>