<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Refining an Ontology by Learning Stakeholder Votes from their Texts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olga Tatarintseva</string-name>
          <email>tatarintseva@znu.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vadim Ermolayev</string-name>
          <email>vadim@ermolayev.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of IT, Zaporozhye National University</institution>
          ,
          <addr-line>66 Zhukovskogo st., 69063, Zaporozhye</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Key terms. KnowledgeEngineeringMethodology</institution>
          ,
          <addr-line>SubjectExpert, SubjectDomain, Metric, Ontology</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Satelliz</institution>
          ,
          <addr-line>158 Lenina st., P.O. Box 317, 69057, Zaporozhye</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>64</fpage>
      <lpage>78</lpage>
      <abstract>
        <p>This paper reports on our experiments evaluating the improvement of OntoElect approach to ontology refinement in the case study with the ICTERI Scope Ontology. OntoElect is based on collecting and assessing the commitment of domain knowledge stakeholders for ontological refinement offerings. We report the improvement with respect to the previous results. Our first experiment evaluates the change in the quality of ontology due to the involvement of domain knowledge stakeholders in semantic annotation of their papers, compared to the previous study in which the annotations were done by knowledge engineers. Our second experiment checks if the result became better after the introduction of the automated term extraction from the full texts of ICTERI papers. Extracted terms are compared to the manual annotations. The results of the experiments verify the proposed ontology changes and are further used for the ICTERI Scope ontology refinement.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>ICTERI Scope ontology</kwd>
        <kwd>ontology engineering</kwd>
        <kwd>domain knowledge stakeholder</kwd>
        <kwd>term mining</kwd>
        <kwd>evaluation</kwd>
        <kwd>refinement</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Maintaining an ontology in its lifecycle that fits all the requirements of the subject
domain stakeholders is a complicated task in ontology engineering which does not
have a complete solution so far. The problem is to a large extent in devising a
methodology for ontology refinement that enables a complete and timely account for those
requirements and maps them to the updated revision of the ontology. One
complication is that the stakeholders who own the requirements need to be committed to
provide their inputs for ontology refinement. Furthermore, those inputs need to be
measured and applied correspondingly to the utility of their contribution and in a
harmonized way to ensure the consistency and validity of result.</p>
      <p>
        This paper reports on the improvement of our OntoElect approach for iterative
ontology refinement [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. The approach has been proposed in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] using the allusion of
elections in which different “ontology offerings” compete for the commitment of the
pool of the relevant domain knowledge stakeholders being the “electorate”. OntoElect
has been basically validated in an experiment reported in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] where the approach was
detailed by offering voting metrics for the ICTERI ontology built and refined
iteratively based on semantic annotation of the pool of papers of ICTERI 2011 conference.
      </p>
      <p>
        The results of our previous experiment suggested several important technical
aspects [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for improving OntoElect methodology as a whole and the accuracy of our
measurements in particular. Some of those aspects have been addressed in the work
reported in this paper.
      </p>
      <p>Firstly, our previous experiment was based substantially on the manual annotation
of papers. A knowledge engineer assigned key terms or suggested missing terms
based on her personal interpretation of the abstract of a paper. By that we mimicked
voting by paper authors while keeping them free of extra annotation effort. The
lowlights of this approach to annotation were that:
 Domain knowledge stakeholders (paper authors) were in fact not involved in the
workflow and therefore not motivated to be committed to the resulting ontology
refinement
 The quality of semantic annotations we obtained has been perceived as fairly low
because (a) done by a knowledge engineer who is not a subject expert with respect
to the annotated paper; (b) the source for this work was just an abstract, not a
paper, and its meaning has been interpreted by a knowledge engineer.</p>
      <p>To overcome those shortcomings we first decided to involve the authors more
actively by requesting that they themselves semantically annotate their submissions to
ICTERI 20121. It has also been considered as promising to refine the approach by
automated extraction of terms from the papers authored by our subject domain
knowledge stakeholders. Here we present the results of our experiments which
checked how these two refinements helped improving the quality and adequacy of the
ICTERI Scope ontology.</p>
      <p>Further, the document corpus used in the previous experiment was fairly small in
size for assuring reliable judgements about the opinion of the stakeholder community.
For improving on that we continued the collection of ICTERI papers which has been
extended by all papers of ICTERI 2012.</p>
      <p>
        We first repeat the previous experiment [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] based however on the document corpus
of ICTERI 2012 papers semantically annotated by their authors. We then focus on
answering the question about annotation quality by: (i) performing automated term
extraction from the full texts of our complete document corpus (ICTERI 2011 and
2012); and (ii) comparing the results of automated term extraction to the outputs of
manual semantic annotation.
1 See http://isrg.kit.znu.edu.ua/icteriwiki/index.php/ICTERI-Terms
      </p>
      <p>The remainder of the paper is structured as follows. Section 2 briefly reviews the
related work in relevant fields. Section 3 outlines the OntoElect approach to ontology
refinement and presents the case study dealing with ICTERI Scope Ontology as well
as the document corpus at our disposal. Section 4 sets up our experiments by
describing the workflow, evaluation metrics, and used tools. Section 5 presents and discusses
the results of our experiments. The paper is further concluded and our plans for the
future work are outlined.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        One of the possible ways to check if a conceptualization of a domain is correct and
complete is to evaluate the model against the interpretation of the meaning of the
representative set of relevant documents. The document corpus will be relevant and
representative if it covers the majority of the views by the domain knowledge
stakeholders. Their interpretations may be collected and further analysed for refining the
ontology using different techniques which may be sought in several areas of research
and development. In this section we briefly outline the related work in the relevant
fields of research and refer to our previous publication [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for a more in-depth and
detailed coverage.
      </p>
      <p>
        One of the popular relevant research areas studying how interpretations are
collected is collaborative or social tagging and annotation. A good survey of the field is
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] where the use of tags for different purposes and associated shortcomings are
analysed. Semantic annotation and tagging approaches further refine social tagging
techniques by offering the collections of terms that are taken from taxonomies,
folksonomies, or thesauri [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Hybrid approaches for collaborative tagging and annotation
aiming at the enrichment of seed knowledge representations by a user community are
reported for example in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        One of the promising approaches focused, besides collecting interpretations or
subjective conceptualizations, on motivating more people to take part in developing or
refining ontologies is offering a game with a purpose to intended users. Following this
approach, ontology development or refinement can be implicitly embedded in a game
software. There ontology elements are created, updated, and validated implicitly in
the background [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Gaming approach has also been tried for evaluating how well
ontological specifications fit to the interpretations of random users (FACTory Game
by Cycorp, http://game.cyc.com/). Several game scenarios have been developed [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
for ontology building and refinement, ontology matching, annotating content using
lightweight ontologies. Those are similar to our OntoElect approach. Both approaches
offer possibilities to identify whenever users start to agree on and share commitment
to certain ontological items.
      </p>
      <p>
        Social and gaming approaches that involve the direct participation of human
stakeholders are complemented by the plethora of research results in automated knowledge
extraction or ontology learning. This strand of research involves the stakeholders
indirectly – through making use of their professional outputs, like authored texts. A
comprehensive survey of the techniques used to learn ontologies from texts is [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In
the second experiment we present in this paper only term extraction using the
TerMine tool [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] has been performed.
      </p>
      <p>
        Yet one more important aspect in developing or refining an ontology is the re-use
of the other ontologies or their most relevant parts to the developed ontology. In this
context the ontology meaning summarization approach [#] makes good sense for
helping an ontology engineer choose the most relevant and valuable parts for re-use.
The approach is based on detecting the “key concepts” of an ontology under analysis
which best characterize its meaning. The key concepts are determined using a
combination of criteria from lexical statistics, taxonomy graph analysis, and popularity
based on a number of hits. Especially in using the popularity and coverage metrics,
this approach coincides well with our approach (OntoElect). OntoElect is however
used not for summarizing but refining an ontology based, among other things, on
assessing the coverage and popularity of the Key Terms. Besides that the mechanisms
of obtaining the measures are different. From the other hand, OntoElect does not yet
consider ontology re-use as one of important mechanisms for refinement. Hence,
combining some features of [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] in OntoElect may be enriching.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>ICTERI Case Study</title>
      <p>
        The idea of OntoElect approach [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] was inspired by public election campaigns. Just as
the leader in a public election campaign gets the major part of the electorate’s
commitment to win, the extent of the domain knowledge stakeholders' commitment hints
about the quality and completeness of the ontology. Following this allusion, the votes
of the domain knowledge stakeholders for alternative ontology offerings are collected
and used as the measure of their commitment. The ontology offering that collects the
biggest share of votes could therefore be considered as the best and most complete.
      </p>
      <p>In our case study the OntoElect approach is applied for refining the ICTERI Scope
Ontology in the iterative ontology engineering experiment. Our domain knowledge
stakeholders are the authors of ICTERI papers. Ontology offerings in the reported
work are the strucrtural contexts2 in the five thematic areas of the ICTERI scope
offered to the authors for choosing the appropriate key terms to annotate their papers.</p>
      <p>As this data had to be selected we decided to simulate the opinions of the electorate
by annotating the papers of ICTERI 2011 manually. For this we extracted the terms
which were specified as the list of ICTERI Key Terms if it was possible. In some
cases we had to add Missing Concepts (also called Missing Key Terms) for the
papers, if such terms did not exist in the list.</p>
      <p>
        So, we received three semantic annotation types:
 KeyWord – for the key words, which were selected by the authors
 KeyTerm – for the terms which were selected manually and were found in the list
of the ICTERI terms
2 A structural context, as suggested e.g. in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], is composed of a central concept with all his
domain and object properties and the concepts connected to the central concept by these
object properties.
 MissingConcept - for the terms, which did not exist in the list of the ICTERI terms,
but were covered during annotation
      </p>
      <p>One use of the particular term was considered to be one vote for the selected term.
The votes were normalized as frequencies of use. Such information allowed us to
measure the popularity of each semantic context, circumscribe the most frequently
demanded part of the ontology and make suggestion about the completeness of the
ontological offerings.</p>
      <p>For the papers of ICTERI 2012 we requested that the authors annotate their papers
not only using the freely chosen key words, but also using the terms found in ICTERI
scope ontology. As the result the corpus for the futher analysis was increased. The
data provided by the authors can be accepted as more authentic than that which was
obtained ourselves, as they are the real domain experts for the field they study. The
analysis of the received data is presented in Section 5.</p>
      <p>But even when we use the information presented by the authors, and the results of
our manual annotation we can't guarantee that this information is accurate enough for
applying it to the ontology refinement process. To obtain the experiment we needed to
have results received in different ways because we wanted to achieve the impartial
assessment of OntoElect approach. Before applying the results in ontology refining
process we decided to check them with freely available tool for text mining.</p>
      <p>For our experiment we chose one of the services provided by the National Centre
for Text Mining (NaCTeM). As reported in the official website of NaCTeM3 it is the
first publicly-funded text mining centre in the world. It provides text mining services
in response to the requirements of the UK academic community. NaCTeM is operated
by the University of Manchester.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Experimental Set-up and Tools</title>
      <p>To control the results of the experiment we have to understand which main questions
we are going to answer after its realization and how to measure these results.</p>
      <p>
        Our measurable objectives for the experiment have been formulated as follows [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]:
 Does the ontology fit to the requirements of the subject experts in the domain?
      </p>
      <p>The fitness of the ontological offering will be measured as a ratio of the average
frequency of use of the available Key Terms (positive votes) to the similar for the
missing Key Terms (negative votes). Special attention will be paid to the freely
chosen key words that are identical to the available Key Terms. Those will be considered
as extra positive votes for the semantic context of the Key Term.
 Is there a particular part in the ontology that is the most important for the
stakeholders?</p>
      <p>The importance of an ontology fragment comprising particular concepts will be
measured as frequency of use of these concepts (positive votes). Fragments of
different importance will also be presented as percentiles.
3 Official web site of National Centre for Text Mining http://www.nactem.ac.uk/
 Is there a part in the ontology that could be dropped as the stakeholders do not
really require it?</p>
      <p>Similarly to importance, these ontology fragments will be outlined using low
frequency of use percentiles.
 What would be a most valuable addition to the ontology that will substantially
improve stakeholders’ commitment to it?</p>
      <p>The papers have been annotated using missing Key Terms and freely chosen
keywords. Those missing Key Terms that are frequently used will form the core of this
effective extension. If some of the keywords are also used frequently by the authors
they may become good candidates for the inclusion in the effective extension as well.
Special attention will be paid to the freely chosen key words that are identical to the
missing Key Terms. Those will reinforce the votes on the addition to the ontology.</p>
      <p>
        The flow of activities has been organized in three consecutive phases as presented
in Fig. 1. The description of each phase in details is presented in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <sec id="sec-4-1">
        <title>Conference</title>
        <p>(MEaansyaCgehmaiern)t System ICTERI Semantic MediaWiki</p>
        <sec id="sec-4-1-1">
          <title>Submissions Data Corpus</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>Legend:</title>
          <p>- Parse
- Add Wiki mark-up
- Import to ICTERI
Wiki</p>
          <p>PHASE 1
- Activity is performed using the developed
tool
- Activity is performed manually
ICTERI Key Terms
t)se
o
(sv
n
o
itt
a
o
n
n
a
c
ti
n
a
m
e
S
s
m
r
e
T
y
e
K
g
issn
i
M
- Analyse votes
using SMW queries ICTERI papers
PHASE 3
- ontology
engineer
- Annotate using</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>ICTERI MediaWiki ICTERI Scope</title>
        <p>Ontology concepts
- Collect missing
concepts
- Collect free key
words</p>
        <p>PHASE 2</p>
        <sec id="sec-4-2-1">
          <title>ICTERI paper pages - 2011: ontology engineer</title>
          <p>At phase 1 we have extracted the semi-structured information about the papers
accepted for ICTERI 2012 and transformed these into the collection of paper articles in
the ICTERI Wiki. At Phase 2 we extracted the freely chosen KeyWords and the
KeyTerms from the ICTERI Scope ontology assigned by the authors and added these
to the semantic annotations of the papers. In several cases we detected considerable
meaning gaps between the extracted key words and Key Terms when annotated the
papers manually. Therefore, we opted to add the missing Key Terms to the
corresponding semantic annotations. As a result of this Phase the semantic relationships
between the pages of Category:Paper and the pages of Category:Concept
have been specified as semantic properties. These semantic properties allowed us to
receive all the measurements planned for the evaluation experiment. These
measurements have been done using different Semantic MediaWiki queries at Phase 3.</p>
          <p>
            Compared to the previous year experiment [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ], we automated the extraction of the
frequency of use statistics which made the process less error prone and faster. For that
the SMWAskAPI4 extension of the Semantic MediaWiki has been used. This
extension supports semantic queries of #ask and enables the use of the corresponding API
for executing Semantic MediaWiki ask queries.
          </p>
          <p>Each page of the ICTERI Wiki uses semantic tagging. An example of the Semantic
properties specified for pages in Category:Paper is given in Fig. 2.</p>
          <p>For our analysis we used the pages from Category:Paper and
Category:Workshop with the property hasPublicationYear equal to 2012, namely the
values of the semantic properties hasKeyWord, hasKeyTem, and MissingConcept.
The scripts for analyzing these values were coded in Python. Some steps were also
implemented using shell scripting. As outputs we have received:
 The list of KeyWords, KeyTerms, and MissingConcepts for each article, if they
were defined
 The overall number of the papers according to the values of the properties
hasPublicationYear, and the selected Category
 The number occurrences of each KeyWord, KeyTerm and MissingConcept.</p>
          <p>The analysis and discussion of the results is given in Section 5.</p>
          <p>
            To perform our second experiment we applied the Term Management System
named TerMine, which identifies key phrases in text. It uses C-value [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ], a
domainindependent method for automatic term recognition (ATR) which combines linguistic
and statistical analyses with the emphasis on the statistical part. The linguistic
analysis enumerates all candidate terms in a given text by applying part-of-speech tagging,
extracting word sequences based on adjectives/nouns, and stop-list. The statistical
4 See the description of the SMWAskAPI on
          </p>
          <p>SMWAskAPI
http://www.mediawiki.org/wiki/Extension:
analysis assigns a candidate term to a termhood by using the following four
characteristics:
 The occurrence frequency of the candidate term
 The frequency of the candidate term as part of other longer candidate terms
 The number of these longer candidate terms
 The length of the candidate term</p>
          <p>The data corpus for this term extraction and analysis was the merge of the pools of
ICTERI 20115 and ICTERI 20126 papers published in the respective proceedings, and
consisted of 63 papers. The papers from both proceedings volumes have been merged
in a single file and uploaded for processing by TerMine. The workflow for the second
experiment is pictured in Fig. 3.</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>ICTERI 2011 proceedings</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>ICTERI 2012 proceedings</title>
        </sec>
        <sec id="sec-4-2-4">
          <title>ICTERI proceedings</title>
        </sec>
        <sec id="sec-4-2-5">
          <title>File</title>
          <p>upload</p>
        </sec>
        <sec id="sec-4-2-6">
          <title>Term</title>
          <p>extraction
using
TerMine</p>
        </sec>
        <sec id="sec-4-2-7">
          <title>Statistical results for further analysis</title>
          <p>
            The data processed in the pipeline is illustrated by the example of a single
paper [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] in Fig. 4.
          </p>
          <p>
            The results of term mining were provided in several forms. All the terms defined in
the text were highlighted by colour markings (upper part of Fig. 4a). The information
about the overall number of the terms mined from the text was also given (433 terms
listed – in the bottom of Fig. 4a). The terms were also presented in the table view,
each preceded with the assigned rank number and followed by the statistical score
measure (lower part of Fig. 4a). The rank of a term means the position of each term in
the table sorted by the score; the rank values of the terms with the same score are
equal. The scores were computed automatically using the Term Recognition
technique [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ] which uses the information about the frequencies of term occurrence. This
approach is essentially a shallow bag of terms extraction technique – therefore the
output needs to be post-processed as described using our single paper data example.
For this example the number of extracted terms was 433 which is obviously too many.
To compare, the authors were advised to assign 3-5 KeyTerms to their papers which
best describe its meaning. Among those extracted terms that we needed to sort out
were also names, affiliations, cities, etc, which had no semantic relationship to the
5 http://ceur-ws.org/Vol-716/
6 http://ceur-ws.org/Vol-848/
meaning of the paper. Also, it has been assumed that the terms with a low number of
occurrences in text have a negligent semantic contribution and may also be filtered
out – so only the higher ranked part of the term list may be considered.
          </p>
        </sec>
        <sec id="sec-4-2-8">
          <title>Term list purifying</title>
          <p>Rank Term
1 ontology engineering
2 knowledge representation
3 subject expert
4 intended user
5 ontology element
6 psi suite
6 stakeholder commitment
6 semantic technology
9 active involvement
9 semantic web
9 ontological context
(b) 11 high-ranked terms listed</p>
          <p>While post-processing the list of mined terms we decided to leave only the terms,
which were used more than 5 times, and had score more than 5 points. Applying this
threshold returned 11 of 433 terms for the example outlined in Fig. 4., which
constitutes only 2.54 per cent of the overall number of the mined terms. The manual check
of the example however indicates that these 11 high ranked terms indeed contribute
most significantly to describing the semantics of the corresponding paper (Fig. 4b).</p>
          <p>The right part of Fig. 4 allows to compare the result of term extraction (Fig. 4b)
with the output of manual semantic annotation (Fig. 4c) for the selected example
paper. A mechanical comparison reveals substantial difference, which however is not
that big after manual mapping of the extracted terms to the concepts of the ICTERI
Scope ontology. In fact there is a subset of extracted terms that could be directly
mapped into the Key Terms of the ontology: subject expert; ontology engineering (as
a methodology). Another group is relevant to the assigned KeyWords: ontology,
stakeholder commitment, ontology engineering. Some are synonymic in the context
of this paper: stakeholder and intended user. Some represent the meaning which is too
fine-grained for a semantic annotation: ontology element. And, which is most
important, some are the new valid candidates for the inclusion into the ICTERI Scope
ontology: knowledge representation, semantic technology, semantic web.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results and Discussion</title>
      <p>In this section we present the results of the experiment. The set up of all its stages is
described in Section 4. The discussion of the experiment results is structured along
the measurable items.</p>
      <p>The frequency of use diagrams (Fig. 5 and 6) are built in regard to the total amount
of the papers and the number of occurrences of a particular term.</p>
      <p>
        Similar work was reported in our previous publication [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In it we described the
mechanism of ranging. The actual experiment is based on the previous results. But
they changed as the document corpus we are working with has increased and the data
to work with has changed. As we consider OntoElect approach in the case study of
iterative refinement of the ICTERI Scope Ontology such changes are greatly
important. The diagrams which show these changes are shown below:
 For the KeyWords which were selected by the authors manually (only those, which
were chosen by at least two authors, Fig. 5)
 For the KeyTerms, which were selected from the ICTERI ontology terms (Fig. 6)
      </p>
      <p>We did not provide a frequency of use diagram comparing the Missing Key Terms
because the difference in the results of 2011 and 2012 is tiny and could be neglected.
We provided the comparison analysis for KeyWords and Missing Key Terms lists
(Fig. 7). To compute the range of use of each term we divided each frequency of use
index by the frequency value of the most popular term, which range of use was taken
as 100 per cent. These terms are not the part of the ontology, but are the most possible
candidates.</p>
      <p>Academia</p>
      <p>Agent</p>
      <p>Approach
BioInspiredApproach
BusinessIntel igence</p>
      <p>Capability
Characteristic
Col aboration</p>
      <p>Competence
CompetenceFormationProcess</p>
      <p>Computation
Computing
Cooperation</p>
      <p>Data
DecisionSupport
Dev elopment</p>
      <p>Didactics
Env ironment
Ex perience
FormalMethod</p>
      <p>Grid</p>
      <p>GridInfrastructure
HighPerformanceComputing</p>
      <p>ICTComponent
ICTEnv ironment</p>
      <p>ICTTool
ImmaterialArtifact</p>
      <p>Industry
InformationCommunicationTechnology</p>
      <p>InformationTechnology</p>
      <p>Infrastructure
Integration</p>
      <p>Intel igence
Know ledgeEngineeringMethodology</p>
      <p>Know ledgeEngineeringProcess</p>
      <p>Know ledgeEv olution
Know ledgeManagementMethodology</p>
      <p>Know ledgeManagementProcess</p>
      <p>Know ledgeRepresentation
Know ledgeTechnology
Know ledgeTransfer</p>
      <p>LaborMarket</p>
      <p>LinkedData
MachineIntel igence</p>
      <p>Management
MathematicalModel
MathematicalModeling</p>
      <p>Method
Methodology
PSI-ULO:Object</p>
      <p>PSI-ULO:Process
PSI-ULO:ProcessPat ern</p>
      <p>Qualification</p>
      <p>Quality AssuranceProcess
Qualty AssuranceMethodology</p>
      <p>Reasoning
Requirement</p>
      <p>Research
SemanticWebServ ice
Social y InspiredApproach</p>
      <p>Softw areComponent
Softw areEngineeringProcess</p>
      <p>Softw areSy stem
SpecificationProcess
StandardizationProcess</p>
      <p>SubjectDomain</p>
      <p>SubjectEx pert
TeachingMethodology</p>
      <p>TeachingPat ern
TeachingProcess</p>
      <p>Technology</p>
      <p>Tool
VerificationProcess</p>
      <p>WebServ ice
m
r
e
T
y
e
K
d
e
s
U</p>
      <p>Metric</p>
      <p>Model
ModelBasedSoftw areDev elopmentMethodology</p>
      <p>MultiAgentSy stem</p>
      <p>Need
Nominativ eData
0
4
8
12
16
20
24</p>
      <p>28</p>
      <p>FrequencyOfUse, %
KeyT erm 2011</p>
      <p>KeyT erm 2012</p>
      <p>The results which we received at the previous step can be merged and we can
receive the new version of the potential ontology offering. But before doing this we will
look into the results of the text mining experiment.</p>
      <p>Using TerMine tool for automation of the knowledge mining process we received
some interesting results. The overall number of the found terms was 8487. All of the
selected terms were graduated and received the position in the rank table. Studying
the results it is obvious that the number of the terms selected by the data mining tool
is too big.</p>
      <p>It was decided to leave in the rank table only those terms which have the score of
10 and more. The popularity of the terms which were picked up is evident. The total
number of such terms is 157. It is easy to count that it makes up less than 2% of all
the terms proposed by the tool. After refining the list and deletion of the superfluous
information only 140 terms left.</p>
      <p>To compare the frequencies of use provided by the TerMine and our own
calculations we decided to use percentage method (similar to that used for building the
diagram in Fig. 7). We took the maximal value for each group of concepts as 100 per
cent and divided it by the frequency of use value of a particular term. As a result each
term got the value, called the range, which could be compared with the ranges of the
other terms. We analyzed the three groups of mined term matches to the: (i)
KeyWords; (ii) MissingConcepts; and (iii) KeyTerms. As the number of the Missing
Concepts is not too big we decided to combine them with the Key Words in the diagram
(Fig. 8).
Fig. 8. The range of use for terms detected by tool, Key Words and Missing Concepts. Range
values were normalized by the frequency of use of the highest scored extracted term (100)</p>
      <p>After the automated search of the identical terms in the lists of KeyTerms and
terms mined by tool we discovered that some of them were missed as the search did
not use the rules of common sense and the relations described in the ontology. For
example, according to the ICTERI Scope ontology the term Integration subsumes to
PSI-ULO:Process. Knowing this fact we understand that the term IntegrationProcess
is just the same as the term Integration. But this match is not obvious for the simple
search and will not be detected.</p>
      <p>Therefore, to find the matches in the lists of KeyTerms and terms mined by the tool
we decided to undertake a more careful analysis. We scanned the list of the KeyTerms
for matching the ToolTerms manually. Besides for this process we used the whole
pool of 8487 terms mined by the tool. The result is pictured in the diagram (Fig. 9).</p>
      <p>Agent
BusinessIntel igence</p>
      <p>Characteristic
Col aboration</p>
      <p>Competence
CompetenceApproach</p>
      <p>Computation</p>
      <p>Computing
DecisionSupport</p>
      <p>Development</p>
      <p>Didactics
Environment
FormalMethod</p>
      <p>Grid
HighPerformanceComputing</p>
      <p>IctComponent
ICTEnvironment</p>
      <p>IctTool
ImmaterialArtifact</p>
      <p>Industry
InformationCommunicationTechnology</p>
      <p>InformationTechnology</p>
      <p>Infrastructure
Integration</p>
      <p>Intel igence
KnowledgeEngineeringMethodology</p>
      <p>KnowledgeEngineeringProcess</p>
      <p>KnowledgeEvolution
KnowledgeManagementProcess</p>
      <p>KnowledgeRepresentation
Fig. 9. The range of use for the KeyTerms and the terms extracted by the TerMine tool</p>
    </sec>
    <sec id="sec-6">
      <title>Concluding Remarks and Future Work</title>
      <p>The paper reported on the experiment evaluating the improvement of OntoElect
approach to ontology engineering in the case study with the ICTERI Scope Ontology. In
particular, the approach has been used to evaluate the validity of the papers’
annotation by drawing knowledge stakeholders to their own papers’ annotation and by
studying their quality using voting for full texts.</p>
      <p>The experiment consisted of two stages. The first one was based on the
comparison analysis of the papers presented during international conferences ICTERI 2011
and ICTERI 2012. The second one was dedicated to performing automated term
extraction from the full texts and comparing the results of automated term extraction to
the outputs of manual semantic annotation.</p>
      <p>Achieved results stress the important parts of the ontology and those which are
less popular among the authors. The comparison analysis of the first experiment
shows how the situation changed during two years. The KeyWords and
MissingConcepts which have high frequency of use values, especially if they are named in both
lists, are the first candidates to become the new part of the ontology.</p>
      <p>The second experiment shows which ontological offerings agree with the terms
mined by tool and which numeric characteristics these matches have. The terms
extracted by the tool and their matches with the KeyWords and MissingConcepts, which
have range more than 50 per cent, are also good candidates to be added to the
ontology. Besides, the degree to which the extracted terms match the KeyWords and
KeyTerms indicate about the adequacy of paper annotation. Overall the overlap
between the meanings of the extracted terms and the KeyTerms measures the range of
so to say the similarity in the meanings of the papers within the corpus and the
ontological offerings aimed at covering these meanings. The quantitative results of our
experiments still need to be processed and analyzed more thoroughly before deciding
about the implementation of the changes to the ontology. Besides that several other
aspects still need to be researched in our future work.</p>
      <p>Firstly, the document corpus used in the case study, though growing, is still not
very big to allow robustly applying the majority of traditional knowledge extraction
techniques. At the moment it could only be stated that the information we have now is
enough to prove the concept – i.e. the validity of the approach based on the
assessment of and account for domain knowledge stakeholder opinions, implicitly reflecting
their needs. After applying the refinements suggested by the stakeholders, the
ontology still needs to be evaluated and validated using other methods.</p>
      <p>
        Secondly, in this paper we reported about only a partial and shallow way of
extracting knowledge from paper texts. A possible refinement to this preliminary
solution could be sought in using a hybrid iterative knowledge extraction workflow that
incrementally adds ontology elements to the “ontology learning layer cake (c.f.
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ])”.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Tatarintseva</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fensel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Is Your Ontology a Burden or a Gem? - Towards Xtreme Ontology Engineering</article-title>
          . In: Ermolayev,
          <string-name>
            <surname>V.</surname>
          </string-name>
          et al.
          <source>(eds.) Proc. 7-th Int. Conf. ICTERI</source>
          <year>2011</year>
          , Kherson, Ukraine, May 4-
          <issue>7</issue>
          ,
          <year>2011</year>
          , CEUR-WS.org, vol-
          <volume>716</volume>
          , ISSN 1613-
          <issue>0073</issue>
          ,
          <fpage>65</fpage>
          -
          <lpage>81</lpage>
          , online (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Tatarintseva</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borue</surname>
          </string-name>
          , Yu., and
          <string-name>
            <surname>Ermolayev</surname>
          </string-name>
          , V.:
          <article-title>OntoElect Approach for Iterative Ontology Refinement: a Case Study with ICTERI Scope Ontology</article-title>
          . In: Ermolayev,
          <string-name>
            <surname>V.</surname>
          </string-name>
          et al.
          <source>(eds.) Proc. 8-th Int. Conf. ICTERI</source>
          <year>2012</year>
          , Kherson, Ukraine, June 6-10,
          <year>2012</year>
          , CEURWS.org, vol-
          <volume>848</volume>
          , ISSN 1613-
          <issue>0073</issue>
          , 244,
          <string-name>
            <surname>online</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , Han,
          <string-name>
            <surname>J</surname>
          </string-name>
          .:
          <source>Survey on Social Tagging Techniques. SIGKDD Explorations</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ),
          <fpage>58</fpage>
          -
          <lpage>72</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Uren</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iria</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Handschuh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vargas-Vera</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciravegna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Semantic annotation for knowledge management: Requirements and a survey of the state of the art</article-title>
          .
          <source>Science. Services and Agents on the World Wide Web</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ),
          <fpage>14</fpage>
          -
          <lpage>28</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hunter</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gewrber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>HarvANA - Harvesting Community Tags to Enrich Collection Metadata</article-title>
          . In:
          <string-name>
            <surname>Paepcke</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borbiha</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naaman</surname>
            <given-names>M</given-names>
          </string-name>
          <source>(eds.) 8th ACM/IEEE-CS Joint Conference on Digital Libraries</source>
          ,
          <fpage>147</fpage>
          -
          <lpage>156</lpage>
          . ACM New York, New York (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Siorpaes</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hepp</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Games with a Purpose for the Semantic Web</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          <volume>23</volume>
          (
          <issue>3</issue>
          ),
          <fpage>50</fpage>
          --
          <lpage>60</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>W.</given-names>
            , and
            <surname>Bennamoun</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Ontology learning from text: A look back and into the future</article-title>
          .
          <source>ACM Comput. Surv.</source>
          ,
          <volume>44</volume>
          (
          <issue>4</issue>
          ),
          <source>Article</source>
          <volume>20</volume>
          , 36 pages.
          <source>DOI=10.1145/2333112</source>
          .2333115 (
          <year>September 2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Frantzi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Mima</surname>
          </string-name>
          , H.:
          <article-title>Automatic recognition of multi-word terms</article-title>
          .
          <source>Int. J. of Digital Libraries</source>
          <volume>3</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>117</fpage>
          -
          <lpage>132</lpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Peroni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>d'Aquin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Identifying Key Concepts in an Ontology, through the Integration of Cognitive Principles with Statistical and Topological Measures</article-title>
          .
          <source>In: Proc. 3rd Asian Semantic Web Conference (ASWC</source>
          <year>2008</year>
          ),
          <source>Dec 08-11</source>
          ,
          <year>2008</year>
          , Bangkok, Thailand (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Copylov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keberle</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jentzsch</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matzke</surname>
          </string-name>
          , W.-E..
          <article-title>Using Contexts in Ontology Structural Change Analysis.</article-title>
          . In: Ermolayev,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Gomez-Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-M.</given-names>
            ,
            <surname>Haase</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Warren</surname>
          </string-name>
          ,
          <string-name>
            <surname>P</surname>
          </string-name>
          , (eds.)
          <source>CIAO</source>
          <year>2010</year>
          ,
          <article-title>CEUR-WS</article-title>
          , vol.
          <volume>626</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Bennamoun</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Ontology Learning from Text: a Look Back and into the Future</article-title>
          .
          <source>ACM Comput. Surv.</source>
          ,
          <volume>44</volume>
          (
          <issue>4</issue>
          ),
          <source>Article</source>
          <volume>20</volume>
          , 36 p., http://doi.acm.
          <source>org/10</source>
          .1145/2333112.2333115 (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>