<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Knowledge Acquisition for Knowledge Management: Position Paper</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christopher Brewster</string-name>
          <email>C.Brewster@dcs.shef.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Ciravegna</string-name>
          <email>F.Ciravegna@dcs.shef.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yorick Wilks</string-name>
          <email>Y.Wilks@dcs.shef.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, The University of Sheffield</institution>
          ,
          <addr-line>Regent Court, 211 Portobello Street, S1 4DP, Sheffield</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>2. Ontology Learning</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Knowledge is considered to be “the information needed to
make business decisions”1, and so knowledge management
is the “essential ingredient of success” for 95 per cent of
CEOs [Manchester 1999]. A company’s value depends
increasingly on “intangible assets”2 which exist in the minds
of employees, in databases, in files and in a myriad
documents. Knowledge management technologies capture
this intangible element in an organisation; and make it
universally available. The most widely used method of
mapping the knowledge of a domain is to use an ontology
describing such a domain. Ontologies can act as an index to
the memory of an organisation and facilitate semantic
searches and the retrieval of knowledge from the corporate
memory as it is embodied in documents and other archives.
There are many real-world examples where the utility of
ontologies as maps or models of specific domains has been
repeatedly proven.</p>
      <p>Human</p>
      <p>Introspection
KNOWLEDGE
ELICITATION</p>
      <p>Human
Introspection
KNOWLEDGE
BASE
DEFINITION</p>
      <p>Human</p>
      <p>Introspection
KB
POPULATION
Figure 1 A basic knowledge base development cycle (current reality)</p>
      <p>For knowledge to be managed it must first of all be
captured or acquired in some useful form, e.g. stored in an
ontology. Knowledge acquisition (KA) is a complex process
which traditionally is extremely expensive. Yahoo currently
employs over 100 people to keep its category hierarchy up
to date [Dom 1999]. An appropriate model of knowledge
acquisition under the current paradigm would note that it is
1 Philip Crawford, European Vice-President of Oracle Corp. cited
by [Philip 1999]
2 A term coined by the industry consultant Karl-Erik Sveiby
an entirely manual process. The typical process of KA is
outlined in Fig. 1: initially knowledge is elicited either from
a set of documents or from an expert of the domain. Then an
ontology is defined, finally a set of instances are associated
to concepts in the knowledge base. It is widely agreed that
the key factors which impede the wider use of ontologies
both in research and commercial applications are time, cost
and subjectivity. Time and money committed to ontology
development is substantial as the years of development at
Cyc have shown and the manpower requirements of Yahoo
prove. Subjectivity is inevitable, given any method which
depends on individual introspection or elicitation for data
collection.</p>
      <p>Although efforts exist to automate KA most of these are
isolated to specific aspects of the construction cycle, little
effort has been spent on creating a comprehensive set of
tools for KA. In this paper, we propose a set of techniques
to largely automate the process of KA, by using
technologies based on Information Extraction (IE),
Information Retrieval and Natural Language Processing. We
aim to reduce all the impeding factors mention above and
thereby contribute to the wider utility of the knowledge
management tools. In particular we intend to reduce the
introspection of knowledge engineers or the extended
elicitations of knowledge from experts by extensive textual
analysis using a variety of methods and tools, as texts are
largely available and in them – we believe – lies most of an
organization’s memory.
Ontology construction aims to capture knowledge in a
usable format. The nature and granularity of the ontology
depends on its eventual use; typically IS-A hierarchies are
central but other emphases are often encountered.</p>
      <p>The process of ontology construction may be divided into
three stages the first two of which contribute to the learning
of the ontology structure and the third is used to populate
the knowledge base with instances. These stages are
illustrated in the rest of this section.
refinement [Järvelin and Kekäläinen 2000]; iii) passed on to
the next development stage below.</p>
      <sec id="sec-1-1">
        <title>2.1 Taxonomy construction</title>
        <p>
          We propose to introduce automation in the stage of
taxonomy construction mainly in order to eliminate or
reduce the need for extensive elicitation of data. In the
literature approaches to construction of taxonomies of
concepts have been proposed [Brown et al. 1992, McMahon
and Smith 1996, Sanderson and Croft 1999]. Such
approaches either use a large collection of documents as
their sole data source, or they can attempt to use existing
concepts to extend the taxonomy [Agirre et al.2000, Scott
1998]. We intend to develop a semi-automatic method that,
starting from a seed ontology sketched by the user, produces
the final ontology via a cycle of refinements by eliciting
knowledge from a collection of texts. In this approach the
role of the user should only be that of proposing an initial
ontology and validate/change the different versions
proposed by the system. We believe an ontology
construction method should a) permit multiple placement of
terms in the structure, b) allow rapid recalculation of the
structure, c) provide monothetic labels for nodes, d) allow
the input of seed ontologies for further expansion.
We intend to integrate a methodology for automatic
hierarchy definition
          <xref ref-type="bibr" rid="ref14">(such [Sanderson and Croft 1999])</xref>
          with
a method for the identification of terms related to a concept
in a hierarchy
          <xref ref-type="bibr" rid="ref15">(such as [Scott 1998])</xref>
          .
        </p>
        <p>The advantage of this integration is that as knowledge is
continually changing, we can reconstruct an appropriate
domain specific ontology very rapidly. This does not
preclude incorporating an existing ontology and using the
tools to extend and update it on the basis of appropriate
texts. Finally an ontology defined in this way has the
particular advantage that it overcomes the well-known
‘Tennis problem’ associated with many predefined
ontologies such as WordNet, i.e where terms closely related
in a given domain are structurally very distant such as ball
and court.</p>
        <p>In addition we intend to employ classic Information
Extraction techniques such as Sheffield’s named entity
recognition system [Humphreys 1998] in order to
preprocess the text, as the identification of complex terms
such as proper names, dates, numbers, etc, allows to reduce
data sparseness in learning [Ciravegna 2000].</p>
        <p>We plan to introduce many cycles of ontology learning
and validation. At each stage the defined ontology can be: i)
validated/corrected by a user/expert; ii) used to retrieve a
larger set of appropriate documents to be used for further</p>
      </sec>
      <sec id="sec-1-2">
        <title>2.2 Learning Other Relations</title>
        <p>This stage proceeds to build on the skeletal ontology in
order to specify, as much as possible without human
intervention, relations among concepts in the ontology,
other than ISAs. In order to flesh out the concept relations,
we need to identify relations such as synonymy, meronymy,
antonymy and other relations. We plan to integrate a variety
of methods from the literature, e.g. by using recurrences in
verb subcategorisation as a symptom of general relations
[Basili et al. 1998], by using Morin’s user-guided approach
to identify the correct lexico/syntactic environment [Morin
1999], and by using methods such as [Hays 1997] to locate
specific cases of synonymy.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Populating the Knowledge Base</title>
      <p>
        Once the ontology has been learnt, there is the problem of
retrieving instances in order to populate the resulting
knowledge base. This is a key issue, in order to use the
ontology as index to the organisation’s memory, for
example by allowing semantic searches and the retrieval of
knowledge from the corporate memory. In many
environments KB population is performed manually by a
user via instance identification in a text corpus. We plan to
automate this process as much as possible by using a
combination of text classification (TC)
        <xref ref-type="bibr" rid="ref4">(e.g. [Ciravegna et al
1999])</xref>
        and Adaptive Information Extraction
        <xref ref-type="bibr" rid="ref5">([Ciravegna
2001])</xref>
        . Text classification is useful in order to identify the
scenario to apply to a specific set of texts, while IE will
identify (i.e. index) the instances in the texts. Both TC and
IE should be adaptive, as it is generally not possible to ask a
user to develop rules himself for each scenario/application.
Again, automating the process will both reduce the cost of
instance identification and the subjectivity involved in the
human identification.
      </p>
    </sec>
    <sec id="sec-3">
      <title>4. Conclusion and future work</title>
      <p>Knowledge is only of value when it can be used effectively
and efficiently. The management of knowledge is a key
element in extracting its value. In this position paper we
have outlined how we are addressing the issue of
automating the Knowledge Acquisition process in order to
reduce both required time and cost of KA, and subjectivity
in the resulting ontology. Overall we believe, this will make
knowledge management not only more acceptable in a
commercial environment but also contribute to the overall
productivity of the economy.</p>
      <p>The work outlined above is being undertaken by
the University of Sheffield in the context of AKT
(Advanced Knowledge Technologies,
http://www.aktors.org), a multi-million pound, six-year
project involving the University of Southampton, the Open
University, the University of Edinburgh, the University of
Aberdeen, and the University of Sheffield. AKT will extend
knowledge management technologies to exploit the
potential of the semantic web, covering the use of
knowledge over its entire lifecycle, from acquisition to
maintenance and deletion. It began in October 2000 and will
comprehensively addresses six main challenges, which are
fundamental bottlenecks to knowledge management:
• acquisition • reuse
• modelling • publication
• retrieval/extraction • maintenance
The work at Sheffield will on provide a library of Natural
Language Processing based tools for different types of
knowledge management tasks.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgement</title>
      <p>This work is supported under the Advanced Knowledge
Technologies (AKT) Interdisciplinary Research
Collaboration (IRC), which is sponsored by the UK
Engineering and Physical Sciences Research Council under
grant number GR/N15764/01. The views and conclusions
contained herein are those of the authors and should not be
interpreted as necessarily representing official policies or
endorsements, either express or implied, of the EPSRC or
any other member of the AKT IRC.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Agirre et al. 2000]
          <string-name>
            <given-names>Eneko</given-names>
            <surname>Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Olatz</given-names>
            <surname>Ansa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Hovy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Martínez</surname>
          </string-name>
          ,
          <article-title>Enriching very large ontologies using the WWW</article-title>
          ,
          <source>in Proceedings of the ECAI 2000 workshop “Ontology Learning” 2000</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Basili et al. 1998] Basili,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Catizone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stevenson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Velardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vindigni</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Wilks</surname>
          </string-name>
          <article-title>An Empirical Approach to Lexical Tuning</article-title>
          .
          <source>In Proc. of the Adapting Lexical and Corpus</source>
          Resources to Sublanguages and Applications Workshop,
          <article-title>held jointly with 1st LREC Granada</article-title>
          , Spain,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Brown et al. 1992] Brown, Peter F.,
          <string-name>
            <surname>Vincent</surname>
            <given-names>J. Della</given-names>
          </string-name>
          <string-name>
            <surname>Pietra</surname>
          </string-name>
          , Petere V. DeSouza, Jenefer C. Lai, Robert L.
          <string-name>
            <surname>Mercer</surname>
          </string-name>
          ,
          <year>1992</year>
          <article-title>Class-based n-gram models of natural language</article-title>
          ,
          <source>Computational Linguistics</source>
          ,
          <volume>18</volume>
          ,
          <fpage>467</fpage>
          -
          <lpage>479</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Ciravegna et al. 1999]
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Ciravegna</surname>
          </string-name>
          , Alberto Lavelli, Nadia Mana , Luca Gilardoni, Silvia Mazza, Johannes Matiasek,
          <string-name>
            <given-names>William</given-names>
            <surname>Black</surname>
          </string-name>
          , Fabio Rinaldi,
          <source>David Mowatt "Classifying Texts Integrating Pattern Matching and Information Extraction" Proceedings of the Sixteenth International Joint Conference on Artificial Intelligence (IJCAI99)</source>
          , Stockholm,
          <year>August</year>
          ,
          <year>1999</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Ciravegna 2001]
          <article-title>Fabio Ciravegna "Adaptive Information Extraction from Text by Rule Induction and Generalisation"</article-title>
          <source>in Proceedings of the Seventeenth International Joint Conference on Artificial Intelligence (IJCAI</source>
          <year>2001</year>
          ), Seattle,
          <year>August 2001</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Dom 1999]
          <article-title>Byron Dom, Automatically finding the best pages on the World Wide Web (CLEVER), In Search Engines and Beyond: Developing efficient knowledge management systems</article-title>
          ,
          <source>April 19-20</source>
          <year>1999</year>
          , Boston, Mass
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Hays 1997]
          <string-name>
            <given-names>Paul R.</given-names>
            <surname>Hays</surname>
          </string-name>
          , Collocational Similarity:
          <article-title>Emergent Patterns in Lexical Environments</article-title>
          , Dissertation submitted to the School of English, University of Birmingham
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Hearst 1992]
          <article-title>Marti Hearst, Automatic Acquisition of Hyponyms from Large Text Corpora</article-title>
          , COLING
          <volume>92</volume>
          ,
          <string-name>
            <surname>Nantes</surname>
          </string-name>
          ,
          <year>1992</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Humphreys</source>
          <year>1999</year>
          ]
          <string-name>
            <given-names>K.</given-names>
            <surname>Humphreys</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gaizauskas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Azzam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Huyck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mitchell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wilks</surname>
          </string-name>
          .
          <article-title>Description of the University of Sheffield LaSIEII System as used for MUC-7</article-title>
          .
          <source>In Proceedings of the Seventh Message Understanding Conference (MUC-7)</source>
          . Morgan Kaufmann. 1999
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>[Järvelin and Kekäläinen</source>
          <year>2000</year>
          ]
          <article-title>Kalervo Järvelin and Jaana Kekäläinen IR evaluation methods for retrieving highly relevant documents</article-title>
          .
          <source>In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>48</lpage>
          , Athens, Greece,
          <volume>24</volume>
          .-
          <fpage>28</fpage>
          .7.
          <year>2000</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Manchester 1999] Philip, Manchester, “Survey - Knowledge
          <source>Management” Financial Times</source>
          ,
          <volume>28</volume>
          April,
          <year>1999</year>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[McMahon and Smith</source>
          <year>1996</year>
          ]
          <article-title>McMahon</article-title>
          ,
          <string-name>
            <given-names>John G</given-names>
            ,
            <surname>Francis</surname>
          </string-name>
          <string-name>
            <given-names>J.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <year>1996</year>
          <article-title>Improving Statistical Language Models Performance with Automatically Generated Word Hierarchies</article-title>
          ,
          <source>Computational Linguistics</source>
          ,
          <volume>22</volume>
          (
          <issue>2</issue>
          ),
          <fpage>217</fpage>
          -
          <lpage>247</lpage>
          , ACL/MIT
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Morin 1999]
          <article-title>Emmanuel Morin, Using Lexico-Syntactic patterns to Extract Semantic Relations between Terms from Technical Corpus</article-title>
          , TKE
          <volume>99</volume>
          ,
          <fpage>268</fpage>
          -
          <lpage>278</lpage>
          , Innsbruck, Austria,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <source>[Sanderson and Croft</source>
          <year>1999</year>
          ]
          <article-title>Mark Sanderson and Bruce Croft, Deriving concept hierarchies from text</article-title>
          ,
          <source>in Proceedings of the 22nd ACM SIGIR Conference</source>
          ,
          <volume>206</volume>
          -
          <fpage>213</fpage>
          ,
          <year>1999</year>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Scott 1998]
          <article-title>Mike Scott, Focusing on the Text and Its Key Words</article-title>
          ,
          <source>TALC 98 Proceedings, Oxford, Humanities Computing Unit</source>
          , Oxford University,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>