<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>E. Rilo. An Empirical Study of Auto- [Rilo,
traction in Three Domains. AI Journal</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <pub-date>
        <year>1996</year>
      </pub-date>
      <volume>85</volume>
      <issue>101</issue>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>1URL:http://www.cyc.com/</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>3. Story annotation. Either a journalist or a knowl</title>
      <p>Knote (the Planet knowledge editor),
edge engineer manually annotates the story using
7. Story retrieval and query answering. The Lois
interface supports integrated access to the story
archive</p>
    </sec>
    <sec id="sec-2">
      <title>The paper is organised as follows: Section 2 briey</title>
      <p>describes our suite of tools in order to give some
backbelieve it is important to mix the syntactic and
semanin this case \murder", in any tense. Trigger words can
sentence is also matched using the extra words \by" and
Conversely, some grammar constructions (such as dates)
ambiguities, for example, ontologies can provide us with
\with."
visiting a place then this someone should be a person."
of objects by syntactic features or by means of the KMi
tactically identify some fragments of the sentences. We
scribe the use of template-driven IE to populate an
ontic. The semantic checking is often necessary to resolve
tion 5 gives conclusions and directions for future work.
sentences that contain keywords (target-string) encoded
information is the subject or the direct object of a verb
a dieren t set of keywords. By using this keywords we
select only relevant sentences from the whole text. then
in our templates. For our specic event type we have
tology. The template matching itself is supported
seaxioms of common sense knowledge such \if someone is
tent.</p>
      <p>Planet ontology (dened in section 2).
lem of populating the ontology. Section 4 describes the
use of ontology to cope with the ambiguity in the
idendescribed in 1996 For example, if the targeted [Rilo, ].
an ontology engine to supply the necessary semantic
conthen the best trigger word should be the main verb. The
based on some of the event types we use. Finally,
Secbe reliably identied using linguistic rules lik e the ones
we parse this sentences, and nally recognize the t ypes
ground. In Section 3 we present a solution to the
probmantically by referring to the ontology, but also
conclassication process embedded in our system and the
tains some lightweight NLP techniques in order to
syntication of objects in the story, along with examples
can be recognized robustly. Overall, our primary
contriFor our IE we use a template-driven approach
similar to, and based upon, that of Rilo. We identify the
bution is to integrate a template-driven IE engine with
In particular, the main aim of this paper is to
dewhich are relevant to our institute. The Ontology
Licommunication between members of KMi [Domingue
journalists. In the case of KMi Planet it contains stories
and Motta, 2000].
the ontology library. The Story library contains the text
ing activities:
brary contains several existing ontologies, in particular
of the stories that have been provided to Planet by the
nario supported by KMi planet, and supports the
followTwo primary components are the story library and
the KMi ontology.</p>
      <p>KMi Planet is a Web-based news server that facilitates
augmented the basic publish/nd sce- PlanetOnto2</p>
    </sec>
    <sec id="sec-3">
      <title>2. Story reading. A Planet reader browses through</title>
      <p>the latest stories using a standard Web browser,
pre-process the email-text: tokenise the text and
join expressions such as dates into a single token
a new domain the entire knowledge process of dening
Then the system decides to classify the story according
For the templates we use a prolog-like notation: the
ontology. For example, \visiting-a-place-or-people",
\academic-conference", \academic-workshop",
\eventa story more than one type of event can be recognised.
mining which kind of templates are necessary in our set
involving-project", etc. Each event has several templates
ple templates. However, problems might occur when in
of templates. Currently, in order to apply our system to
with the following criteria: how many feature for each
sentence is converted to a list of words (or tokens) and
keywords, some variables which are allowed to match
then matched against a template list containing some
We have devised several type of events in KMi planet
used to recognize it. So far we have a library of 100
simThe nature of the domain is very important in
detertemplates has to be repeated.
type were encontered in the story.
one or more words, and some underscores \ " that can
dened in our KMi planet on tology
1. classify the story according an event type
prend the objects in the story using further partial
parsing
identication
3. produce OCML code
4 Examples of classication and object
Figure 1: Email Story
on syntactic features, e.g. we assume titles such as
ample, project names will be recognized this way.
domain this means classify them as person, place,
date, etc. This identication is done based partially
mantic knowledge obtained by querying the server
\Dr." can be used to recognize some people.
How\Web-onto" for the KMi planet ontology. For
exidentify each of the objects in the story. In our
ever, some other classications are based upon
sejournalists use slang, break conventions for
capitalizaknowledge engineer for conrmation. For each type of
snapshot the extracted information is presented to the
formation shown in screen snapshot of Figure 2 In this
side of the form there is a menu that presents to the
event a dieren t form is presented from the information
writing in such an archive can vary wildly. All these
tion, incorrectly use punctuation, etc. Also, styles of
ample, if a word is capitalised and not starting a sentence
nalist from KMi sent the story (presented in Figure 1)
company. However, as discussed earlier we expect such
user with possibles instantiations for each eld. The
suge-mail message) and see if it matches any of our
domainorganization-being-visited. Therefore, in the right hand
Our domain (KMi planet) consists of email message,
specic templates. If no extraction template applies to
Let us take the following example: Suppose that a
jourmeans that irrelevant text can be processed very quickly.
Our IE tool will take the story and produce the
in\Ltd" then it can be tagged as a proper noun of type
The main process in the classication is to tak e each
make it diÆcult to apply heuristics to identify proper
use templates.
sentence in the text (in our case a story written in a
a sentence, then no information will be extracted; this
to the e-mail server
nouns. NLP approaches typically use heuristics, for
exstored in the ontology. Figure 2 also has some elds
and hence is generally unstructured for example: the
then it is a proper name. If a string contains \&amp; Co" or
methods to be too fragile for our domain, and so instead
that cannot be found in the story such as
people-orgested instantiations are obtained from the ontology.
system might encounter objects of type: visitor, place
instance, for the event \visiting-a-place-or-people" the
object’s types.
predened objects that should be found in the story . For
of the 40 types of events according with the type of the
Classication is performed in the follo wing steps:
and date. Currently, our system handles 22 dieren t
objects that are found. For each event type we have a
We classify the story or document as belonging to any
2. [ ,X, , \visited", Y, \from", Z , ] matches the
automatically recognized by syntactic features of
month, day, and year.
&amp; place (Y), but also require \date(Z)"; dates are
sentence word list. Again, we demand visitor(X)
for matching to the visited templates above::
would also be a match against the template, and
extract the same information. In contrast, the misspelled
tirely and extract no information.
Let us take another example of a candidate sentence
initial \Tooo" might cause a complete parser to fail
en</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>EKAW'00 Workshop on Ontologies and Text</source>
          , Juan-
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Les-Pins</surname>
          </string-name>
          , France, oct
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          et al.,
          <string-name>
            <surname>J-U. Kietz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Maedche</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Volz</surname>
          </string-name>
          . [Kietz 2000]
          <article-title>A method for semi-automatic ontology acquisition</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>