<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Artequakt: Generating Tailored Biographies with Automatically Annotated Fragments from the Web</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sanghee Kim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harith Alani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wendy Hall</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul H. Lewis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David E. Millard Nigel R. Shadbolt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark J. Weal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Intelligence, Agents, Multimedia Group, University of Southampton</institution>
          ,
          <addr-line>SO17 1BJ</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Artequakt project is working towards automatically generating narrative biographies of artists from knowledge that has been extracted from the Web and maintained in a knowledge base. An overview of the system architecture is presented here and the three key components of that architecture are explained in detail, namely knowledge extraction, information management and biography construction. Conclusions are drawn from the initial experiences of the project and future plans are described.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The growth of the World Wide Web (Web) and the corpus of
documents that it covers increased the demand for content to be
annotated. Such annotation facilitates systematic search and discovery of
knowledge and intelligent information processing. Annotating
existing Web documents forms one of the basic barriers towards realising
the Semantic Web ([
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]).
      </p>
      <p>Annotations can be roughly classified into two types. The first is
concerned with identifying textual entities in documents that match
information already existing in a knowledge base, e.g. the word
‘Rembrandt’ in the document is matched to a painter’s name in the
knowledge base. Such annotations are normally restricted to the type
and amount of information held in the knowledge base. The other
type of annotation is involved in locating new factual information
in documents based on a given domain classification structure, e.g.
‘Rembrandt’ in the document is the ‘name’ of a ‘Painter’, where
Painter is a class in the ontology with the relation name. This new
fact can be asserted in the knowledge base. This second type is the
main approach taken to annotation in the Artequakt project.</p>
      <p>
        Previous work on annotation has demonstrated the value of
coupling Natural Language Processing (NLP) with ontologies ([
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]). The ontology can guide the annotation task by restricting it
to a specific domain and, unlike “rigid templates”, can provide it
with knowledge inference and conceptual browsing facilities [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
An ontology-based approach for annotation needs to deal with the
issues of duplicate information across documents, managing
ontology change, and redundant annotations [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>Annotation can exist in different forms and be used in a variety of
ways. One interesting possibility is to use it to restructure the
original source material in new ways, producing a dynamic presentation
tailored to the users needs.</p>
      <p>
        Previous work on the production of dynamic presentations has
highlighted the difficulties of maintaining a rhetorical structure
across a dynamically assembled sequence [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], as a consequence
there has been a focus on dynamic presentation decisions as opposed
to narrative ones [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Where dynamic narrative is present it has been
based around robust story-schema such as the format of a news
program (a sequence of atomic bulletins) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>It is our belief that by building a story-schema layer on top of
an ontology we can create dynamic stories within a certain domain.
By populating the ontolgy through automatic annotation software we
could allow those stories to be constructed from the vast wealth of
information that exists on the World Wide Web.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>The Artequakt Project</title>
      <p>The Artequakt project aims to implement such a system around the
domain of artists and their paintings, automatically producing
tailored biographies of artists from fragments of information extracted
from the Web. This is not an attempt to out-perform hand-crafted
biographies, but rather to gather information from a wide variety
of sources and target it specifically at the interests of a particular
reader. The first stage of this project consists of developing an
ontology for the domain of artists and paintings. A selection of
information extraction tools and techniques are being developed and applied
that attempt to automatically generate annotated content from online
documents based on the project’s ontology and WordNet lexicons.
The annotations are stored in a knowledge base and will be analysed
for duplications. In the second stage, narrative construction tools are
being developed to query the knowledge base through an
ontologyserver to search and retrieve relevant facts or textual paragraphs and
generate a specific biography. The automatic generation of tailored
biographies is concerned with two areas of focus. Firstly, providing
biographies for artists where there is sparse information available,
distributed across the web. This may mean constructing text from
basic factual information gleaned, or combining text from a number
of sources with differing interests in the artist. Secondly, the project
aims to provide biographies that are tailored to the particular interests
and requirements of a given reader. These might range from rough
stereotyping such as “A biography suitable for a child” to specific
reader interests such as “I’m interested in the artist’s use of colour in
their oil paintings”.</p>
      <p>The expertise and experience of three separate projects are drawn
together under the umbrella of the Artequakt project. These are:
The Artiste project - A European project working on a distributed
database of art images in collaboration with partners that include
the Louvre, the Uffizzi Gallery, the National Gallery and the
Victoria and Albert Museum.</p>
      <p>The Equator IRC - An EPSRC funded Interdisciplinary Research
Centre that, amongst many other activities, is investigating the use
of narrative techniques in information structuring and
presentation.</p>
      <p>The AKT IRC - An EPSRC funded Interdisciplinary Research
Centre looking at all aspects of the knowledge lifecycle.</p>
      <p>Although focussing on artists and their paintings, the techniques
being developed could be applied to other domains.</p>
      <p>This paper will examine the overall proposed Artequakt
architecture, looking at the three main component parts, namely, knowledge
extraction, knowledge representation and storage and narrative
generation.
2</p>
    </sec>
    <sec id="sec-3">
      <title>ARCHITECTURE OVERVIEW</title>
      <p>Figure 1 illustrates the systems architecture used for the initial
Artequakt demonstrator. Three key areas can be identified.</p>
      <p>The first concerns the knowledge extraction tools. These are to be
used to extract factual information items together with sentences and
paragraphs from web documents that might be manually selected or
obtained automatically using appropriate search engine technology.
The fragments of information are passed to the ontology server along
with metadata derived from the vocabulary of the ontology.</p>
      <p>The second key area is the information management and storage.
The information is being stored by the ontology server and
consolidated into a knowledge base, focused on artists and paintings.</p>
      <p>The final key area, is the narrative generation. The Artequakt
servlet will take requests from a reader via a simple web interface.
The reader request will usually include an artist for whom to generate
a biography in a particular style (chronology, through the paintings
etc.) and also any user information; for example, the narrative might
be generated specifically for a child or an art historian. The server
then uses story templates to render a narrative from the information
stored in the knowledge base. The rest of this paper will examine
these three areas in more detail.</p>
      <p>Input Web Pages</p>
      <p>The Biography is
rendered as a web page
Reader selects an artist
and a biography style
Knowledge
Extraction</p>
      <p>Tools</p>
      <p>Artequakt</p>
      <p>Servlet
Ontology</p>
      <p>Server</p>
    </sec>
    <sec id="sec-4">
      <title>KNOWLEDGE EXTRACTION</title>
      <p>
        The aim of the knowledge extraction section is to extract and identify
factual information from the Web-based documents and to structure
it appropriately for entry into the knowledge base. Much of the
information from the Web is in the form of natural language documents.
One of the promising approaches to providing easy access to such
documents is centred on information extraction that reduces them
into tabular structures from which the fragments of documents can be
retrieved as answers to queries. However, the effort and time needed
for annotating a large number of texts and the prerequisite of
acquiring background knowledge that stipulates which types of information
are extractable, are major challenges toward exploiting such
extraction techniques for practical purposes([
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]). Work such as ([
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ])
investigated the application of machine learning techniques in order
to automatically identify patterns from annotated example texts.
      </p>
      <p>In particular, whilst many attempts have been made to extract
information from the Web by using manually annotated texts, no robust
and reliable methodologies are yet available. Documents from the
Web use limitless vocabularies, structures and/or composition styles
for defining approximately the same content, implying that it is of
little use to make efforts to locate recurrent syntactic patterns. For
example, although content similarity between two biographic
documents might be expected, expressions used for both sources may
vary dramatically.</p>
      <p>
        These observations have led us initially to use a natural
languagebased extraction approach for a comparatively deeper content
understanding from which various clues concerning semantic and
syntactic features can be obtained. The use of an ontology coupled with a
general-purpose lexical database (WordNet [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]) as a guidance tool
for creating interesting relations is another dimension of our initial
approach aiming at minimising reliance on domain-specific
extraction rules. Figure 2 shows extraction results based on the
example of ‘Rembrandt’s father was a miller who died in 1630’. Two
biographic pieces of information about ‘Rembrandt’s father’ (i.e.
‘job title (miller)’ and ‘date of death (1630)’), were captured as well
as the fact that ‘Rembrandt’ is a person and he is the son of a dead
miller.
3.1
      </p>
    </sec>
    <sec id="sec-5">
      <title>Natural Language Information Extraction</title>
      <p>
        The capability of recognising a named entity without the annotation
effort of humans or without the need to create extraction rules is one
of the objectives of our approach. The idea is to make use of
generalpurpose lexical databases and to exploit the knowledge from
syntactical and semantic analysis to clarify the types and structures of given
information. Although the proposed approach may not be as
sophisticated as manually annotated definitions, its contribution lies in its
extensibility and practical nature (acceptable performance). We use a
paragraph as a unit of semantic analysis instead of a sentence, since
much of the critical information used for interpreting text is scattered
in different sentences (as observed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]). Downloaded documents
from the Web are first divided into paragraphs, which are then
broken down into a group of sentences. The paragraphs are analysed as
follows:
1. Syntactical analysis: A sentence is decomposed into a set of
grammatically related phrases (e.g. a verb-phrase, or a noun-phrase).
We have used the Apple Pie Parser, which is a bottom-up
probabilistic chart parser and is freely available [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
2. Semantic analysis:
Identification of main components: each compound sentence is
decomposed into simplified structures, each of which contains
one clause, i.e. a simple sentence. Each clause is clustered as
one of three parts: subject, verb, and object. Temporal
properties are inferred from a verb tense (e.g. ‘past’, ‘present’), and
associated with the sentence. A writing style (e.g. ‘first-person’,
‘third-person’) can be derived from the personal pronoun if it
exists in the sentence’s subject.
      </p>
      <p>
        Recognition of named entity: two resources are used for
determining whether or not a given word denotes a person’s name.
The first is syntactical tags, which are obtained as the result
of the syntactical analysis carried out by the Apple Pie Parser.
The second is gazetteers of people names, which are available
as part of the GATE (General Architecture for Text
Engineering, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) package. GATE provides text files which contain
person names associated with gender attributes. A name which is
not defined in GATE’s text files will still be extractable if it
is tagged as a proper noun. Heuristics and grammar rules are
applied in order to extract only proper personal nouns.
      </p>
      <p>Resolution of pronoun references (anaphoric references): a
personal pronoun refers to a specific person, and acts as a subject
(‘he’ or ‘she’), an object (‘him’ or ‘her’), or a marker of
possession defining who owns a particular thing (‘his’ or ‘hers’).
Currently we are using a simple resolution function that runs
at reasonably fast speed obtaining the best-guessed referent.
Three attributes (gender, number, and structural information)
are considered in determining the right referent.</p>
      <p>Adding a missing subject: a clause can inherit a subject from
a main clause, since it is syntactically dependent on the main
clause.</p>
      <p>In Figure 2, the given example ‘Rembrandt’s father was a miller
who died in 1630’ is divided into two clauses. The same subject (i.e.
‘Rembrandt’s father’) is assigned to both clauses since the second
clause is dependent on the first one. At this stage, ‘Rembrandt’ was
successfully recognised as a person’s name. Gazetteers provided by
GATE do not contain the name ‘Rembrandt’, whereas syntactic tags
for this sentence mark it as a proper noun.
3.2</p>
    </sec>
    <sec id="sec-6">
      <title>Relation Extraction</title>
      <p>To create a binary relationship between two extracted individual
facts, knowledge about the pre-defined semantic relations will be
required. Consulting the ontology, which specifies various
relationships among classes, will act as a basis for decisions concerning
which relations to use. A query is submitted to the ontology server to
obtain such knowledge.</p>
      <p>In order to reduce the problem of linguistic variation between
relations defined in the ontology and the extracted facts, we will use
three lexical chains (synonyms, hypernyms, and hyponyms) as
defined in WordNet. For example, the concept of ‘depict’ is matched
with ‘portray’ (synonym) and ‘represent’ (hypernym). In order to
reduce over- and under-generalisation, we will consider only one-level
of hypernyms and hyponyms when a given word is a verb.</p>
      <p>The types of information are identified by tracing the hierarchies
of hypernyms. For example, as shown in Figure 2, ‘miller’ is
extracted as the job of Rembrandt’s father since the hypernyms map to
‘worker’. Factual data, such as a date or a city name, are extracted
by using a date parsing program coupled with a simple grammar and
the hypernyms defined in WordNet. In cases, where there are
multiple matches, all relations are represented in outputs.</p>
      <p>In Figure 2, relation extraction for both clauses is determined by
the categorisation results of verbs (i.e. ‘be’ and ‘die’). The ‘be’ verb
poses a rather difficult case, since its semantic meaning is
heavily dependent on other phrases, i.e. subject and object. According
to WordNet definitions, one of its senses states ‘work in a specific
place, with a specific subject or in a specific function’. Since its
synonyms (i.e. ‘work’ and ‘follow’) are matched with ‘work’, we
exploit this relation to further examine whether or not it is related to
‘job-information’.</p>
      <p>In the second clause, since ‘die’ can be converted to the noun
format ‘death’, the verb ‘die’ matches with two potential relations
(‘date of death’ and ‘place of death’). In this case ‘date of death’
was chosen since the ‘1630’ was extracted from the same sentence
and instantiated as date information.</p>
      <p>The output from this section is an XML-formated representation
of the facts, paragraphs, sentences and keywords identified in the
knowledge extraction process. The XML files are sent to the
ontology server to populate the knowledge base.
4</p>
    </sec>
    <sec id="sec-7">
      <title>KNOWLEDGE REPRESENTATION AND</title>
    </sec>
    <sec id="sec-8">
      <title>STORAGE 4.1</title>
    </sec>
    <sec id="sec-9">
      <title>Artequakt Ontology</title>
      <p>
        An ontology is a conceptualisation of a domain into a machine
readable format [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. For Artequakt the requirement is to build an ontology
to represent the domain of artists and artefacts. This ontology is
being implemented in Prote´ge´, which is a graphical ontology editing
tool [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The main part of this ontology is being constructed from
selected sections in the CIDOC Conceptual Reference Model (CRM
- [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) ontology. CRM was developed by ICOM/CIDOC2
Documentation Standards Group to represent an ontology for cultural heritage
information. It was built to facilitate the transformation of existing
disparate museum and cultural heritage information sources into one
coherent source.
      </p>
      <p>The CRM ontology is designed to represent artefacts, their
production, ownership, location, etc. This ontology was modified for
Artequakt and is being enriched with additional classes and relationships
to represent a variety of information related to artists, their personal
http://www.cidoc.icom.org/
information, family relations, relations with other artists, details of
their work, etc. The Artequakt ontology also allows the storage of
textual paragraphs or sentences along with their source URLs so that
at a later point they can be reorganised using the ontology as a guide.
4.2</p>
    </sec>
    <sec id="sec-10">
      <title>Automatic Ontology Population</title>
      <p>
        There is an increasing interest in building ontologies to provide a
variety of knowledge services. Populating ontologies with knowledge is
labour intensive and time consuming. Semi-automatic approaches to
ontology population have been followed by for example [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] where
relationships can be added automatically between instances if these
instances already exist in the knowledge base, otherwise user
intervention will be needed. OntoAnnotate [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and OntoMat [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] are
supporting tools of user-driven ontology-based annotations, where the
produced annotations can be fed back to the ontology.
      </p>
      <p>In this project we are investigating the possibility of moving
towards a fully automatic approach of feeding the ontology with
knowledge extracted from the web. As mentioned in section 3.2, this
information is extracted with respect to the Artequakt ontology, and
provided as XML files, one per document, using tags mapped directly
from names of classes and relationships in the ontology. When a new
XML file is produced (Figure 3(a)), it will be sent to the Artequakt
ontology server which launches a program to parse the received file
and populate the ontology with the newly provided knowledge
(Figure 3(b)).</p>
      <p>The ontology server is based on Java sockets and connected to
the Artequakt knowledge base through the Prote´ge´ API. A limited
inference engine is being built on this server to allow querying and
the retrieval of specific information from the ontology, for example
to get all paragraphs that mention the date of birth of a specific artist,
get the artist of a painting, get all available facts about an artists, etc.
4.3</p>
    </sec>
    <sec id="sec-11">
      <title>Consolidating the Knowledge-Base</title>
      <p>
        When analysing web documents about selected artists, it will be
inevitable that we extract duplicated information or even contradictory
information. Handling such information is challenging for automatic
ontology population approaches. Staab et al[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] stressed the
problem of creating duplicate objects when extracting from different
documents. They relied on manually assigned object-identifiers to avoid
duplication. Our approach is attempting to identify and eliminate
duplications automatically using a two-stage consolidation process.
      </p>
      <p>The first stage is for the Artequakt ontology server to add all
extracted information to the knowledge base regardless of what is
already stored. This results in the creation of multiple instances of
artists with possibly the same information (e.g. multiple instances
of Rembrandt). The challenge is to identify which of these instances
refer to the same artist, and which ones refer to genuinely different
artists who happen to have the same name or information.</p>
      <p>The second stage is to run a consolidation process to identify
possible duplicate instances in the knowledge base, searching for clues
in the rest of information available about these instances. This is why
it is best to feed the new information to the knowledge base first
(stage 1), which provide the consolidation process with more
information to compare with.</p>
      <p>The consolidation process involves applying a set of heuristics.
Information extraction tools are sometimes only able to extract
fragments of information about an artist, especially if the source
document or paragraph is small or difficult to analyse. This results in the
creation of new instances with only one or two facts associated with
each, for example two artist instances with the name Rembrandt, but
one instance has a location relationship to Holland, while the other
has a date of birth relationship to 1609. One heuristic to apply here is
to merge such shallow instances into one instance of Rembrandt with
both location and date of birth relations, keeping the original source
URLs of each fact.</p>
      <p>Another heuristic is if two instances of same-name artists have
equal values for their date and place of birth and death relationships,
then these instances are likely to be duplicates, in which case they can
be fused together as one instance, otherwise the two instances will
stay separate. Such a heuristic helps to distinguish between
samename artists. The amount and type of information overlap between
instances can be used to calculate a confidence value to indicate
whether certain instances can be merged or left separate.</p>
      <p>Another challenge in information consolidation is to identify exact
matches. Identical information can exist in different versions. For
example consider the sentences:</p>
      <p>Rembrandt was born in the 17th century in Leiden.</p>
      <p>Rembrandt was born in 1606 in the Netherlands.</p>
      <p>Rembrandt was born on July 15 1606 in Holland.</p>
      <p>The sentences above provide similar information about an artist,
written in different formats and specificity levels. To match the above
sentences it will be necessary to enrich the current ontology with
proper temporal and geographical representations. Some format
varieties can be dealt with at the extraction level. For example the
information extraction tools being used in this project can identify and
extract dates in different formats, and provide it as day, month, year,
decade, etc. This information could be fed to the temporal ontology
and reasoned over to match between different time frames.</p>
      <p>
        There has been much work on developing databases and gazetteers
of place names, such as the Thesaurus of Geographic Names (TGN,
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]), Alexandria Digital Library (ADL, [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]), and WordNet which
also provides some geo-information. Such sources can be integrated
with the current ontology to provide knowledge on geographical
hierarchies, place name variations, and other spatial information [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-12">
      <title>NARRATIVE GENERATION</title>
      <p>
        While machines benefit from using structured ontologies to exchange
information, human beings need a more intuitive interface. One of
the most natural ways to do this is by story telling. There is a wealth
of critical and philosophical thought concerning narrative that can be
drawn on to assist in constructing a story (in this case a biography)
from the raw information gathered. Figure 4 shows one way of
viewing the layers that make up a narrative as proposed by Bal [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The
raw facts and chronological collection of events in any particular tale
is called the Fabula. For any given Fabula we could present the facts
from different perspectives and in different sequences to produce a
Story. We could then render any given Story into several different
forms or Narratives (e.g. a film or novel).
      </p>
      <p>In Artequakt the knowledge base can be thought of as our
underlying fabula. To produce the eventual narrative (in our case pages
of html) we need to first arrange sub-elements of the fabula into a
sensible sequence and produce a story.
5.1</p>
    </sec>
    <sec id="sec-13">
      <title>Biography Templates</title>
      <p>The story structures we are using are human authored biography
templates that contain queries into the knowledge base.
&lt;Paragraph&gt;</p>
      <p>&lt;url&gt;http://search.ebi.eb.com/ebi/article/
0,6101,36822,00.html &lt;/url&gt;</p>
      <p>&lt;text&gt; Rembrandt Harmenszoon van Rijn was
born on July 15, 1606, in Leiden, the
Netherlands… Rembrandt left the University of
Leiden to study painting. … He was influenced
by the work of Caravaggio and was fascinated
by the work of many other Italian artists. &lt;/text&gt;</p>
      <p>…..
&lt;sentence&gt;Rembrandt Harmenszoon van Rijn</p>
      <p>was born on July 15 1606 in Leiden
&lt;Painter&gt;
&lt;name&gt;Rembrandt Harmenszoon van</p>
      <p>Rijn&lt;/name&gt;
&lt;date_of_birth&gt;15 july 1606&lt;/date_of_birth&gt;
&lt;place_of_birth&gt;Leiden Netherlands</p>
      <p>&lt;/place_of_birth&gt;
&lt;/Painter&gt;
&lt;/sentence&gt;
…..
&lt;sentence&gt; He was influenced by the work of</p>
      <p>Caravaggio
&lt;Painter&gt;
&lt;name&gt;rembrandt&lt;/name&gt;
&lt;inspired_by&gt;Caravaggio&lt;/inspired_by&gt;
&lt;/Painter&gt;
&lt;/sentence&gt;
……
&lt;/Pragraph&gt;
(a)
(b)</p>
      <p>Implementation</p>
      <p>HTML Pages</p>
      <p>Contextual Templates</p>
      <p>Ontology + Knowledge Base</p>
      <p>
        Previous work has stored queries into an ontological space as the
destination of navigational links [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], by following the links the user
causes the queries to be executed (and views the results). With
Artequakt these basic links have evolved into more complex structures
that arrange the queries into a sequence (a biography template).
      </p>
      <p>
        The templates are written in XML using the Fundamental Open
Hypermedia Model (FOHM) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], which is capable of
representing a variety of hypermedia structures including tours and links. The
XML files are then loaded into the Auld Linky contextual structure
server [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which provides pattern matching facilities over the
structures via HTTP.
      </p>
      <p>Any given biography template may be constructed from several
sub-structures. The basic structure used is the Sequence. This
represents a list of queries that have to be instantiated from the knowledge
base and inserted into the biography in order. These queries are
authored using the vocabulary of terms defined within the ontology.
Other structures allow more complex effects. A Concept structure
contains several queries, any of which may be used at this point in
the biography. A Level of Detail (LOD) structure is similar to a
concept, but there is an ordering between the queries that corresponds
to preference (i.e. preferably the highest numbered query should be
used, if that’s not possible the next highest, and so on). These
structures may be nested (e.g. a sequence of concepts).</p>
      <p>Some queries may retrieve paragraphs directly while others query
the consolidated ontology for specific facts and construct sentences
dynamically from the results. This can be useful for facts that have
been inferred (and therefore there is no corresponding paragraph), or
when there is no paragraph that fits the literary form of the rest of the
biography (e.g. the biography is in third person, but all the available
paragraphs are in first person).</p>
      <p>The templates also contain contextual information on which parts
of the biography structure are appropriate in different contexts
(specified as a list of tag value pairs inside a context object). For
example imagine that the user has specified that they do not have a good
knowledge of artists. The template structure can specify that parts
of the structure are only available to people with a good knowledge.
Thus, when the user queries Linky for the template, the inappropriate
parts that require this are pruned away.</p>
      <p>Figure 5 shows an example structure being pruned. In this case
a query into the ontology concerning artistic influences (here it
would resolve into a sentence about Caravaggio) is removed because
it would not make sense to a user who did not have a reasonable
knowledge of artists. The resulting paragraph reads:</p>
      <p>‘Rembradt Harmenszoon van Rijn was born on July 15 1606 in
Leiden. Rembradt’s father was a miller who died in 1630. His early
work was devoted to showing the lines, light and shade, and color of
the people he saw about him.’</p>
      <p>
        In this way the biography structures will be tailored to the needs of
each individual user. For our prototype we are concentrating on broad
user classification (child/adult, etc) but it would also be possible to
incorporate more sophisticated user modelling techniques (such as
training sets [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]).
      </p>
      <p>Once it has been retrieved from Linky the template has to be
instantiated, by making each query in turn and then rendering the
results into a html page for display.</p>
      <p>1</p>
      <p>2
Rembrandt's
father was a miller
who died in 1630
3
6</p>
    </sec>
    <sec id="sec-14">
      <title>CONCLUSION &amp; FUTURE WORK</title>
      <p>In this paper we have described the basic architecture and initial work
in the Artequakt project. Our aim is to be able to generate
automatically tailored biographies from a knowledge base which has been
automatically populated by annotating text fragments extracted from
Web documents.</p>
      <p>We are currently working on completing the initial prototype
system by integrating the three main components identified in the
architecture. We will then be able to assess the effectiveness over real
data sources and begin the process of refining the constituant parts to
improve the overall quality of the biographies served by the system.</p>
      <p>Although some of the research issues in this process are
particularly challenging, the final objective is to have an architecture in
place which will allow us to explore some of the research issues that
have arisen so far in more detail; for example, more comprehensive,
automatic consolidation of knowledge bases, better techniques for
knowledge extraction and more sophisticated narrative structuring of
the knowledge fragments. To this end, progress has been made in the
identification of an approach and the building of a prototype
demonstrator for the project.</p>
    </sec>
    <sec id="sec-15">
      <title>ACKNOWLEDGEMENTS</title>
      <p>The work presented here is part of a larger project and we would
particularly like to note the contributions of Hugh Glaser,
Srinandan Dasmahapatra and David De Roure. This research is funded in
part by EU Framework 5 IST project “Artiste” IST-1999-11978,
EPSRC IRC project “Equator” GR/N15986/01 and EPSRC IRC project
“AKT” GR/N15764/01.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Alani</surname>
          </string-name>
          ,
          <source>Spatial and Thematic Ontology in Cultural Heritage Information Systems, Ph.D. dissertation</source>
          , Computer Studies Department University of Glamorgan, U.K.,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bal</surname>
          </string-name>
          , Narratology: Introduction to the
          <source>Theory of Narrative</source>
          , University of Toronto Press,
          <year>1978</year>
          . Trans. Christine van Boheemen.
          <source>Torento</source>
          .
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Cole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zaenen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Zue</surname>
          </string-name>
          .
          <article-title>Survey of the state of the art in human language technology</article-title>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Craven</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>DiPasquo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Freitag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mitchell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Nigam</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Slattery</surname>
          </string-name>
          , '
          <article-title>Learning to construct knowledge bases from the world wide web</article-title>
          .',
          <string-name>
            <surname>Artificial</surname>
            <given-names>Intelligence</given-names>
          </string-name>
          , (
          <issue>1-2</issue>
          ),
          <fpage>69</fpage>
          -
          <lpage>113</lpage>
          , (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Crofts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.M.</given-names>
            <surname>Dionissiadou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Stiff</surname>
          </string-name>
          , '
          <article-title>Definition of the cidoc object-oriented conceptual reference model'</article-title>
          ,
          <source>Technical report</source>
          , International Organization for Standardization, (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Tablan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ursu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Dimitrov</surname>
          </string-name>
          , '
          <article-title>Developing language processing components with gate (user's guide)'</article-title>
          ,
          <source>Technical report</source>
          , University of Sheffield, U.K., (
          <year>2002</year>
          ). available in http://www.gate.ac.uk/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Guarino</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Giaretta</surname>
          </string-name>
          ,
          <article-title>Ontologies and Knowledge bases: towards a terminological clarification. Towards Very Large Knowledge Bases: Knowledge Building and Knowledge Sharing</article-title>
          ., IOS Press,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Handschuh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Maedche</surname>
          </string-name>
          , '
          <article-title>Cream - creating relational metadata with a component-based, ontology-driven annotation framework'</article-title>
          ,
          <source>in In Proceedings of the First International Conference on Knowledge Capture</source>
          , pp.
          <fpage>76</fpage>
          -
          <lpage>83</lpage>
          , Canada, (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Harpring</surname>
          </string-name>
          , '
          <article-title>Proper words in proper places: The thesaurus of geographic names</article-title>
          .',
          <source>MDA Information, (3)</source>
          ,
          <fpage>5</fpage>
          -
          <lpage>12</lpage>
          , (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.L.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Frew</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , '
          <article-title>Geographic names. the implementation of a gazetteer in a georeferenced digital library</article-title>
          .',
          <string-name>
            <surname>Digital</surname>
            <given-names>Library</given-names>
          </string-name>
          , (
          <volume>1</volume>
          ), (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kahan and M.-R. Koivunen</surname>
          </string-name>
          , '
          <article-title>Annotea: An open rdf infrastructure for shared web annotations'</article-title>
          ,
          <source>in In Proceedings of The Tenth International World Wide Web Conference, WWW10</source>
          , pp.
          <fpage>623</fpage>
          -
          <lpage>632</lpage>
          , (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luparello</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Roudaire</surname>
          </string-name>
          , '
          <article-title>Automatic Construction of Personalised TV News Programs'</article-title>
          ,
          <source>in In Proceedings of the Seventh ACM Conference on Multimedia, Orlando, Florida</source>
          , pp.
          <fpage>323</fpage>
          -
          <lpage>332</lpage>
          , (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Maedche</surname>
          </string-name>
          , G. Neumann, and
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          ,
          <article-title>Bootstrapping an OntologyBased Information Extraction System</article-title>
          .,
          <source>Intelligent Exploration of the Web</source>
          , Springer / Physica Verlag,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Mancini</surname>
          </string-name>
          , 'From Cinematographic to Hypertext Narrative',
          <source>in In Proceedings of the Eleventh ACM Conference on Hypertext and Hypermedia</source>
          , San Antonio, Texas, USA, pp.
          <fpage>236</fpage>
          -
          <lpage>237</lpage>
          , (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.T.</given-names>
            <surname>Michaelides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.E.</given-names>
            <surname>Millard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.J.</given-names>
            <surname>Weal</surname>
          </string-name>
          , and D. DeRoure, '
          <article-title>Auld leaky: A contextual open hypermedia link server'</article-title>
          , in Hypermedia:Openness,
          <string-name>
            <given-names>Structural</given-names>
            <surname>Awareness</surname>
          </string-name>
          , and
          <source>Adaptivity (Proceedings of OHS-7, SC-3 and AH-3)</source>
          ,
          <source>Published in Lecture Notes in Computer Science</source>
          ,
          <source>(LNCS 2266)</source>
          , Springer Verlag,
          <source>Heidelberg (ISSN 0302-9743)</source>
          , pp.
          <fpage>59</fpage>
          -
          <lpage>70</lpage>
          , (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.E.</given-names>
            <surname>Millard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Moreau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.C.</given-names>
            <surname>Davis</surname>
          </string-name>
          , and S. Reich, '
          <article-title>FOHM: A Fundamental Open Hypertext Model for Investigating Interoperability Between Hypertext Domains'</article-title>
          , in HT00, pp.
          <fpage>93</fpage>
          -
          <lpage>102</lpage>
          , (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>G.A.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Beckwith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fellbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gross</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Miller</surname>
          </string-name>
          , '
          <article-title>Introduction to wordnet: An on-line lexical database'</article-title>
          ,
          <source>Technical report</source>
          , University of Princeton, U.S.A., (
          <year>1993</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Musen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. W.</given-names>
            <surname>Fergerson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. E.</given-names>
            <surname>Grosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. F.</given-names>
            <surname>Noy</surname>
          </string-name>
          , M. Grubez´Y, and
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Gennari</surname>
          </string-name>
          , '
          <article-title>Component-based support for building knowledgeacquisition systems'</article-title>
          ,
          <source>in In Proceedings of the Conference on Intelligent Information Processing of the International Federation for Processing World Computer Congress</source>
          , Beijing, (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Pazzani</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Billsus</surname>
          </string-name>
          , '
          <article-title>Learning and revising user profiles:the identification of interesting web sites'</article-title>
          ,
          <source>Machine Learning</source>
          ,
          <fpage>313</fpage>
          -
          <lpage>331</lpage>
          , (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L.</given-names>
            <surname>Rutledge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bailey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. V.</given-names>
            <surname>Ossenbruggen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hardman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Geurts</surname>
          </string-name>
          , '
          <article-title>Generating Presentation Constraints from Rhetorical Structure'</article-title>
          ,
          <source>in In Proceedings of the Eleventh ACM Conference on Hypertext and Hypermedia</source>
          , San Antonio, Texas, USA, pp.
          <fpage>19</fpage>
          -
          <lpage>28</lpage>
          , (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sekine</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Grishman</surname>
          </string-name>
          , '
          <article-title>A corpus-based probabilistic grammar with only two non-terminals'</article-title>
          ,
          <source>in In Proceedings of the Fourth International Workshop on Parsing Technology</source>
          , pp.
          <fpage>216</fpage>
          -
          <lpage>223</lpage>
          , (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Staab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maedche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Handschuh</surname>
          </string-name>
          , '
          <article-title>An annotation framework for the semantic web'</article-title>
          ,
          <source>in In Proceedings of the First International Workshop on MultiMedia Annotation</source>
          , Japan, (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vargas-Vera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Motta</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Domingue</surname>
          </string-name>
          , '
          <article-title>Knowledge extraction by using an ontology-based annotation tool'</article-title>
          ,
          <source>in In Proceedings of the Workshop on Knowledge Markup and Semantic Annotation</source>
          , KCAP'01,
          <string-name>
            <surname>Canada</surname>
          </string-name>
          , (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>M.J. Weal</surname>
            ,
            <given-names>G.J.</given-names>
          </string-name>
          <string-name>
            <surname>Hughes</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          <string-name>
            <surname>Millard</surname>
          </string-name>
          , and L. Moreau, '
          <article-title>Open Hypermedia as a Navigational Interface to Ontological Information Spaces'</article-title>
          ,
          <source>in In Proceedings of the Twelth ACM Conference on Hypertext and Hypermedia</source>
          , Arhus, Denmark, pp.
          <fpage>227</fpage>
          -
          <lpage>236</lpage>
          , (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>R.</given-names>
            <surname>Yangarber</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Grishman</surname>
          </string-name>
          , '
          <article-title>Machine learning of extraction patterns from unannotated corpora: Position statement'</article-title>
          ,
          <source>in In Proceedings of Workshop on Machine Learning for Information Extraction</source>
          , pp.
          <fpage>76</fpage>
          -
          <lpage>83</lpage>
          , ECAI,Berlin, (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>