<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Controlled Natural Language for Semantic Annotation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Brian Davis</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pradeep Varma</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Siegfried Handschuh</string-name>
          <email>g@deri.org</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Dragan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hamish Cunningham</string-name>
          <email>hamish@dcs.shef.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Digital Enterprise Research Institute, National University of Ireland</institution>
          ,
          <addr-line>Galway</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>She eld NLP Group, University of She eld Extended Abstract</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Richly interlinked, machine-understandable data constitute the basis for the Semantic Web and by extension the Social Semantic Desktop[2]. Manual semantic annotation is a complex and arduous task both time-consuming and costly often requiring specialist annotators. (Semi)-automatic annotation tools attempt to ease this process by detecting instances of classes within text and relationships between classes, however their usage often requires knowledge of Natural Language Processing(NLP) and/or formal ontological descriptions. This challenges researchers to develop user-friendly annotation environments within the knowledge acquisition process. Controlled Natural Languages (CNL)s o er an incentive to the novice user to annotate, while simultaneously authoring, his/her respective documents in a user-friendly manner,yet shielding him/her from the underlying complex knowledge representation formalisms. CNLs have already been successfully applied within the context of ontology authoring, yet very little research has focused on CNLs for semantic annotation. We describe a user friendly semantic annotator, based on Controlled Language for Information Extraction (CLIE) tools, which permits non-expert users to semi-automatically both author and annotate meeting minutes and status reports using controlled natural language.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>Controlled Natural Languages and Semantic Annotation</title>
      <p>
        \Controlled Natural Languages are subsets of natural language whose grammars
and dictionaries have been restricted in order to reduce or eliminate both
ambiguity and complexity.3" The use of CNLs for ontology authoring and population
is by no means a new concept and it has already evolved into quite an active
research area[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A natural overlap exists between tools used for both ontology
      </p>
      <sec id="sec-2-1">
        <title>3 http://www.ics.mq.edu.au/~rolfs/controlled-natural-languages/</title>
        <p>
          creation and semantic annotation, for instance the CLIE technology permits
ontology creation and population by mapping both concept de nitions and
instances of concepts to a ontological representation using CLOnE - Controlled
Language for Ontology Editing[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. However, there is a subtle di erence between
the process of ontology creation and population and that of semantic
annotation. We describe semantic annotation as \a process as well as the outcome of the
process. Hence it describes i) the process of addition of semantic data or
metadata to the content given an agreed ontology and ii) it describes the semantic
data or metadata itself as a result of this process"[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Of particular importance
here is the notion of the addition or association of semantic data or metadata to
content - in this context a semantic note on the Semantic Desktop. As with any
annotation environment, a major drawback is that in order to create metadata
about a document, the author must rst create the content and second
annotate the content, in an additional a posteriori, annotation step. In the context of
our annotator we seek to merge both authoring and annotation steps into one.
Consequently, the user authors parts of his/her notes in CNL while
simultaneously creating relation metadata to describe its content. Very little research is
available with respect to CNLs for semantic annotation. For instance, Project
HALO4 was a research venture sponsored by Vulcan Inc5. It aimed to develop, a
\Digital Aristotle\- a comprehensive, automated tutor and research assistant.A
CNL for semantic annotation was implemented as part of the project, yet no
public material describing the CNL is available for scienti c scrutiny.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>A Use Case for Controlled Natural Language for</title>
    </sec>
    <sec id="sec-4">
      <title>Semantic Annotation</title>
      <p>CNLs cannot o er a panacea for semi-automatic annotation since it is unrealistic
to expect users to annotate every textual resource using CNL, however there are
certain use-cases where CNLs can o er an attractive alternative as a means for
semi-automatic semantic annotation, particularly in contexts, where controlled
vocabulary or terminology is implicit such as health care patient records or
business vocabulary. Our use case focuses on administrative tasks such taking
minutes during a project team meeting and weekly status reports. Very often
such note taking tasks can be repetitive and boring. In our scenario the user is a
member of a research group which in turn is part of an integrated EU research
project. Based on pre-de ned templates, the user simultaneously authors and
annotates his/her meeting minutes or status reports in CNL, using a semantic
note taking tool - SemNotes6, which is an application available for
NepomukKDE7 - the KDE instance of the Social Semantic Desktop. The metadata is
available for immediate use after creation for querying and aggregation, whereby
4 http://www.projecthalo.com/
5 http://www.vulcan.com
6 http://smile.deri.ie/projects/semn
7 http://nepomuk.kde.org/
the retrieved RDF triples can be passed to a Natural Language Generator to
produce tailored textual reports and summaries.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Implementation</title>
      <p>In our scenario, the CNL annotator is realised within a Semantic Note. The CNL
is anchored to existing semi-structured data such as a AgendaTitle, Scribe or
ActionItem based on prede ned meeting minutes or status report templates.
The annotator is based on CLIE. The CNL itself is very similar to the CLOnE
language with signi cant modi cations. The annotator architecture contains a
standard GATE pipeline8(see Figure 1) which includes the following language
processing resources: The GATE English tokeniser, the Hepple POS tagger, a
morphological analyser, a gazetteer list component for recognising useful
keyphrases, such as structured elements from the templates and reserved CNL
phrases. Any sentences for example, preceded by a Comment: element are
considered candidates for controlled language parsing. Any remaining tokens from the
CNL sentence which are not recognised as reserved CNL key-phrases are used
as names to generate links to ontological objects(See Figure 2). This is followed
by a standard Named Entity(NE) transducer in order to recognise useful NEs,
a preprocessing JAPE9 nite state transducer(FST) for identifying quoted
strings, chunking Noun Phrases(NPs) and additional preprocessing. A second
gazetteer list look up is applied which identi es trigger phrases associated with
NEs which intersect with quoted and unquoted NP annotation spans. Additional
feature values are then added to the NP chunks to indicate the appropriate class
to link an NP chunk as an instance to. The last FST parses the CNL from
the text and generates the metadata. The current tool is bootstrapped via the
Nepomuk Core Ontologies10 and currently the application creates/populates a
meeting minutes/status report ontology which references the users Personal
Information Model Ontology(PIMO) 11, via the GATE Ontology API. We intend
to modify the code to write directly to Nepomuk KDE RDF store.</p>
      <sec id="sec-5-1">
        <title>8 General Architecture for Text Engineering, See http://gate.ac.uk/ 9 Java Annotations Pattern Engine 10 http://www.semanticdesktop.org/ontologies/ 11 http://www.semanticdesktop.org/ontologies/2007/11/01/pimo/</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>
        Incentive for the user to annotate his/her respective documents plays an
important role for the realisation of both the Semantic Web and Social Semantic
Desktop. We have described a Semantic Annotator which allows non expert users
to simultaneously create content within, and add relational metadata to, notes
on the Semantic Desktop, using CNL. Furthermore, our annotator has already
been implemented and wrapped as a plugin for a semantic note taking tool.
Finally, we intend to complete the integration with Nepomuk KDE and
evaluate the user-friendliness of our annotator based on the previously successful
empirical methods employed in CLOnE [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Brian</given-names>
            <surname>Davis</surname>
          </string-name>
          , Ahmad Ali Iqbal, Adam Funk, Valentin Tablan, Kalina Bontcheva, Hamish Cunningham, and
          <string-name>
            <given-names>Siegfried</given-names>
            <surname>Handschuh</surname>
          </string-name>
          .
          <article-title>Roundtrip ontology authoring</article-title>
          . In Amit P. Sheth, Ste en Staab,
          <string-name>
            <surname>Mike Dean</surname>
            ,
            <given-names>Massimo</given-names>
          </string-name>
          <string-name>
            <surname>Paolucci</surname>
          </string-name>
          , Diana Maynard,
          <string-name>
            <surname>Timothy W. Finin</surname>
          </string-name>
          , and Krishnaprasad Thirunarayan, editors,
          <source>International Semantic Web Conference</source>
          , volume
          <volume>5318</volume>
          of Lecture Notes in Computer Science, pages
          <volume>50</volume>
          {
          <fpage>65</fpage>
          . Springer,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>S.</given-names>
            <surname>Decker</surname>
          </string-name>
          .
          <article-title>The social semantic desktop: Next generation collaboration infrastructure</article-title>
          .
          <source>Information Services and Use</source>
          ,
          <volume>26</volume>
          (
          <issue>2</issue>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Siegfried</given-names>
            <surname>Handschuh</surname>
          </string-name>
          .
          <article-title>Creating Ontology-based Metadata by Annotation for the Semantic Web</article-title>
          .
          <source>PhD thesis</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>P. R.</given-names>
            <surname>Smart</surname>
          </string-name>
          .
          <article-title>Controlled natural languages and the semantic web</article-title>
          .
          <source>Technical report</source>
          , School of Electronics and Computer Science, University of Southampton,
          <year>2008</year>
          ,(Unpublished).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>