<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Aid to spatial navigation within a UIMA annotation index</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicolas Hernandez</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>In order to support the interoperability within UIMA workows, we address the problem of accessing one annotation from another when the type system does not specify an explicit link between the two kinds of objects but when a semantic relation between them can be inferred from a spatial relation which connects them. We discuss the limitations of the framework and brie y present the interface we have developed to support such navigation.</p>
      </abstract>
      <kwd-group>
        <kwd>Apache UIMA</kwd>
        <kwd>Type System interoperability</kwd>
        <kwd>Annotation Index</kwd>
        <kwd>Spatial navigation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        One of the main ideas in using document analysis frameworks such as Apache
Unstructured Information Management Architecture 1 (UIMA) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is to move away
from handling directly the raw subject of analysis. The idea is to enrich the raw
data with descriptions which can be used as basis for the processing of
subsequent components. In the UIMA framework, the descriptions are typed feature
structures. The component developer de nes a type system which informs about
the features of a type (set of (attribute, typed value) pairs) as well as how the
types are arranged together (through inheritance and aggregation relations).
Annotations are feature structures attached to speci c regions of documents.
      </p>
      <p>In this paper, we address the problem of accessing one annotation from
another when the type system does not specify an explicit link between the two
kinds of objects but when a semantic relation between them can be inferred from
a spatial relation which connects them. The situation is a case of
interoperability issue which can be encountered when developing a component (e.g. a term
extractor) that uses analysis results produced by two components developed by
di erent developers (e.g. part-of-speech and lemma information being both hold
by distinct annotations at the same spans).</p>
      <p>
        In practice, most of the existing type systems de ne annotation types which
inherit from the built-in uima.tcas.Annotation type [
        <xref ref-type="bibr" rid="ref4 ref5 ref7">4, 5, 7</xref>
        ]. This type contains
begin and end features which are used to attach the description to a speci c
region of the text being analysed. Thanks to these features, it is possible to
cross-reference the annotations which extend this type.
      </p>
      <sec id="sec-1-1">
        <title>1 http://uima.apache.org</title>
        <p>In this paper, we argue that the Apache UIMA Application Programming
Interface (API) is not enough intuitive for a Natural Language Processing (NLP)
developer. We argue that it has some restrictions which prevent from a complete
and free navigation among the annotations. We also argue that the API is forcing
the way of developing algorithms. In section 2, we de ne the kind of spatial
navigation we would like to perform within an annotation index. In section 3,
we describe the Apache UIMA solutions to index and explore the annotations
within the indexes. In section 4, we discuss the API and show its limitation to
access annotations by spatial relations. Finally, in section 5, we brie y present the
structures and the interface we have developed to support a spatial navigation.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>The spatial navigation problem</title>
      <p>By spatial relations we mean that we assume that the annotations in a text can
be located in a two-dimensional space: One axis to represent the position in the
text linearity and an orthogonal axis to represent the covering degrees between
the annotations. Indeed annotations can cover, be covered by, precede, or follow
(contiguously or not) other annotations. The spatiality may inform about the
semantic relations. Two annotations at the same span may mean that they are
di erent aspects of the same object. They can have complementary features or
one of them can be the property of the other. One annotation covering some
others may mean that the former is made of the others, and in the opposite,
that the others are part of the former. The semantic interpretation of the spatial
relations may depend on the considered linguistic paradigm.</p>
      <p>To give examples of situations we are dealing with, let's assume the following
type system: Document, Source (information about the document such as its
URI), Sentence, Chunk, Word (having a feature whose value informs about the
lemma), POS (having a feature whose value informs about the part-of-speech)
and NamedEntity. Let's also assume that all these types do not hold explicit
references to each other and that there is no inheritance relation between them. In
that context, examples of access we would like to carry out are to get :The words
of a given sentence (can be interpreted as a made of relation); The sentence of
a given word (is part of relation); The words which are a verb (is a property
of relation); The named entities followed by a word which is a verb and has
the lemma visit (followed by relation, . . . ). Indeed, we would like to be able to
navigate within an annotation index from an annotation to its covering/covered
annotations or to the spatially following/preceding annotation of a given type
having such or such properties. We de ned this problem as a navigation problem
within an annotation index.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Accessing the annotations in the UIMA framework</title>
      <p>The problem of accessing the annotations depends on the means o ered by the
framework2 to build annotation indexes and to navigate within them.
2 See the Reference Guide http://uima.apache.org/d/uimaj-2.4.0/references.</p>
      <p>html and the Javadoc http://uima.apache.org/d/uimaj-2.4.0/apidocs.
3.1</p>
      <sec id="sec-3-1">
        <title>De ning feature structures indexes</title>
        <p>Adding a feature structure to a subject of analysis corresponds to the act of
indexing the feature structure. By default, an unnamed built-in bag index
exists which holds all feature structures which are indexed. The framework de nes
also a built-in annotation index, called AnnotationIndex, which automatically
indexes all feature structures of type (and subtypes of) uima.tcas.Annotation.
As reported in the documentation, "the index sorts annotations in the order in
which they appear in the document. Annotations are sorted rst by increasing
begin position. Ties are then broken by decreasing end position (so that longer
annotations come rst). Annotations that match in both their begin and end
features are sorted using a type priority". If no type priority is de ned in the
component descriptor3, the order of the annotations sharing the same span in the text is
unde ned in the index. The UIMA API provides getAnnotationIndex methods
to get all the annotations of that index (subtypes of uima.tcas.Annotation) or
the annotations of a given subtype. The UIMA framework allows also to de ne
indexes and possibly to sort the feature structures within them.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Parsing the annotation index</title>
        <p>The UIMA API o ers several methods to parse the AnnotationIndex. Given
an annotation index, the iterator method returns an object of the same name
which allows to move to the rst (respectively the last) annotation of the index,
the next (respectively the previous) annotation (depending on its position in
the index) or to a given annotation in the index. It is also possible to get an
unambiguous iterator to navigate among contiguous annotations in the text. In
practice, this iterator consists of getting successively the rst annotation in the
index whose begin value is higher than the end of the current one. We will call
this mechanism the rst-contiguous-in-the-index principle.</p>
        <p>The subiterator method returns an iterator whose annotations fall within
the span of another annotation. It is possible to specify whether the returned
annotations should be strictly covered (i.e. both begin and end o sets covered)
or if it concerns only its begin o set. Subiterator can also be unambiguous.
Annotations at the same span may be not returned depending on the order in
the index as well as the type priority de nition.</p>
        <p>The constrained iterator allows to iterate over feature structures which satisfy
given constraints. The constraints are objects that can test the type of a feature
structure, or the type and the value of its features.</p>
        <p>The tree method returns an AnnotationTree structure which contains nodes
representing the results of doing recursively a strict, unambiguous subiterator
over the span of a given annotation. The API o ers methods to navigate within
the tree from the root node. From any other nodes, it is possible to get the
children nodes, the next or the previous sibling node, and the parent node.
3 In a UIMA work ow, a component is interfaced by a text descriptor that indicates
how to use the component.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Limitations of the UIMA framework</title>
      <p>The de nition of an index is usually done in the component descriptor. The
de ned index can only contain one speci c type (and subtypes) of feature
structures. So, to get an index made of two distinct types, the trick would be to
declare them as subtypes of the same common type in the type system, and get
the index of this super type. This can lead to make a less consistent type system
from a linguistic point of view, but this is still coherent with the UIMA approach
of doing whatever you need in your component.</p>
      <p>The framework allows also so to sort the feature structures of a de ned
index. There are some restrictions. The sorting key, which should be a feature
of the indexed type, can only be a string or a numerical value. Only the natural
way of sorting such elements is available. There is no way to declare its own
comparator to set the order between two elements. To sort on a di erent kind of
key, the developer has to come down to the available systems. In addition, the
type system may need to be modi ed to add a feature to play the role of the
sorting key, which can also make the type system less consistent.
4.2</p>
      <sec id="sec-4-1">
        <title>Navigation limitations within an annotation index</title>
        <p>Iterator With an ambiguous iterator, the result of a move to the previous/next
annotation in the index may not correspond to the annotation which precedes/follows
spatially in the text. It can also be a covering or a covered one. In Table 1a,
the preceding of Word3 is the covering Chunk4. Unambiguous iterators force the
methods to return only spatially contiguous annotations. In practice, the method
does not always return the expected result. When called on the full annotation
index, it starts from the rst annotation in the index. In Table 1a, it only returns
the Document annotation and no more next annotation. When calling a
unambiguous iterator on a typed annotation index, the e ect of the
rst-contiguousin-the-index principle will be remarkable if some annotations occur at the same
span. In that situation, the developer has no access to all the annotations which
e ectively follow/precede spatially the current annotation. In Table 1a, an
unambiguous iteration over the Chunk type returns Chunk1, Chunk2 and Chunk3.
Chunk4 and Chunk5 are not reachable. To iterate unambiguously over
annotations of distinct types (e.g. Named Entities and POS to get the Named Entities
followed by a verb), the developer has to create a super-type over them and call
the iterator method on this super-type. The super-type may not have linguistic
consistency and the iterator will still su er from the limitation we have
previously mentioned. Another drawback of the unambiguous iterator can be noticed
when iterating an index in reverse order. If two overlapping annotations precede
the current one, the one returned will be the one whose begin o set is the
smallest and not the one with the highest end value, lower than the begin value of
the current one. The iterator follows the rst-contiguous-in-the-index principle
in the normal order. Finally, the API does not allow to iterate over the index
and in the text spatiality in the same time. It is not possible to switch from an
ambiguous iterator to an unambiguous one (and vice-versa).</p>
        <p>Subiterator is the kind of method to get the covered annotations of another
one, like the words of a given sentence. Its major drawback is that, without a
type priority de nition, there is no assurance that annotations occurring at the
same text span will t an expected conceptual order. In Table 1a, an
ambiguous subiterator over each chunk annotation for getting the words returns the
Source annotation for Chunk1, nothing for Chunk2, and the expected words
(and more to lter) for the all remaining Chunks. Concerning the unambiguous
subiterator, the rst-contiguous-in-the-index principle causes to hide some
annotations. In Table 1a, when applying an unambiguous subiterator over each
chunk, then Chunk1 and Chunk2 return the same bad result as previously. Chunk3
only returns Chunk4 and Chunk5 annotations while Chunk4 and Chunk5 return
the right word annotations. To subiterate unambiguously over a set of speci c
types, a super-type, which encompasses both the covered and the covering types,
has to be de ned in the type system. But the problem of the unambiguous
iteration remains.</p>
        <p>Constraints objects aim at testing one given feature structure at a time.
The framework does not allow to de ne dynamic constraints. This means that
the values to test cannot be instantiated relatively to the feature structure in
the index. A constraint cannot be set to select annotations whose begin feature
value is higher than the end feature value of another one. Rather, we have to
specify at the creation the exact value to be higher than. Constraints objects are
complex to understand and to set. It requires, for example, seven lines of code
for creating an iterator which will get the annotations with a lemma feature.
Constraint iterators remain iterators with the same limitations.</p>
        <p>The Tree method returns an object close to the kind of structure we would
like to manipulate to navigate within. Unfortunately, it can only give the children
of a covering annotation. So to get the parent of an annotation, a trick could be
to build the tree of the whole document by taking the most covering annotation
as the root, then to browse the tree until nding the desired annotations for
nally getting its parent. But in any case, there is no way to get directly a node
and the structure will still su er from the remarks we made about unambiguous
subiterators (consequently some annotations may not be present in the tree).</p>
        <p>Missing Methods The existing methods partially answer the problem and
some navigation methods are missing. There is no dedicated method: to
superiterate and to get the annotations covering a given annotation; to move to the
rst/next annotation of a given type (respectively the last/previous annotation
of a given type); or to get partially-covering preceding or following annotations.</p>
        <p>All these remarks lead the developers to use preferentially ambiguous
iterators and subiterators, even if, this causes to write more code to search the index
backward/forward and tests to lter the desired annotations.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Supporting the spatial navigation</title>
      <p>To support a spatial navigation among the annotations we propose to index
the annotations by their o sets in a structure called LocatedAnnotationIndex,
and to merge the annotations occurring at the same spans in a structure called
LocatedAnnotation. Table 1b illustrates the transformation of the AnnotationIndex
depicted in Table 1a into a LocatedAnnotationIndex. Figure 1 shows the spatial
links which interconnect the LocatedAnnotation.</p>
      <p>The LocatedAnnotationIndex is a sorted structure which follows the same
sorting order than the AnnotationIndex: From a given LocatedAnnotation,
covering and preceding LocatedAnnotations are located backward in the
index, and the covered and following LocatedAnnotations forward in the index.
The structure allows to access directly to a LocatedAnnotation thanks to a pair
of begin/end o sets. The rst characteristic of a LocatedAnnotation is to list all
the annotations occurring at the same o sets. This prevents from having to de ne
a type priority for handling the limitation of the subiterator. The structure comes
with several kinds of links to navigate both within the LocatedAnnotationIndex
and spatially in the text. Indeed, the structure has links to visit its spatial vicinity
(parent/children/following/preceding) LocatedAnnotation. The structure has
also links to access the previous/next element in the index. The contiguous
spatial vicinity of each LocatedAnnotation is computed when the LocatedAnnotationIndex
is built. The API also o ers some methods to dynamically search LocatedAnnotation
containing annotations of a given type among the ancestor/descendant or self.
Similarly, it is also possible to search the rst/last (respectively following/preceding)
LocatedAnnotation containing annotations of a given type.</p>
      <p>
        In terms of memory consumption, the built LocatedAnnotationIndex takes
approximatively as much memory as its AnnotationIndex; only the local
vicinity of each LocatedAnnotation is kept in memory. The CPU time for building
the index depends on the AnnotationIndex size. Some preliminary tests
indicate that the time increases by a factor of three when doubling the size of the
annotated text. It takes about 2 seconds for building the index of a 50-sentences
text analysed with sentences, chunks and words.
With the prospect of developing a pattern matching engine over annotations,
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] have addressed some design considerations for navigating annotation lattices.
They have so exposed a language for specifying spatial constraints among
annotations. An engine has been implemented within the UIMA framework. Due
to this technical choice, the design of the language and its implementation may
su er from the drawbacks we have enumerated. Indeed there is no example of
patterns which involve annotation types without inheritance relation. In
addition, as pointed out in the perspectives of the authors, it is not clear how the
engine will behave when handling multiple annotations over the same spans
without the guarantee of a consistent type priority. The LocatedAnnotation
structure is a solution to the need of de ning type priorities. More generally, the
methods of our API can play the role of the navigation devices required to the
development of a pattern matching engine.
      </p>
      <p>uimaFIT4 is a well-known library which aims at simplifying the UIMA
developments. One appealing navigation option it o ers is similar to our API.
Some methods are designated to move from one annotation to the closest
(covering/covered/preceding/following) ones by specifying the type of the annotations
to get. In practice, nevertheless, the implementation relies on the UIMA API and
may have some of the restrictions. The selectFollowing method, for example,
follows the rst-contiguous-in-the-index mechanism. In Table 1a, it returns only
the Chunk 1 to 3, and misses the 4th and 5th, when calling it successively to get
the following chunk from the rst chunk.
7</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and perspectives</title>
      <p>
        Solving the interoperability issues in the UIMA framework is a serious problem
[
        <xref ref-type="bibr" rid="ref1 ref6">1, 6</xref>
        ]. Our opinion is to give the means to developers to do what they want. We
show that the UIMA API presents some limitations regarding the spatial
navigation within annotations in a text. We also show that by adapting his problem
de nition to the framework requirements the developer may succeed to
accomplish his task. But the adaptation has a cost in development time and requires
skills in the framework. To overcome this problem, we have developed a library
which transforms an AnnotationIndex into a navigable structure which can be
used in a UIMA component. It is available in the uima-common project5. Our
perspectives are twofold: Reducing the processing time and adding a mechanism
for updating the LocatedAnnotationIndex.
      </p>
      <sec id="sec-6-1">
        <title>4 http://uimafit.googlecode.com</title>
        <p>5 https://uima-common.google.com</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kano</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McNaught</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Attwood</surname>
            ,
            <given-names>T.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Day</surname>
            ,
            <given-names>P.J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keane</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jackson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pettifer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Towards interoperability of european language resources</article-title>
          .
          <source>Ariadne</source>
          <volume>67</volume>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Boguraev</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ne</surname>
            ,
            <given-names>M.S.:</given-names>
          </string-name>
          <article-title>A framework for traversing dense annotation lattices</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>44</volume>
          (
          <issue>3</issue>
          ),
          <volume>183</volume>
          {
          <fpage>203</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ferrucci</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lally</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Uima: an architectural approach to unstructured information processing in the corporate research environment</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>10</volume>
          (
          <issue>3-4</issue>
          ),
          <volume>327</volume>
          {
          <fpage>348</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Gurevych</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , Muhlhauser,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , Muller,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Steimle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Weimer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Zesch</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          :
          <article-title>Darmstadt knowledge processing repository based on uima</article-title>
          .
          <source>In: First Workshop on UIMA at GSCL</source>
          . Tubingen,
          <string-name>
            <surname>Germany</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hahn</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buyko</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomanek</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McNaught</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsuruoka</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>An annotation type system for a data-driven nlp pipeline</article-title>
          .
          <source>In: The LAW at ACL 2007</source>
          . pp.
          <volume>33</volume>
          {
          <issue>40</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hernandez</surname>
          </string-name>
          , N.:
          <article-title>Tackling interoperability issues within uima work ows</article-title>
          .
          <source>In: LREC</source>
          . pp.
          <volume>3618</volume>
          {
          <issue>3625</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kano</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCrohon</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsujii</surname>
          </string-name>
          , J.:
          <article-title>Integrated NLP evaluation system for pluggable evaluation metrics with extensive interoperable toolkit</article-title>
          .
          <source>In: SETQA-NLP</source>
          . pp.
          <volume>22</volume>
          {
          <issue>30</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>