<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Storing UIMA CASes in a relational database</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Georg Fette</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Toepfer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frank Puppe</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Comprehensive Heart Failure Center, University Hospital Wuerzburg</institution>
          ,
          <addr-line>Straubmuehlweg 2a, Wuerzburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science VI, University of Wuerzburg</institution>
          ,
          <addr-line>Am Hubland, Wuerzburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the UIMA text annotation framework the most common way to store annotated documents (CAS) is by serializing the document to XML and storing this XML in a le in the le system. We present a framework to store CASes as well as their type systems in a relational database. This does not only provide a way to improve document management but also the possibility to access and manipulate selective parts of the annotated documents using the database's index structures. The approach has been implemented for MSSQL and MySQL databases.</p>
      </abstract>
      <kwd-group>
        <kwd>UIMA</kwd>
        <kwd>data management</kwd>
        <kwd>relational databases</kwd>
        <kwd>SQL</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        UIMA [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] has become a well known and often used framework for processing text
data. The main component of the UIMA infrastructure is the CAS (Common
Analysis Structure), a data structure which combines the actual data (the text
of a document), annotations on this data and the type system the annotations
are based on. In many UIMA projects CASes are stored as serialized XML- les
in le folders with the corresponding type system le in a separate location.
In this storage mode the resource management to load which CAS with which
type system lies in the responsibility of the programmer who wants to perform an
operation on speci c documents. However, manual management of les in folders
on local machines or network folders can quickly become confusing and messy
especially when projects get bigger. We present a framework to store CASes as
well as their corresponding type systems in a relational database. This storage
mode provides the possibility to access the data in a centralized, organized way.
Furthermore the approach provides all bene ts that come along with relational
databases including search indices on the data, selective storage, retrieval and
deletion as well as the possibility to perform complex queries on the stored data
in the well known SQL language.
      </p>
      <p>The structure of the paper is as follows: Section 2 describes the related work,
Section 3 describes the technical details of the database storage mechanism,
Section 4 illustrates query possibilities using the database, Section 5 demonstrates
performance experiences with the framework and Section 6 concludes with a
summary of the presented work.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The only approach known to the best knowledge of the authors where CASes are
stored in a database is the Julielab DB Mapper [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] which serialized CASes to a
PostgreSQL database. However, the mechanism does not store the CASes' type
systems nor does it support features like referencing of annotations by features
or derivation of annotation types. Other approaches use indices to improve query
performance but do not allow to reconstruct the annotated documents from the
index (Lucene based: LUCAS [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Fangorn [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]; relational database based: XPath
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], ANNIS [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]; proprietary index based: TGrep/TGrep2 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], SystemT [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). The
indices still need the documents to be stored in the le system. Furthermore some
of the mentioned indices only allow specialized search capabilities (e.g. emphasis
on parse trees) which are provided by the respective search index and cannot
search directly on the UIMA data structures. In contrast to these approaches our
system allows searches on arbitrary type systems by formulating queries closely
related to the involved annotation and feature types.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Database storage</title>
      <p>The storage mechanism is based on a relational database for which the table
model is illustrated in Figure 1. The schema can be subdivided in a document
related part (left), an annotation instance part (middle) and a type system
related part (right). Documents are stored as belonging to a named collection and
can be manipulated (retrieved, deleted, etc.) as a group, e.g. deleting all
annotations of a speci c type. Annotated documents can be handled individually
by loading/saving a single CAS or by processing a whole collection by creating
a collection reader/writer. In either way any communication (loading/saving)
can (but need not) be parametrized so that only desired annotation types are
loaded/saved, thus speeding up processing time, reducing memory consumption
and facilitating debugging processes. A type system, instead of being stored in an
XML le and containing a xed type system, can be retrieved from the database
in di erent task speci c ways. One way is by requesting the type system which
is needed to load all the annotated documents belonging to a certain collection.
Other possibilities are by providing a set of desired type names or by providing
a regular expression determining all desired type names. The storage mechanism
is able to store the inheritance structures of UIMA type systems as well as
referencing of annotations by features of other annotations. For further information
on the technical aspects we refer to the documentation of the framework3.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Querying</title>
      <p>A bene t from storing data in an SQL database is the database index and the
well established SQL query standard. The database can be queried for counts of
occurrences of speci c annotation types, counts of covered texts of annotations
or even complex annotation structures in the documents. We want to exemplify
this with a query on documents which have been annotated with a dependency
parser using the type system shown in Figure 2.
3 http://code.google.com/p/uima-sql/</p>
      <p>To query for all words governing the word walk, we have to look for tokens
with the desired covered text, nd the tokens governing those tokens and return
their covered text. The SQL command for this task is shown in Figure 3. An
abstraction layer to cover the complexity could be put on top (like a graph querying
language), but even in the presented way with standard SQL the capabilities of
the database engine can serve as a useful tool to improve corpus analysis.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Performance</title>
      <p>To run a performance test on the storage engine we created a corpus of 1000
documents, each consisting of 1000 words. The words were taken from a
dictionary of 1000 randomly created words, each of 8 characters length. From each
document we created a CAS and added annotations so that each word was
covered, with the annotations covering 1 to 5 successive words. Each annotation
was given two features, one String feature with a value randomly taken from the
word dictionary and a Long feature containing a random number. All documents
were stored and then loaded again. This was done with the database engine as
well as with a local le folder on the same hard drive the database les were
located on. In a second experiment the same documents where loaded again
and we added an annotation of another type with a Long feature containing a
random number to each document. After adding the additional annotation the
documents were stored again. In a third experiment we wanted to query for the
frequencies of annotations covering each of the words from the word dictionary.
For le system storage this was done by accumulating the annotation counts
during an iteration over all serialized CASes, for database storage this was done
by performing a single SQL query for each of the words from the dictionary.</p>
      <p>In Table 1 we can observe that the time needed for database storage is quite
long but reading is as fast as from the le system. Storing to the database during
the second experiment was faster than in the rst one, because this time only the
additional annotations had to be incrementally stored. Storage to the le system
again performed about ve times faster than to the database but the bene t of
being able to incrementally store only the additional annotations can be clearly
observed. Physical storage space consumption is larger for database storage but
that shouldn't pose a major problem as hard disc space is not an overly expensive
resource nowadays. Query performance in the database is about 20 times faster
than using le system storage illustrating the bene t of the database approach.</p>
      <p>DB
FileSystem</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>We have presented a framework to store/retrieve CASes and perform analysis
queries on them using a relational database. We examined the save, load and
query speed compared to regular le based storage and presented examples how
to use the database index structures to analyze annotations in the corpus. We
hope to be able to improve the storage speed of the database engine so that the
choice between le system storage and database storage will not be in uenced
by the still quite large di erence in speed performance.</p>
      <p>This work was supported by grants from the Bundesministerium fuer Bildung
und Forschung (BMBF01 EO1004).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>Designing and evaluating an xpath dialect for linguistic queries</article-title>
          .
          <source>In: 22nd International Conference on Data Engineering</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ferrucci</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lally</surname>
            ,
            <given-names>A.D.A.M.:</given-names>
          </string-name>
          <article-title>Uima: an architectural approach to unstructured information processing in the corporate research environment</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>10</volume>
          (
          <issue>3-4</issue>
          ),
          <volume>327</volume>
          {
          <fpage>348</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ghodke</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Fangorn: A system for querying very large treebanks</article-title>
          .
          <source>In: COLING (Demos)</source>
          . pp.
          <volume>175</volume>
          {
          <issue>182</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hahn</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buyko</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landefeld</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Muhlhausen,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Poprat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tomanek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Wermter</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>An overview of JCoRe, the JULIE lab UIMA component repository</article-title>
          .
          <source>In: LREC'08 Workshop `Towards Enhanced Interoperability for Large HLT Systems: UIMA for NLP`</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Krishnamurthy</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reiss</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vaithyanathan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
          </string-name>
          , H.:
          <article-title>Systemt: a system for declarative information extraction</article-title>
          .
          <source>SIGMOD Rec</source>
          .
          <article-title>(</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Rohde</surname>
            ,
            <given-names>D.L.T.</given-names>
          </string-name>
          :
          <article-title>Tgrep2 user manual (</article-title>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Zeldes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Ludeling,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Ritz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Chiarcos</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Annis: a search tool for multi-layer annotated corpora</article-title>
          .
          <source>In: Proceedings of Corpus Linguistics</source>
          <year>2009</year>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>