<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bluima: a UIMA-based NLP Toolkit for Neuroscience</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Renaud Richardet</string-name>
          <email>renaud.richardet@epfl.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jean-Cedric Chappelier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Telefont</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Blue Brain Project, EPFL</institution>
          ,
          <addr-line>1015, Lausanne</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes Bluima, a natural language processing (NLP) pipeline focusing on the extraction of neuroscienti c content and based on the UIMA framework. Bluima builds upon models from biomedical NLP (BioNLP) like specialized tokenizers and lemmatizers. It adds further models and tools speci c to neuroscience (e.g. named entity recognizer for neuron or brain region mentions) and provides collection readers for neuroscienti c corpora. Two novel UIMA components are proposed: the rst allows con guring and instantiating UIMA pipelines using a simple scripting language, enabling non-UIMA experts to design and run UIMA pipelines. The second component is a common analysis structure (CAS) store based on MongoDB, to perform incremental annotation of large document corpora.</p>
      </abstract>
      <kwd-group>
        <kwd>UIMA</kwd>
        <kwd>natural language processing</kwd>
        <kwd>NLP</kwd>
        <kwd>neuroinformatics</kwd>
        <kwd>NoSQL</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Bluima started as an e ort to develop a high performance natural language
processing (NLP) toolkit for neuroscience. The goal was to extract structured
knowledge from biomedical literature (PubMed1), in order to help neuroscientists
gather data to specify parameters for their models. In particular, focus was set
on extracting entities that are speci c to neuroscience (like brain regions and
neurons) and that are not yet covered by existing text processing systems.</p>
      <p>
        After careful evaluation of di erent NLP frameworks, the UIMA software
system was selected for its open standards, its performance and stability, and
its usage in several other biomedical NLP (bioNLP) projects; e.g. JulieLab [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
ClearTK [22], DKPRo [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], cTAKES [28], ccp-nlp, U-Compare [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], SciKnowMine
[26], Argo [25]. Initial development went fast and several existing bioNLP models
and UIMA components could rapidly be reused or integrated into UIMA without
the need to modify its core system, as presented in Section 2.1.
      </p>
      <p>Once the initial components were in place, an experimentation phase started
where di erent pipelines were created, each with di erent components and
parameters. Pipeline de nition in verbose XML was greatly improved by the use</p>
    </sec>
    <sec id="sec-2">
      <title>1 http://www.ncbi.nlm.nih.gov/pubmed</title>
      <p>of UIMAFit [21] (to de ne pipelines in compact Java code) but ended up
being problematic, as it requires some Java knowledge and recompilation for each
component or parameter change. To allow for a more agile prototyping,
especially by non-specialist end users, a pipeline scripting language was created. It
is described in Section 2.2.</p>
      <p>Another concern was incremental annotation of large document corpus. For
example, when running an initial pre-processing pipeline on several millions of
documents, and then annotating them again at a later time. The initial strategy
was to store the documents on disk, and overwrite them every time they would
be incrementally annotated. Eventually, a CAS store module was developed to
provide a stable and scalable strategy for incremental annotation, as described
in Section 2.3. Finally, Section 3 presents two case studies illustrating the
scripting language and evaluating the performance of the CAS store against existing
serialization formats.
2</p>
      <sec id="sec-2-1">
        <title>Bluima Components</title>
        <p>Bluima contains several UIMA modules to read neuroscienti c corpora, perform
preprocessing, create simple con guration les to run pipelines, and persist
documents on the disk.
2.1</p>
        <sec id="sec-2-1-1">
          <title>UIMA Modules</title>
          <p>
            Bluima's typesystem builds upon the typesystem from JulieLab [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ], which was
chosen for its strong biomedical orientation and its clean architecture. Bluima's
typesystem adds neuroscienti c annotations, like CellType, BrainRegion, etc.
          </p>
          <p>
            Bluima includes several collection readers for selected neuroscience
corpora, like PubMed XML dumps, PubMed Central NXML les, the BioNLP
2011 GENIA Event Extraction corpus [24], the Biocreative2 annotated corpus
[
            <xref ref-type="bibr" rid="ref16">16</xref>
            ], the GENIA annotated corpus [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ], and the WhiteText brain regions corpus
[
            <xref ref-type="bibr" rid="ref8">8</xref>
            ]. A PDF reader was developed to provide robust and precise text
extraction from scienti c articles in PDF format. The PDF reader performs content
correction and cleanup, like dehyphenation, removal of ligatures, glyph mapping
correction, table detection, and removal of non-informative footers and headers.
          </p>
          <p>
            For pre-processing, the OpenNLP-wrappers developed by JulieLab for
sentence segmentation, word tokenization and part-of-speech tagging [31] were used
and updated to UIMAFit. Lemmatization is performed by the domain-speci c
tool BioLemmatizer [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ].Abbreviation recognition (the task of identifying
abbreviations in text) is performed by BIOADI, a supervised machine learning model
trained on the BIOADI corpus [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ].
          </p>
          <p>
            Bluima uses UIMA's ConceptMapper [29] to build lexical-based NERs
based on several neuroscienti c lexica and ontologies (Table 1). These lexica
and ontologies were either developed in-house or were imported from existing
sources. Bluima wraps several machine learning-based NERs, like OSCAR4
[
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] (chemicals, reactions), Linnaeus [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] (species), BANNER [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] (genes and
proteins), and Gimli [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ] (proteins).
          </p>
          <p>There are several approaches2 to write and run UIMA pipelines (see Table 2).
All Bluima components were initially written in Java with the UIMAFit library,
that allows for compact code. To improve the design and experimentation with
UIMA pipelines, and enable researchers without Java or UIMA knowledge to
easily design and run such pipelines, a minimalistic scripting (domain-speci c)
language was developed, allowing UIMA pipelines to be con gured with text les,
in a human-readable format (Table 3). A pipeline script begins with the de nition
of a collection reader (starting with cr:), followed by several annotation engines
(starting with ae:)3. Parameter speci cation starts with a space, followed by the
2 Other interesting solutions exist (e.g. IBM LanguageWare, Argo), but are not open
source.
3 If not package namespace is speci ed, Bluima loads Readers and Annotator classes
from the default namespace.
parameter name, a column and its value. The scripting language also supports
embedding of inline Python and Java code, reuse of a portion of a pipeline with
include statements, and variable substitution similar to shell scripts. Extensive
documentation (in particular snippets of scripts) is automatically generated for
all components, using the JavaDoc and the UIMAFit annotations.
2.3</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>CAS Store</title>
          <p>A CAS store was developed to persist annotated documents, resume their
processing and add new annotations to them. This CAS store was motivated by the
common use case of repetitively and incrementally processing the same
documents with di erent UIMA pipelines, where some pipeline steps are duplicated
among the runs. For example, when performing resource-intensive operations
(like extracting the text from full-text PDF articles, or performing syntactic
parsing), one might want to perform these preliminary operation once, store these
results, and subsequently perform di erent experiments with di erent UIMA
modules and parameters. The CAS store thus allows to perform the
preprocessing only once, to then persist the annotated documents, and to perform the
various experiments in parallel.</p>
          <p>MongoDB4 was selected as the datastore backend. MongoDB is a scalable,
high-performance, open-source, schema-free (NoSQL), document-oriented
database. No schema is required on the database side, since the UIMA typesystem
acts as a schema, and data is validated on-the- y by the module. Every CAS is
stored as a MongoDB document, along with its annotations. UIMA annotations
and their features are explicitly mapped to MongoDB elds, using a simple and
declarative language. For example, a Protein annotation is mapped to a prot
eld in MongoDB. The mappings are used when persisting and loading from
the database. As of this writing, annotations are declared in Java source les.
In future versions, we plan to store mappings directly in MongoDB to improve
exibility. Persistence of complex typesystem has not been implemented yet, but
could be easily added in the future.</p>
          <p>Currently, the following UIMA components are available for the CAS store:
{ MongoCollectionReader reads CAS from a MongoDB collection. Optionally,
a ( lter) query can be speci ed;
{ RegexMongoCollectionReader is similar to MongoCollectionReader but
allows specifying a query with a regular expression on a speci c eld;
{ MongoWriter persists new UIMA CASes into MongoDB documents;
{ MongoUpdateWriter persists new annotations into an existing document;
{ MongoCollectionRemover removes selected annotations in a MongoDB
collection.</p>
          <p>With the above components, it is possible within a single pipeline to read an
existing collection of annotated documents, perform some further processing, add
more annotations, and store theses annotations back into the same MongoDB
documents.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 http://www.mongodb.org/</title>
      <sec id="sec-3-1">
        <title>Case Studies and Evaluation</title>
        <p>A rst experiment to illustrate the scripting language was conducted on a large
dataset of full-text biomedical articles. A second simulated experiment
evaluates the performance of the MongoDB CAS store against existing serialization
formats.
3.1</p>
        <p>Scripting and Scale-Out
# collection reader configured with a list of files (provided as external params)
cr: FromFilelistReader
inputFile: $1
# processes the content of the PDFs
ae: ch.epfl.bbp.uima.pdf.cr.PdfCollectionAnnotator
# tokenization and lematization
ae: SentenceAnnotator
modelFile: $ROOT/modules/julielab_opennlp/models/sentence/PennBio.bin.gz
ae: TokenAnnotator
modelFile: $ROOT/modules/julielab_opennlp/models/token/Genia.bin.gz
ae: BlueBioLemmatizer
# lexical NERs, instantiated with some helper java code
ae_java: ch.epfl.bbp.uima.LexicaHelper.getConceptMapper("/bbp_onto/brainregion")
ae_java: ch.epfl.bbp.uima.LexicaHelper.getConceptMapper("/bams/bams")
# removes duplicate annotations and extracts collocated brainregion annotations
ae: DeduplicatorAnnotator
annotationClass: ch.epfl.bbp.uima.types.BrainRegionDictTerm
ae: ExtractBrainregionsCoocurrences
outputDirectory: $2</p>
        <p>Bluima was used to extract brain region mention co-occurrences from
scienti c articles in PDF. The pipeline script (Table 3) was created and tested on
a development laptop. Scale-out was performed on a 12-node (144-core)
cluster managed by SLURM (Simple Linux Utility for Resource Management). The
383,795 PDFs were partitioned in 767 jobs. Each job was instantiated with the
same pipeline script, using di erent input and output parameters. The
processing completed in 809 minutes (' 8 PDF/s).
The MongoDB CAS store (MCS) has been evaluated against 3 other available
serialization formats (XCAS, XMI and ZIPXMI). For each, 3 settings were
evaluated: writes (CASes are persisted to disk), reads (CASes are loaded from their
persisted states), and incremental (CASes are rst read from their persisted
states, then further processed, and nally persisted again to disk). Writes and
reads were performed on a random sample of 500,000 PubMed abstracts and
annotated with all available Bluima NERs. Incremental annotation was performed
on a random sample of 5,000 PubMed abstracts and incrementally annotated
with the Stopwords annotator. Processing time and disk space was measured on
a commodity laptop (4 cores, 8GB RAM).</p>
        <p>In terms of speed, the MCS signi cantly outperforms the other formats,
especially for reads (Figure 1). The MCS disk size is signi cantly smaller than XCAS
and XMI formats, but almost 4 times larger than the compressed ZIPXMI
format. The incremental annotation is signi cantly faster with MongoDB, and does
not require duplicating or overwriting les, like with the other serialization
formats. The MCS could be scaled up in a cluster setup, or using solid states drives
(SSDs). Writes could probably be improved by turning MongoDB's "safe mode"
option o . Furthermore, by adding indexes, the MCS can act as a searchable
annotation database.
4</p>
      </sec>
      <sec id="sec-3-2">
        <title>Conclusions and Future Work</title>
        <p>In the process of developing Bluima, a toolkit for neuroscienti c NLP, we
integrated and wrapped several specialized resources to process neuroscienti c
articles. We also created two UIMA modules (scripting language and CAS store).
These additions proved to be very e ective in practice and allowed us to leverage
UIMA, an enterprise-grade framework, while at the same time allowing an agile
development and deployment of NLP pipelines.</p>
        <p>In the future, we will open-source Bluima and add more models for NER and
relationship extraction. We also plan to ease the deployment of Bluima (and its
scripting language) on a Hadoop cluster.
20. Natale, D.A., Arighi, C.N., Barker, W.C., Blake, J.A., Bult, C.J., Caudy, M.,
Drabkin, H.J., D'Eustachio, P., Evsikov, A.V., Huang, H., Nchoutmboube, J.,
Roberts, N.V., Smith, B., Zhang, J., Wu, C.H.: The protein ontology: a structured
representation of protein forms and complexes. Nucleic Acids Res. 39(Database
issue), D539{545 (Jan 2011)
21. Ogren, P.V., Bethard, S.J.: Building test suites for UIMA components. NAACL</p>
        <p>HLT 2009 p. 1 (2009)
22. Ogren, P.V., Wetzler, P.G., Bethard, S.J.: ClearTK: a UIMA toolkit for statistical
natural language processing. Towards Enhanced Interoperability for Large HLT
Systems: UIMA for NLP p. 32 (2008)
23. Osborne, J., Flatow, J., Holko, M., Lin, S.M., Kibbe, W.A., Zhu, L.J., Danila, M.I.,
Feng, G., Chisholm, R.L.: Annotating the human genome with disease ontology.</p>
        <p>BMC Genomics 10(Suppl 1), S6 (Jul 2009)
24. Pyysalo, S., Ohta, T., Rak, R., Sullivan, D., Mao, C., Wang, C., Sobral, B., Tsujii,
J., Ananiadou, S.: Overview of the ID, EPI and REL tasks of BioNLP shared task
2011. BMC Bioinformatics 13(Suppl 11), S2 (Jun 2012)
25. Rak, R., Rowley, A., Black, W., Ananiadou, S.: Argo: an integrative, interactive,
text mining-based workbench supporting curation. Database: the journal of
biological databases and curation 2012 (2012)
26. Ramakrishnan, C., Baumgartner Jr, W.A., Blake, J.A., Burns, G.A., Cohen, K.B.,
Drabkin, H., Eppig, J., Hovy, E., Hsu, C.N., Hunter, L.E.: Building the scienti c
knowledge mine (SciKnowMine1): a community-driven framework for text mining
tools in direct service to biocuration. malta. Language Resources and Evaluation
(2010)
27. Ranjan, R., Khazen, G., Gambazzi, L., Ramaswamy, S., Hill, S.L., Schurmann,
F., Markram, H.: Channelpedia: an integrative and interactive database for ion
channels. Frontiers in neuroinformatics 5 (2011)
28. Savova, G.K., Masanz, J.J., Ogren, P.V., Zheng, J., Sohn, S., Kipper-Schuler,
K.C., Chute, C.G.: Mayo clinical text analysis and knowledge extraction system
(cTAKES): architecture, component evaluation and applications. Journal of the
American Medical Informatics Association 17(5), 507{513 (2010)
29. Tanenblatt, M.A., Coden, A., Sominsky, I.L.: The ConceptMapper approach to
named entity recognition. In: LREC (2010)
30. Thompson, P., et al.: The BioLexicon: a large-scale terminological resource for
biomedical text mining. BMC Bioinformatics 12(1), 397 (2011)
31. Tomanek, K., Wermter, J., Hahn, U.: A reappraisal of sentence and token splitting
for life sciences documents. Studies in health technology and informatics 129(Pt
1), 524{528 (2006)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bairoch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Apweiler</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barker</surname>
            ,
            <given-names>W.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boeckmann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gasteiger</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magrane</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The universal protein resource (UniProt)</article-title>
          .
          <source>Nucleic acids research 33(suppl 1)</source>
          ,
          <source>D154{D159</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rhee</surname>
            ,
            <given-names>S.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An ontology for cell types</article-title>
          .
          <source>Genome Biology</source>
          <volume>6</volume>
          (
          <issue>2</issue>
          ) (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bowden</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubach</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <source>NeuroNames 2002. Neuroinformatics</source>
          <volume>1</volume>
          (
          <issue>1</issue>
          ),
          <volume>43</volume>
          {
          <fpage>59</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bug</surname>
            ,
            <given-names>W.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ascoli</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grethe</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fennema-Notestine</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laird</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larson</surname>
            ,
            <given-names>S.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rubin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shepherd</surname>
            ,
            <given-names>G.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turner</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>The NIFSTD and BIRNLex vocabularies: building comprehensive ontologies for neuroscience</article-title>
          .
          <source>Neuroinformatics</source>
          <volume>6</volume>
          (
          <issue>3</issue>
          ),
          <volume>175</volume>
          {
          <fpage>194</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Campos</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>Gimli: open source and high-performance biomedical name recognition</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>14</volume>
          (
          <issue>1</issue>
          ),
          <volume>54</volume>
          (Feb
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>De</surname>
            <given-names>Castilho</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.E.</given-names>
            ,
            <surname>Gurevych</surname>
          </string-name>
          , I.:
          <article-title>DKPro-UGD: a exible data-cleansing approach to processing user-generated discourse</article-title>
          . In:
          <article-title>Onlineproceedings of the First Frenchspeaking meeting around the framework Apache UIMA, LINA CNRS UMR (</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : WordNet. Theory and Applications of Ontology: Computer Applications p.
          <volume>231</volume>
          {
          <issue>243</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>French</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lane</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavlidis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Automated recognition of brain region mentions in neuroscience literature</article-title>
          .
          <source>Front Neuroinformatics 3 (Sep</source>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Gerner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nenadic</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Linnaeus: A species name identi cation system for biomedical literature</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <volume>85</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hahn</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buyko</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomanek</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mcnaught</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsuruoka</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>An Annotation Type System for a Data-Driven NLP Pipeline (</article-title>
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hahn</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buyko</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landefeld</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Muhlhausen,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Poprat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tomanek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Wermter</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.:</surname>
          </string-name>
          <article-title>An overview of JCoRe, the JULIE lab UIMA component repository</article-title>
          .
          <source>In: Proceedings of the LREC</source>
          . vol.
          <volume>8</volume>
          , p.
          <volume>1</volume>
          {
          <issue>7</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Imam</surname>
            ,
            <given-names>F.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larson</surname>
            ,
            <given-names>S.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grethe</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandrowski</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martone</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          :
          <article-title>NIFSTD and NeuroLex: a comprehensive neuroscience ontology development based on multiple biomedical ontologies and community involvement (</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Jessop</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , et al.:
          <article-title>OSCAR4: a exible architecture for chemical text-mining</article-title>
          .
          <source>Journal of Cheminformatics</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ),
          <volume>41</volume>
          (Oct
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ohta</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tateisi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsujii</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>GENIA corpus{a semantically annotated corpus for bio-textmining</article-title>
          .
          <source>Bioinformatics</source>
          <volume>19</volume>
          , i180{
          <source>i182 (Jul</source>
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Kontonatsios</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korkontzelos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolluru</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Deploying and sharing u-compare work ows as web services</article-title>
          .
          <source>J. Biomedical Semantics</source>
          <volume>4</volume>
          ,
          <issue>7</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Krallinger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morgan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leitner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tanabe</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilbur</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirschman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valencia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Evaluation of text-mining systems for biology: overview of the second BioCreative community challenge</article-title>
          .
          <source>Genome Biology</source>
          <volume>9</volume>
          (
          <issue>Suppl 2</issue>
          ),
          <source>S1</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Kuo</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          , et al.:
          <article-title>BioAdi: a machine learning approach to identifying abbreviations and de nitions in biological literature</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>10</volume>
          (
          <issue>Suppl 15</issue>
          ),
          <source>S7 (Dec</source>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Leaman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , et al.:
          <article-title>BANNER: an executable survey of advances in biomedical named entity recognition</article-title>
          .
          <source>In: Paci c Symposium on Biocomputing</source>
          . vol.
          <volume>13</volume>
          , p.
          <volume>652</volume>
          {
          <issue>663</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , et al.:
          <article-title>BioLemmatizer: a lemmatization tool for morphological processing of biomedical text</article-title>
          .
          <source>Journal of Biomedical Semantics</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ), 3 (Apr
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>