<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Annotation of protein residues based on a literature analysis: cross-validation against UniProtKB</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kevin Nagel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Jimeno</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tom Oldfield</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dietrich Rebholz-Schuhmann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>European Bioinformatics Institute, Wellcome Trust Genome Campus</institution>
          ,
          <addr-line>Hinxton, Cambridge, CB10 1SD</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Background: A protein annotation database, such as the Universal Protein Resource (UniProtKB), is a valuable resource for the validation and interpretation of predicted 3D structure patterns in proteins. Previously, results have been on point mutation extraction methods from biomedical literature which can be used to support the consuming work of manual database curation. However, these methods were limited on point mutation extraction and do not extract features for the annotation of proteins at the residue level. Results: This work introduces a system that identifies protein residue sites in abstract texts and annotate them with features extracted from the context. The performances of all text mining modules were evaluated against a manually annotated corpus. The identified annotation features can be attributed to at least one of six targeted categories, e.g. enzymatic reaction. Extracted results were cross-validated against UniProtKB and for 13 annotations of residues that have not been confirmed in the UniProtKB a manual assessment was performed. Conclusions: This work proposes a solution for the automatic extraction of protein residue annotation from biomedical articles. The presented approach is an extension to other existing systems in that a wider range of residue entities are considered and that features of residues are extracted as annotations.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background</title>
      <p>
        The understanding of the biological function of
proteins remains to be a central challenge in biology.
In protein science, sequence analysis of amino acids
or studies of their spatial distribution have led to
predictions and discoveries of a number of biological
significant patterns and motifs, e.g. metal-binding
sites, catalytic triads, and ligand binding sites [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7">1–7</xref>
        ].
Complementary to these mined data is the
proliferation of protein annotations by extracting
information from biomedical articles in the view of
updating existing databases. Clearly, annotations can be
used to verify data mined sequence/structure
patterns and likewise predicted patterns can be used to
search for association in the database. However, the
major annotation effort at the current stage is the
compilation of features at the protein level, while
the actual target should be at the residue level,
because biological function! s can be mapped to a
defined group of residues in proteins (function sites).
This is also reflected in the field of automatic
information extraction from literature, where solutions
have been published for the extraction of
interactions of proteins [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], subcellular protein
localisation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], pathway discovery [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], and function
annotation with Gene Ontology terminologies [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Few
groups have investigated in point mutation
extraction, but without feature extraction for residue
annotation [
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16 ref17">13–17</xref>
        ].
      </p>
      <p>
        Works have been published that focused on the
extraction of point mutations, which is one type of
a residue entity [
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16 ref17">13–17</xref>
        ]. The point mutation
extraction systems called MEMA [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and MuteXt [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]
use a dictionary lookup approach to detect protein
names and disambiguate multiple protein-residue
pairs with a word distance measurement.
MutationGraB [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the successor of MuteXt, uses a graph
bigram method to calculate the proximity by
weighting the association of word-pairs. Another
application called MutationMiner [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] focuses on the
integration of extracted point mutations into a protein
structure visualisation program.
      </p>
      <p>
        These systems are all dedicated to the
extraction of point mutations, but provide no extraction
of residue annotation. In a recent publication [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ],
an ontological model was proposed that should hold
information extracted from MutationMiner as well
as point mutation annotations. However, the author
did not provide any results of feature extraction nor
was a strategy proposed. Residue annotation
differs from functional annotation of proteins because
the biological role of a residue is described rather
in a biochemical context, which is then revealed in
the function or property of the protein. At present,
there is neither such an ontological model nor a
terminological resource publicly available.
      </p>
      <p>The goal of this research is the identification of
biological function of mined structure patterns of
proteins. For this purpose a novel approach that
combines structure mining and text mining is
proposed. The results of the combined mining study
will be published elsewhere. This paper reports
on the text mining part and introduces a strategy
for the compilation of protein residue annotations
that can be used for the interpretation of
structure patterns. The result demonstrates that
textual information can be captured and used to
augment data in UniProtKB. Because the primary data
resource is Medline, the extraction covers a broad
range of biomedical fields, but is limited to abstract
texts. The biological community benefits from the
extracted annotations, for example, in that data
mined structure patterns can be interpreted
biologically or predicted function in proteins can be better
characterised.</p>
      <p>The contribution of this work is the
automatic extraction of protein residue annotation from
biomedical articles. Contextual information are
exploited to identify features of residues that
correspond to one of six chosen target categories (SCAT,
Table 1). As a result, proteins can be selected with
residues clustered by annotation types, which can
lead to discovery of, for example, evolutionary
relationships.</p>
    </sec>
    <sec id="sec-2">
      <title>Results and Discussion</title>
      <p>The following sections assess first the extraction
system and then the extracted data.</p>
      <sec id="sec-2-1">
        <title>Evaluation of the identification systems for mentions of organism, protein and residues and their associations.</title>
        <p>In order to evaluate the performance of the NER and
the AD systems used in this study, the results were
compared against the results from manual curation
of a set of 100 Medline articles, i.e. the gold standard
corpus (GC) generated as part of this study.</p>
        <p>
          Table 2 (top) shows the performance of each
named entity recognition. With an F1 measure of
0.91 the performance of the residue tagger is within
range of previous works where only the residue was
identified as point mutation [
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16 ref17">13–17</xref>
          ]. On the other
hand, the performance of organism name recognition
was lower with precision of 0.81 and recall of 0.72.
The protein recognition has the lowest performance
(precision = 0.65, recall = 0.60 ). The relatively low
recall is due to permutation and lexical variants in
text that are not covered by the dictionaries.
        </p>
        <p>
          The evaluation of the organism-protein-residue
AD module shows that the algorithm of [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] is
suitable for association detection. The performance has
a precision of 0.83 and a recall of 0.33 (Table 2,
bottom). Two prominent reasons for the low recall is
the correct organism-protein association but with a
mismatch of protein sequence and residue, or the
association of organism and protein was wrong in the
first instance.
        </p>
        <p>The implemented association detection system is
able to extract associations in accordance to
UniProtKB.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Cross-validation of organism-protein association with UniProtKB.</title>
        <p>In this section the evaluation was performed
automatically on a cross-validation test set (XC) derived
from the UniProt corpus (UC). From the 136,566
citations listed in the UniProt a virtually complete set
of 136,559 abstract texts were retrieved from
Medline to build the UC. Subselection from UC to
determine XC resulted in 5,253 abstract texts
representing a range of diverse proteins (Table 3, top).
Corresponding to this test corpus is the set of 70,401
triplet identifiers of UniProtID-TaxonomyID-PMID
(UTP) for the protein-organism association
evaluation and 68,008 triplet identifiers of
UniProtIDResidueID-PMID (URP) for the protein-residue
association (Table 3, middle and bottom).</p>
        <p>With a precision of 0.77 and recall of 0.08 (F1
= 0.14) the result for organism-protein association
extraction indicates that although the system seems
to extract correct relations with a reasonable
number of TP the recall of the solution is too low to
fully judge on the performance. The low recall is
best explained by missing information in the
scientific documents that would confirm the
organismprotein association. The results shows that the
stringent residue-sequence match resulted in a precision
of 1.00 and recall of 0.14 (F = 0.25). The low recall
can be explained by several factors: 1) differences
between the protein sequence index between the
author and the database; 2) changes in the sequence
indexing rules by UniProtKB; 3) sequence variants
which have not been reported in the database yet;
4) false protein-organism association with the
consequence of retrieving the incorrect sequence.</p>
        <p>Notice the evaluation of the extraction system
was done on Medline abstracts for a range of diverse
proteins indexed by UniProtKB as opposed to
previous works with extraction from full texts for a few
protein family examples. Therefore the results
implicate that the extraction from only abstract texts is
possible for a number of different UniProt proteins.</p>
      </sec>
      <sec id="sec-2-3">
        <title>PDB citation enrichment.</title>
        <p>For each PDB protein entry a link to a
corresponding UniProt record is available. The AD
system extracts only relations for proteins recorded
in the UniProtKB. Therefore each Medline record
with a found o-p-r association can be added to
the citation set of the corresponding PDB entry.
At the state of this analysis, the PDB contained
42,943 PDB protein structure with a sub-fraction of
42,653 having a unique corresponding UniProt
protein identifier (11,912). For each of these proteins
the whole Medline was scanned for abstracts with
extracted organism-protein-residue associations.
Figure 1 shows the comparison of the citation sets based
on UniProtKB references and the whole Medline
analysis.</p>
        <p>For 2,535 out of 11,912 proteins the
extraction system found a total of 18,748 corresponding
PMIDs. Analysis with citation indices for this subset
of proteins revealed that 680 out of 18,748 PMIDs
were rediscoveries. The low number of rediscovery
can be explained in that many annotations are done
from sections only available in the full text.
Although the analysis was based on Medline abstract
texts, the extraction was already able to find for 21
percent of the target proteins a large number of
citations. With a precision of 0.83 (determined by
gold standard evaluation) the estimated number of
TP from the novel discovered citations is 15,560. In
context of the 16,560 references of the 2,535
proteins from UniProtKB, the extraction expands the
citation set by 1.94 fold.</p>
        <p>The extraction system can be used to expand the
citation list of UniProtKB/PDB by using only
Medline abstract texts. In this experiment the estimated
number of overlooked citations for a subset of
target proteins provide already a large set for feature
extraction for the annotation of protein residues.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Evaluation of feature extraction.</title>
        <p>The detection of domain specific features was done
by a classification approach which required a labelled
reference set and a defined set of categories. The
precision, recall and F1-measure values were calculated
for each category and summarised in Table 4. Two
sets of categories were tested, each with different but
corresponding semantic categories: (1) the six
targeted categories (SCAT) and (2) the categories listed
in the feautre table in UniProtKB (FCAT).</p>
        <p>For SCAT, the classifiers for structure
component, chemical modification, binding type yielded in
F1 measures of 0.69, 0.61, and 0.67. For FCAT the
top performing classifiers were: motif, variant, and
binding with similar F1 scores (0.62, 0.61, 0.58). The
remaining classifiers are still usable for feature
detection, as they had precision scores comparable to the
top F1 performing classifiers: enzymatic activity and
cellular phenotype from SCAT, modified residue,
active site and site for FCAT. The figures indicate that
the features used here are suitable for feature
detection and their classification. The performance
of feature detection was tested on the gold
standard corpus (GC). Sentences with residue
mentioning were examined and where applicable suitable
features were annotated manually and compared with
the extraction method. The number of validated
and non-validated features was determined and
performance measured.</p>
        <p>The performance shows that the classification
approach for feature detection had a reasonable
coverage for SCAT and FCAT (recall of 0.61 and 0.59 for
SCAT and FCAT) but is imprecise in capturing the
correct annotation (precision of 0.21 for both, Table
5). This is not surprising, considering that features
are expressed throughout the whole sentences, but
have different attachments to named entities.</p>
        <p>The association of residues and features was
based on a syntactical analysis of their verbal and
prepositional relations by using a shallow language
parser. The approach was evaluated by the
performance of detecting all manually annotated
residuefeature pairs within the GC data set. With a
precision of 0.54 and recall of 0.81 the performance of the
shallow parser suggests it is highly usable for residue
annotation extraction (Table 6). The low precision
is explained by the current implementation of the
parser which returns relations with nested
prepositional phrases, thus the calculated precision tends
to have a lower value. extraction performance
decreases when additional extraction modules (NER,
AD, FE) were used. This shows that the extraction
of annotation is greatly sensitive to each extraction
modules.</p>
        <p>Despite the performance of each module can be
improved, the result shows that the extraction
system can deliver residue annotations.</p>
      </sec>
      <sec id="sec-2-5">
        <title>Protein residue annotation extraction and comparison with UniProtKB.</title>
        <p>The extraction system in this study delivered
classified features of protein residues from Medline as
annotations. This section provides examples of the
validity of the drawn annotations by comparing
extracted information from the gold standard corpus
with entries in the UniProtKB.</p>
        <p>Within this experiment, four UniProt proteins
with a total of 19 annotations from seven sentences
and five abstract texts were mined with the
extraction system (Table 7). By comparing the mined
annotations with correspondent entries in the UniProt
six out of 19 annotations were equivalent to
existing information in the database (rediscovery).
Further, the semantic tags of the annotations, provided
by the classification of extracted text features, are
biologically meaningful. For example, “the putative
catalytic triad” is correctly tagged as enzymatic,
because it is a chemical reaction site and therefore a
requirement for enzymatic function. In this example,
the predicted semantic tag is equivalent to the
category active site from the feature table in UniProt. In
another example, “major phosphorylation sites” was
evaluated as rediscovery of the database information
“Phosphothreonine; by MAPK” and
“Phosphoserine; by MAPK” while the predicted tag (structural
com! ponent) and the assigned category in UniProt
(modified residues) are not equivalent. This is still
valid, because both pieces of information describe
the function of the residues as modification site,
while the predicted tag represented this as a
substructure and UniProt emphasises on the
modification of the residues.</p>
        <p>For the remaining 13 extracted annotations there
are no equivalent information represented in the
UniProt. All are tagged with structural component
which is biologically valid, for example, “highly
conserved C-terminal region” is an important
substructure of the protein and the extraction can aid in
determining evolutionary important residues of protein
families. However, the annotation “conserved
phosphopantothenate binding” can arguably be discussed
whether it should be tagged as structural component
or binding.</p>
        <p>In conclusion, the biological significance of the
extracted annotations were studied by comparison
with annotations from UniProt for the extracted
proteins from the gold standard corpus. From the
comparison, the rediscovery data shows that the
used SCAT scheme and its feature sets are able to
capture information correspondent to UniProt
annotations. The predicted semantic tags are biologically
valid and do not necessarily have to be equivalent to
the categories found in the database. On the other
hand, the novel discovery data indicates a potential
contribution of the extraction for the automatic
annotation of protein residues in UniProt.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>The aim of this work was to compile protein
residue features from Medline texts as annotation
for UniProtKB proteins by combining a series of text
mining methods. Although the performances of each
module may not be at optimal level, the generated
data output indicates that the strategy is able to
deliver biological meaningful results. Cross-validation
with UniProtKB analysis indicate that the
extraction contains novel information that can complement
and update the knowledge in UniProtKB and
consequently provide annotations for PDB protein
structures.</p>
      <p>It is important to note that the extraction was
done only on abstract texts from Medline. The
advantage over full text is to exploit a publicly available
broad range of scientific publications but on the cost
on the information level of abstract texts. However,
the results demonstrate that even with abstract texts
a vast amount of annotation can be obtained.</p>
      <p>As with high performing NER, AD, and FE
systems become more available, this conceptual
strategy in protein residue annotation extraction may
yield optimal results for the biological community.</p>
    </sec>
    <sec id="sec-4">
      <title>Methods</title>
      <p>The extraction of protein residue annotation from
text can be divided into three steps: 1) named entity
recognition (NER) and extraction of residue
mentions, 2) association detection (AD) of related named
entities, 3) extraction of annotation features for
associated entities.</p>
      <sec id="sec-4-1">
        <title>NER for protein and species.</title>
        <p>
          Named entity recognition for proteins was based
on an approach that combined dictionary lookup
with fuzzy matching and basic disambiguation [
          <xref ref-type="bibr" rid="ref18 ref19 ref20">18–
20</xref>
          ]. All protein names were collected from
UniProtKB/SwissProt. Names of species were extracted
from the NCBI Taxonomy references from
UniProtKB/SwissProt and then collecting scientific and
common names of the referenced organisms. The
dictionary was complemented with terminologies
describing only the referenced genus and the collection
of full organism name (genus + specie) augmented
with abbreviated genus forms (first letter
abbreviation of genus + specie). Web services for the
identification of protein names and taxa names are available
from the TM infrastructure at the EBI ( [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]).
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Identification of residue mentions from the text.</title>
        <p>
          The extraction of residue mentions follows
approaches of previous publications [
          <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
          ]. Sets
of regular expressions were constructed to identify
three types of protein residue site mentions. The
first basic type is the single protein sequence site
reference which consists of a (wild-type) amino acid
name, followed by the sequence position number
(e.g. “Gly-12”, “arginine 4”, “Tyr74”, “Arg(53)”).
A point mutation is the second type of residue site
where the description details the change of an amino
acid at given position. The common notation is
the wild-type amino acid name, the sequence
position followed by the substitution (e.g. “W77R”,
“Cys560Arg”, “ser-52-&gt;ala”, “ala2-methionine”).
Finally, the third type of residue site describes
either a list of residues or an interaction pair (e.g.
“Tyr 85 to Ser 85”, “Trp27–Cys29”). The common
notation is an amino acid name, sequence position, a
connection symbol or conn! ection word, amino acid
name, and sequence position. In addition to the
abbreviated notation residue sites can be expressed in
grammatical form (e.g. “isoleucine at position 3”,
“substitution of Ala at position 4 to Gly”, “Ser472
to glutamic acid”).
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Identification of associations between mentions of species, proteins and residues.</title>
        <p>
          The identification of a residue can only be validated,
if it is part of the protein sequence as it is reported
in a reference database (e.g., UniProtKB). This
requires that the protein mention in the text is further
supported by evidence for the species under scrutiny
to select the appropriate protein sequence from the
bioinformatics database; that excludes the risk of
using orthologous protein sequences. The
association of organisms with proteins and the proteins
with residues was done based on the algorithm
described by [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. First, specie and protein mentions
were associated by measuring the word distance
between them. Associated proteins and their specie
mention form a pair that correctly specifies the
protein with a unique identifier in the reference database
(UniProtKB). If no match was found, the
association was relaxed to genus matching resulting in a list
of protein identifiers. In case of multiple organisms
matching, word proximity metric was used to pr!
efer the closest word-pair. The identifier was used
to retrieve the protein sequence from the database
in order to validate the residue mention. According
to the algorithm proposed by [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], three cases can be
distinguished: (1) the residue correctly matches the
protein sequence, (2) several alternative sequences
are matching from a list of protein mentions
(identifiers), and (3) no match can be found for the residue
in the available protein sequences. If several protein
sequences were relevant candidates, then again the
word distance metric was used to select the closest
word pairs.
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>Feature extraction for the annotation of residues.</title>
        <p>
          The origin of a biological function of a protein is
group of residues and their experimental
characterisation are reported in scientific publications. In this
study the feature extraction process was divided into
two parts: in the first part the text was processed to
extract NPs that served as candidate features, and
in the second part the extracted candidate features
were classified into categories of annotation features.
Noun phrases are specified as nominal forms in
combination with adjective and adverb mentions (NP =
Det? (Adj—Adv—N)* N ). Even though most NPs
denote terms this is not always true [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
        </p>
        <p>
          In the first part, the abstract text was split into
sentences and annotated with part-of-speech (pos)
tags using the cistagger which has a similar
performance as the treetagger but it has an integration of
a large biomedical terminological resource. Then the
shallow parser described in [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] was applied to
extract verbal and prepositional dependencies. Since
this parser does not deal with prepositional
attachment ambiguity it has been extended with a
prepositional phrase attachment disambiguation module
explained in [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. In the second part, the features
were categorized using the endogenous classification
approach described in [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. Basically, the algorithm
relies only on the mutual information of the lexical
constituents of terms and their assigned categories.
In contrast, the exogenous (corpus-based) approach
requires large amounts of contextual cues which are
difficult to obtain. The endogenous approach is
therefore more reliable to produce results even
under conditions of sparse data. During the training
phase, lexical constituents of multi-word terms were
extracted from a labelled reference set and represent
features for a defined set of categories. The
association between both, the features and the categories,
were estimated based on their mutual information
score and the association between the multi-word
term and a category was computed as the sum of the
associations of its constituents. The categorization
of a multi-word term into one of the categories then
amounts to t! he identification of the best fitting
category for a term based on the term’s components.
The reference set for the relevant multi-word terms
was generated using maximal length noun phrase
(MLNP) analysis based on two different sets of NPs
that were extracted from an whole Medline abstract
texts analyses: the first set consists of NPs that
cooccurred with residue mentions in the same sentence
without nested residue terms (NP(not r)), and the
second set represents NPs with nested residue terms
(NP(r)); since the co-occurrence with a residue may
indicate higher relevance. Once the set of MLNPs
were extracted each NP was manually labelled
using three different categorization schemes. The first
scheme is binary labelling (BCAT) to separate
domain relevant terms from non relevant ones. The
second scheme uses six semantic categories
identified from a study on the manual categorization of
residue annotations based on scientific content from
Medline (bottom-up approach). The identified
categories and their definitions are shown in Table 1
(SCAT). The final set was defined through a
topdown approach by reusing categories described in
the feature table of the UniProtKB data resource
for proteins (FCAT).
        </p>
      </sec>
      <sec id="sec-4-5">
        <title>Generation of evaluation corpora.</title>
        <p>For the evaluation of the extraction system, two test
corpora were generated using the UniProt corpus
(UC). The UC consists of those Medline abstract
texts that are cited in the UniProt database for
relevant protein-residue pairs. The complete corpus was
automatically analysed for organism, protein and
residue mentions and tagged appropriately. A gold
standard corpus (GC) was created through manual
curation since no corpora are available. A random
sample of 100 Medline abstract texts was drawn from
the UC where every abstract had to fulfil the
condition that a mention of an organism, a protein and
a residue was present (tri-co-occurrence). All
mentions of an organism, a protein, the residue, the
associations between the mentions, and the contained
features of the residues (see above) were then
annotated manually from two independent annotators
with domain expertise. For the automatic evaluation
of extracted data a cross-validation corpus (XC) was
derived from UC, because not all database
information are necessarily expressed in abstract texts and
vice versa. Documents in UC were scanned for
trioccurrences of organism-protein-residue mentions in
text, and then analysed if the combinations of the
four identifiers
UniProtID-TaxonomyID-ResidueIDPMID can be found in the database. If at least a
single match was found the document was selected.
For the non-matching combinations the
corresponding annotations were removed from text.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Authors contributions</title>
      <p>Kevin Nagel carried out the experiments, developed
and implemented the methods, assessed the
annotations, and drafted the manuscript. Antonio
Jimeno participated in the development of the
methods and drafted the manuscript. Dietrich
RebholzSchuhmann participated in design of the
experiments, assessed the annotation and drafted the
manuscript. All authors read and approved the final
manuscript.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>We thank Kim Henrick, Michael Ashburner and Rob
Russell for their input in this project.</p>
      <p>KIT_MOUSE
NFKB1_MOUSE
BARD1_HUMAN</p>
      <p>THIO_HUMAN
VPS36_HUMAN</p>
      <p>ETXB_STAAU
Q7LZK5_BITAR</p>
      <p>VPS27_YEAST
GP1BA_HUMAN</p>
      <p>TGFA_HUMAN
XRCC1_HUMAN</p>
      <p>LCK_HUMAN
FINC_HUMAN
HFE_HUMAN
TFR1_HUMAN
ASPP2_HUMAN
UC citation
common citation
RC citation
the Gene Ontology - uncoupling the web. Novartis
Found Symp 2002.</p>
      <p>Citation sets of Uniprot proteins
0
100
200
300
400
500
600
700
800
SCAT
structure component
data
UC
UC
UC
UC
UC
UC
UC
XC
XC
XC
XC
XC
XC
XC
XC
XC</p>
      <p>XC
data
XC
data
XC
o
+
o
+
o
+
+
+
+
+
+
+
+
+
p
+
p
+
p
136,559</p>
      <p>6,532
119,880
129,792
30,732
119,653
27,709
5,253
131,306
5,253
5,253
5,253
5,253
5,253
5,253
5,253
4506</p>
      <p>TaxID
11,348</p>
      <p>0
11,348
11,328
4,743
11,328
4,740
1,536</p>
      <p>0
1,536
1,536
1,536
1,536
1,536
1,536
1,536
1301
o-p
category
structure component</p>
      <p>SCAT
recall
0.8
0.61
0.67
0.36
0.46
enzymatic
“
“
“
“
“
“
“
“
“
“
“
“
“
“
Q93K00
Q9HAB8</p>
      <p>SER19
SER19
ASP123
HIS279
ASP250
GLU55
TRP124
W612
GLY43
SER61
GLY63
GLY66
PHE230
ASN258
ASN59
ALA179
ALA180
ASP183
”
”
”
”
”
”
”
”
”
”
”
”
”
12147465 str comp
9617436</p>
      <p>str comp
12906824 str comp
”
”
”
”
”
”
”
”
”
”
”
”
”
chem mod
negative effect</p>
      <p>mutagen
extracted feature
major
phosphorylation sites for
MAPK
”</p>
      <p>FCAT
mod res
the putative cat- act site
alytic triad
” ”
”
”
putative oxyanion
hole
”
conserved ATP
binding residues
”
”
”
”
”
conserved
phosphopantothenate
binding
”
”
”
”
”
n/a
n/a
nucleophile (by
similarity)
proton acceptor (by
similarity)
proton donor (by
similarity)
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a
n/a</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barker</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thornton</surname>
            <given-names>J</given-names>
          </string-name>
          :
          <article-title>An algorithm for constraint based structural template matching: application to 3D templates</article-title>
          .
          <source>Bioinformatics</source>
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Oldfield</surname>
            <given-names>T</given-names>
          </string-name>
          :
          <article-title>Data Mining the Protien Data Bank: Residue Interactions</article-title>
          .
          <source>Proteins</source>
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Nebel</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herzyk</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gilbert</surname>
            <given-names>D</given-names>
          </string-name>
          :
          <article-title>Automatic generation of 3D motifs for classification of protein binding sites</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kristensen</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ward</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lisewski</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erdin</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fofanov</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kimmel</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavraki</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lichtarge</surname>
            <given-names>O</given-names>
          </string-name>
          :
          <article-title>Prediction of enzyme function based on 3D templates of evolutionarily important amino acids</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Polacco</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Babbitt</surname>
            <given-names>P</given-names>
          </string-name>
          :
          <article-title>Automated discovery of 3D motifs for protein function annotation</article-title>
          .
          <source>Bioinformatics</source>
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Yoon</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ebert</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chung</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DeMicheli</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Altman</surname>
            <given-names>R</given-names>
          </string-name>
          :
          <article-title>Clustering protein environments for function prediction: finding PROSITE motifs in 3D</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Stark</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sunyaev</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Russell</surname>
            <given-names>R</given-names>
          </string-name>
          :
          <article-title>A model for statistical significance of local similarities in structure</article-title>
          .
          <source>J Mol Biol</source>
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Marcotte</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xenerios</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisenberg</surname>
            <given-names>D</given-names>
          </string-name>
          :
          <article-title>Mining literature for protein-protein interactions</article-title>
          .
          <source>Bioinformatics</source>
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Blaschke</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andrade</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ouzounis</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valencia</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>Automatic extraction of biological information from scientific text: Protein-protein interactions</article-title>
          .
          <source>Proc Int Conf Intell syst Mol Biol</source>
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Stapley</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelley</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sternberg</surname>
            <given-names>M</given-names>
          </string-name>
          :
          <article-title>Predicting the subcellular location of proteins from text using support vector machines</article-title>
          .
          <source>Pac Symp Biocomput</source>
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Friedman</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kra</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krauthammer</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rzhetsky</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>GENIES: A natural-language processing system for the extraction of molecular pathways from journal articles</article-title>
          .
          <source>Bioinformatics</source>
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Blaschke</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leon</surname>
            <given-names>EA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krallinger</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valencia</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>Evaluation of BioCreAtIvE assessment of task 2</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lee</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horn</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            <given-names>F</given-names>
          </string-name>
          :
          <article-title>Automatic extraction of protein point mutations using a graph bigram association</article-title>
          .
          <source>PLoS Computational Biology</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Witte</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kappler</surname>
            <given-names>T</given-names>
          </string-name>
          :
          <article-title>Enhanced semantic access to the protein engineering literature using ontologies populated by text mining</article-title>
          .
          <source>Int. J. Bioinformatics Research and Applications</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Baker</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witte</surname>
            <given-names>R</given-names>
          </string-name>
          : Mutation Miner - Textual
          <source>Annotation of Protein Structures. CERMM Symposium</source>
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marcel</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albert</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tolle</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Casari</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirsch</surname>
            <given-names>H</given-names>
          </string-name>
          :
          <article-title>Automatic extraction of mutations from Medline and cross-validation with OMIM</article-title>
          .
          <source>Nucl. Acids Res</source>
          .
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Horn</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lau</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            <given-names>F</given-names>
          </string-name>
          :
          <article-title>Automated extraction of mutation data from the literature: application of MuteXt to G protein-coupled receptors and nuclear hormone receptors</article-title>
          .
          <source>Bioinformatics</source>
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arregui</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaudan</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirsch</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimeno</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>Text processing through Web services: Calling Whatizit</article-title>
          .
          <source>Bioinformatics</source>
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Pezik</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimeno</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            <given-names>D</given-names>
          </string-name>
          :
          <article-title>Static dictionary features for term polysemy identification. Building and evaluating resources for biomedical text mining</article-title>
          ,
          <source>LREC Workshop</source>
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Tsuruoka</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mcnaught</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananiadou</surname>
            <given-names>S</given-names>
          </string-name>
          :
          <article-title>Normalizing biomedical terms by minimizing ambiguity and variability</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2008</year>
          , 9.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Krauthammer</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nenadic</surname>
            <given-names>G</given-names>
          </string-name>
          :
          <article-title>Term identification in the biomedical literature</article-title>
          .
          <source>J Biomed Inform</source>
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Leroy</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez</surname>
            <given-names>J</given-names>
          </string-name>
          :
          <article-title>A shallow parser based on closed-class words to capture relations in biomedical text</article-title>
          .
          <source>J Biomed Inform</source>
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Schuman</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergler</surname>
            <given-names>S</given-names>
          </string-name>
          :
          <article-title>Postnominal Prepositional Phrase Attachment in Proteomics</article-title>
          .
          <source>In Proceedings of the HLT-NAACL BioNLP Workshop on Linking Natural Language and Biology</source>
          ,
          <source>Association for Computational Linguistics</source>
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Cerbah</surname>
            <given-names>F</given-names>
          </string-name>
          :
          <article-title>Exogeneous and endogeneous approaches to semantic categorization of unknown technical terms</article-title>
          .
          <source>COLING</source>
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Gaizauskas</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demetriou</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Artymiuk</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willett</surname>
            <given-names>P</given-names>
          </string-name>
          :
          <article-title>Protein structures and information extraction from biological texts: the PASTA system</article-title>
          .
          <source>Bioinformatics</source>
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Ashburner</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            <given-names>S</given-names>
          </string-name>
          :
          <article-title>On ontologies for biologists:</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Bairoch</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>The ENZYME database in 2000</article-title>
          .
          <source>NAR</source>
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <article-title>Phosphoserine; by MAPK S → E:reduces activity as a cdc2 inhibitor; when associated with E-13</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>