<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mutation tagging with gene identifiers applied to membrane protein stability prediction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rainer Winnenburg</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Conrad Plake</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Schroeder Biotec</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>TU Dresden</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Germany ms@biotec.tu-dresden.de</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>s. Identifying the correct gene improved from 77% to 91% when considering the mutations.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>To demonstrate practical relevance, we set
up a mutation screening for five
membrane proteins from the family of G
proteincoupled receptors to evaluate a solvation
energy based model for the prediction of
stabilising regions in membrane proteins. We
identified 35 mutations in text. 25 out of
35 mutation phenotypes reported in
literature were in compliance with the prediction
of the energy model, which supports a
relation between mutations and stability issues
in membrane proteins.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        Proteins carry out most cellular functions as they are
acting as building blocks for structures, enzymes,
gene regulators, and are involved in cell mobility
and communication
        <xref ref-type="bibr" rid="ref1">(Alberts et al., 2002)</xref>
        . Proteins
may interact briefly with each other in an enzymatic
reaction, or for a long time to form part of a
protein complex. The interactions between proteins
are of central importance for almost all processes
in living cells, and are described by numerous
distinct pathways in databases such as KEGG
        <xref ref-type="bibr" rid="ref23">(Ogata et
al., 1999)</xref>
        . Malfunctions or alterations in such
pathways can be the cause of many diseases, when for
instance the biosynthesis of involved proteins is
repressed or proteins are not interacting the way they
should. The latter can be due to structural changes
in one of the interacting proteins, caused by point
mutations, i.e. single wild type amino acid
substitutions. Indeed, it is already well known that such
mutations are the cause of many hereditary diseases.
Thus the large-scale analysis of point mutation data
in combination with information about protein
interactions, protein structure and disease pathogenesis,
might facilitate the study of still unresolved
phenotypes and diseases.
      </p>
      <p>It is envisaged to provide an automated system
for the interpretation of structure-function relations
in the context of genetic variability data.
Despite the availability of numerous biomedical data
collections, valuable information about
mutationphenotype associations is still hidden in
nonstructured text in the biomedical literature. Thus text
mining methods are implemented to automatically
retrieve these data from the 18 millions of literature
references in PubMed. The extracted knowledge
will be stored in one homogeneous data store and
integrated with already available data from suitable
databases. On the basis of all these combined data,
new hypotheses can be formulated, like the
prediction of phenotypic effects induced by mutations. At
the moment, we are populating a database with
organism specific protein-mutation associations which
we envisage to apply on diverse biological
problems, such as the detection of mutation centred
genedisease associations in human.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Background</title>
      <p>Genomic variation data has already been collected
for many years. Single nucleotide polymorphisms
(SNPs), which make up about 90% of all human
genetic variation and occur every 100 to 300 bases
along the 3-billion-base human genome, are
available as large collections. Single amino acid
polymorphisms (SAPs) are often manually extracted
from literature and curated into databases,
originating from wet lab experiments. Additionally, some
structures of such mutations may be revealed in
crystallography experiments and might eventually
end up as distinct structures in the Protein Database
PDB. Of particular interest is the identification of
mutations which have a strong influence on the
stability of proteins. Therefore, the biomedical
literature can be systematically searched for
information about mutation-phenotype associations by text
mining, which may lead to new insights beyond
information in existing databases. For the text mined
data it is additionally possible to weight or prioritise
information according to their publication date, the
involved authors and the journal. Considering these
meta data can be relevant if for instance an already
published assumption has been proven wrong in a
more recent publication, or for determining whether
a protein is a hot topic or if the information is
already available for years. Furthermore, it is
possible to receive a more detailed view on a protein’s
characteristics, e.g. if a certain interaction only takes
place under specific conditions, or if an interaction is
prevented by the conformational change of a protein
domain triggered by a point mutation.
2.1</p>
    </sec>
    <sec id="sec-4">
      <title>Databases</title>
      <p>
        Data on mutations have been collected for years, for
numerous species and by different organisations for
diverse purposes. There are many efforts to cope
with the data, which is being made available in a
growing number of databases. The Human Genome
Variation society
        <xref ref-type="bibr" rid="ref14 ref4">(Horaitis and Cotton, 2004)</xref>
        promotes the collection, documentation and free
distribution of genomic variation information. New
mutation databases are reported in the Journal Human
Mutation on a regular basis. There are manually
curated databases like OMIM
        <xref ref-type="bibr" rid="ref13">(Hamosh et al., 2002)</xref>
        ,
UniProt Knowledgebase
        <xref ref-type="bibr" rid="ref29 ref30">(Yip et al., 2008; Yip et al.,
2007)</xref>
        , and general central repositories like the
Human Gene Mutation Database
        <xref ref-type="bibr" rid="ref26">(Stenson et al., 2008)</xref>
        ,
Universal Mutation Database
        <xref ref-type="bibr" rid="ref3">(Broud et al., 2000)</xref>
        ,
Human Genome Variation Database
        <xref ref-type="bibr" rid="ref10">(Fredman et al.,
2004)</xref>
        , MutDB
        <xref ref-type="bibr" rid="ref25">(Singh et al., 2007)</xref>
        .
      </p>
      <p>
        Besides these central repositories, there are small
specialised databases, such as the infevers
autoinflammatory mutation online registry
        <xref ref-type="bibr" rid="ref22">(Milhavet et al.,
2008)</xref>
        , the GPCR NaVa database for natural variants
in human G protein-coupled receptors
        <xref ref-type="bibr" rid="ref16">(Kazius et al.,
2007)</xref>
        , or the Pompe disease mutation database with
107 sequence variants
        <xref ref-type="bibr" rid="ref18">(Kroos et al., 2008)</xref>
        .
      </p>
      <p>In contrast, unpublished SNPs normally make
their way into large locus specific data repositories.
Since August 2006, there is a wiki based approach
SNPedia in contrast to classical databases collecting
information on variations in human DNA.
2.2</p>
    </sec>
    <sec id="sec-5">
      <title>Text mining</title>
      <p>
        Despite the availability of numerous biomedical data
collections, valuable information about
mutationphenotype associations is still hidden in
nonstructured text in the biomedical literature. Thus
text mining methods are implemented to
automatically retrieve these data from the 18 millions of
referenced articles in PubMed. Text mining aims to
automatically extract and combine information spread
in several natural language texts and by this
generating new hypotheses. One of the key prerequisites for
finding new facts (e.g. interactions or mutations) is
the named entity recognition (NER) in text, the
assignment of a class to an entity (e.g. protein), as well
as a preferred term or identifier, in case an entry in
a database, such as UniProt, or a controlled
vocabulary like the Gene Ontology (GO)
        <xref ref-type="bibr" rid="ref2">(Ashburner et al.,
2000)</xref>
        exists. For the task of named entity
recognition usually a dictionary is used, which contains a
list of all known entity names of a class (e.g. human
proteins) including synonyms. For the recognition
of patterns (e.g. database identifiers like NM 12345)
regular expression can be defined. For the
analysis of whole sentences, Natural language processing
(NLP) techniques are used, which aim to understand
text on a syntactic and semantic level. This approach
is often paired with systems which are based on a
set of manually defined rules or which make use of
(semi-)supervised machine learning algorithms.
      </p>
      <p>
        Up to now, there have already been diverse
examples for the successful application of text mining to
the mutation retrieval task. Early examples are the
automatic extraction of mutations from Medline and
cross-validation with OMIM
        <xref ref-type="bibr" rid="ref24">(Rebholz-Schuhmann
et al., 2004)</xref>
        , and the work by
        <xref ref-type="bibr" rid="ref14 ref4">(Cantor and Lussier,
2004)</xref>
        , who mined OMIM for phenotypic and
genetic information to gain insights into complex
diseases. More recently,
        <xref ref-type="bibr" rid="ref5 ref6">(Caporaso et al., 2007b)</xref>
        applied their concept recognition system based on
regular expressions on mutation mining task, and the
automatic Extraction of Protein Point Mutations
Using a Graph Bigram association
        <xref ref-type="bibr" rid="ref20">(Lee et al., 2007)</xref>
        was reported to find reliably gene-mutation
associations in full text. For identifying gene-specific
variations in biomedical text,
        <xref ref-type="bibr" rid="ref17">(Klinger et al., 2007)</xref>
        integrate the ProMiner system developed for the
recognition and normalisation of gene and protein names
with a conditional random field (CRF)-based
recognition system. As an answer to the diverse
approaches developed over the past years, a framework
for the systematic analysis of mutation extraction
systems was proposed by
        <xref ref-type="bibr" rid="ref15 ref20 ref27 ref29 ref9">(Witte and Baker, 2007)</xref>
        .
      </p>
      <p>
        More and more groups are working on
mutations in proteins and their involvement in
diseases.
        <xref ref-type="bibr" rid="ref15">(Kanagasabai et al., 2007)</xref>
        developed
mSTRAP (Mutation extraction and STRucture
Annotation Pipeline), for mining mutation annotations
from full-text biomedical literature, which they
subsequently used for protein structure annotation and
visualisation.
        <xref ref-type="bibr" rid="ref28">(Worth et al., 2007)</xref>
        use structure
prediction to analyse the effects of nonsynonymous
single nucleotide polymorphisms (nsSNPs) with regard
to diseases. Focussing on Alzheimer’s disease,
        <xref ref-type="bibr" rid="ref15 ref20 ref27 ref29 ref9">(Erdogmus and Sezerman, 2007)</xref>
        extract mutation-gene
pairs, with estimated 91.3%, and precision at 88.9%.
        <xref ref-type="bibr" rid="ref19">(Lage et al., 2007)</xref>
        realised a human
phenomeinteractome network of protein complexes
implicated in genetic disorders by by integrating
qualitycontrolled interactions of human proteins with a
validated, computationally derived phenotype
similarity score,
3
      </p>
    </sec>
    <sec id="sec-6">
      <title>Methods</title>
      <p>Through the combination of different data from
literature and databases it is possible to derive new
facts, e.g. novel gene-disease associations or the
influence of mutations on protein-protein interactions.
The approach is designed in such a way, that it can in
principle be applied to any kind of genetic data for
answering disease centred questions. For the
moment, we concentrate on collecting available high
quality data on protein point mutations from curated
databases and from peer-reviewed literature. For the
latter we will present a flexible approach for both the
specific and high-throughput retrieval of mutations.
In detail, the following tasks have to be performed:
(1) Identify genes/ proteins in abstracts. (2) From
this subset consider only these which additionally
contain information about mutations. (3) Propose
potential protein - mutation pairs. (4) Filter
proposed pairs by sequence compliance. (5) Utilise
this information for the refinement of the original
gene/protein identifier.
3.1</p>
    </sec>
    <sec id="sec-7">
      <title>Entity recognition</title>
      <p>
        Gene normalisation This module allows for the
automated named entity recognition of genes and
proteins. Our approach performs gene name
disambiguation by using background knowledge to
match a gene with its context against the text as a
whole
        <xref ref-type="bibr" rid="ref11">(Hakenberg et al., 2007)</xref>
        . A gene’s context
contains information on Gene Ontology annotations,
functions, tissues, diseases etc. extracted from the
databases Entrez Gene and UniProt. A comparison
of gene contexts against the text gives a ranking of
candidate identifiers and the top ranked identifier is
taken if it scores above a defined threshold. This
approach has been recently extended for inter-species
normalisation and achieves 81% success rate on a
mixed dataset of 13 species
        <xref ref-type="bibr" rid="ref12">(Hakenberg et al., 2008)</xref>
        .
Mutation tagging We implemented an entity
recognition algorithm (MutationTagger) to
automatically extract protein point mutation mentions from
PubMed abstracts. Wild-type and mutant amino
acid, as well as the sequence position of the
substitution are extracted by means of both a set of regular
expressions for pattern recognition of 1 or
3-letternotations (e.g. E312A or Glu(312)→Ala), and rules
for the more complex identification of textual
mutation descriptions (e.g. Glu312 was replaced with
alanine). Problems concerning the full text
representations (detecting the correct sequence position
of the mutated residue and unravelling
enumerations) have been addressed by additional extraction
algorithms and the implementation of a sequence
check. An evaluation of our method on the test
data from MutationFinder
        <xref ref-type="bibr" rid="ref25 ref27 ref28 ref5 ref6">(Caporaso et al., 2007a)</xref>
        showed comparable success rates of around 89%
Fmeasure for mutation mention extraction.
3.2
      </p>
    </sec>
    <sec id="sec-8">
      <title>Association of entity pairs</title>
      <p>In the process of recognising mutations in text, the
normalisation, i.e. the direct association to specific
proteins, remains a challenge. This is due to the fact
that the abstracts of relevant publications typically
mention more than only one single mutation and
protein. Thus, a mutation-protein association purely
based on their co-occurrence in one abstract is not
sufficient, as it would result in a permutation with a
huge number of false positive predictions. The
problem becomes even more evident, when considering
that both gene and mutation tagging are imperfect,
achieving a precision of 80 to 90% each.</p>
      <p>
        A method is desired, that both disambiguates the
relations of candidate mutations and proteins, and
filters out false positives from the underlying
individual mutation and protein recognition tasks. There
are approaches which apply a word distance
metric for assigning a mutation to its nearest occurring
protein term, which is error prone, as matching
mutation and protein do not necessarily have to occur
close to each other in the abstract or even in the
same sentence. The statistical approach GraB is an
excellent tool for the automatic extraction of
Protein Point Mutations using a Graph Bigram
association
        <xref ref-type="bibr" rid="ref20">(Lee et al., 2007)</xref>
        , achieving good results for
most likely mutation-protein association but alone
would also not fulfil the second aspect of filtering
out false positives.
      </p>
      <p>
        Sequence Checks Mutations are commonly
described as the substitution of a wild-type by a
mutant amino acid at a given position. Our method
compares the wild-type residue as described in a
mutation mention with the UniProt/Swiss-Prot and
PDB protein sequences for all candidate proteins.
It is important to incorporate sequences from both
repositories, as the sequence numbering can differ
and it is not always evident from a publication’s
abstract, which numbering the mutation notation refers
to. To map UniProt IDs to PDB and vice versa, we
used PDB cross-references in
UniProtKB/SwissProt from http://beta.uniprot.org/docs/pdbtosp
and the residue specific comparison between
PDB and SwissProt sequences as provided by
http://www.bioinf.org.uk/pdbsws/
        <xref ref-type="bibr" rid="ref21">(Martin, 2005)</xref>
        .
Only associations between mutations and proteins
with matching sequences are considered.
3.3
      </p>
    </sec>
    <sec id="sec-9">
      <title>Annotation pipelines</title>
      <p>The developed mutation retrieval pipeline can be
accessed through two different interfaces (see
Figure 1), which offer dependent on the annotation task,
either a systematic or quick and flexible solution.
The following approaches have been implemented:
• Organism-centred approach (database)
All available mutations for a given organism
will be retrieved in one single literature
screening and stored in the Mutation database. This
approach relies on the large-scale identification
of gene mentions in PubMed abstracts, which
have to be compiled for organisms of interest
prior to a mutation screening. As of now, gene
mention data is available for human, mouse,
and yeast. However, data for additional
relevant organisms will be added on a regular basis
in the near future.
• Protein-centred approach (on-the-fly)</p>
      <p>It is possible to retrieve relevant data for a
single gene or a list of genes/ proteins for any
organism. For this purpose, the gene
identification part performed by the gene normaliser
is replaced by a direct full text search in the
PubMed library using the Entrez Programming
Utilities. Again, the result is a set of abstracts,
which is subsequently processed by the
MutationTagger.</p>
    </sec>
    <sec id="sec-10">
      <title>3.4 Improvement of gene normalisation</title>
      <p>As described above, we defined the input set of
documents for the organism-centred mutation mining
approach by scanning the whole PubMed database
for abstracts mentioning at least one gene or protein
of a pre-defined species. For this filtering step, we
relied on the gene normalisation techniques of our
gene normaliser, which was applied to all PubMed
abstracts in advance and has shown 85% F-measure
for human genes and slightly lower for other species.
However, the gene normalisation proposes by
default only one single identifier per gene mention,
even if a set of different candidate identifiers was
computed. According to internal ranking
mechanisms, only the top scoring candidate is
considered. This leads to a possible scenario, where in
some cases the correct identifier is ranked lower and
would be neglected for any subsequent data
procession. In case of our mutation mining algorithm, we
assume that some mutations cannot be associated to
the correct protein, because the gene tagging task
already failed.</p>
      <p>On the other hand, it should be possible to
improve the performance of both entity recognition
techniques for genes and mutations by combining
the results. The idea is to run both approaches with
low precision thus receiving a high recall,
permutate all elements of both sets, and then consider
the intersection of all combinations that fit.
Mutation and gene product are considered to be a valid
pair, if the wild-type residues at the mutated
position in the protein sequence and in the reported
mutation match (as described in section 3.1). For all
proposed gene identifiers, protein sequences are
obtained and checked for compliance with the reported
wild type amino acid. The score of identifiers that
show a match are increased, which might lead to
a re-ranking of the identifiers for one gene entity.
This could further improve the original gene
normalisation approach for candidate entities which are
reported to show a mutation.</p>
      <p>Example As shown in Figure 2 our gene normaliser
identified CCP (human crystallin, gamma D;
EntrezGene ID 1421) as the top candidate gene name for
abstract PMID 8142383. The mutation tagger
identified a replacement of tryptophan with glycine at
position 191 as the only mutation mentioned in the
paper. None of the protein sequences retrieved for
human CCP showed a tryptophan residue at position
191, which means that this gene identifier was not
supported by mutation information. However,
besides human crystallin, there was also
cytochromec peroxidase in yeast (EntrezGene ID 853940)
proposed as an alternative identifier, which received a
lower score. As the product of this gene showed
a tryptophan residue at postion 191 (according to
PDB sequencing) the score was increased making
it the new top candidate. Indeed, manual curation
of the corresponding literature confirmed, that the
only gene mentioned in the abstract is cytochrome-c
peroxidase in yeast. The same positive re-ranking
finding the correct gene identifier through
mutation information was shown for human TP53 in
paper 11254385, and human amylase alpha in paper
15182367.
4</p>
    </sec>
    <sec id="sec-11">
      <title>Results</title>
      <p>Mutation database In order to establish a
mutation database, which will eventually store all protein
point mutations mentioned in PubMed abstracts for
all organisms of interest, a first platform has been
realised, comprising a MySQL database, which can
be accessed by a web-interface.</p>
      <p>To populate the database, in a first step the
PubMed corpus is filtered for abstracts mentioning
at least one gene or protein using the named entity
recognition algorithm as described in Section 3.1,
which is currently working for the three organisms
human, mouse, and yeast. This led to a set of set of
3,443,566 abstracts proposing more than 10 millions
of potential protein candidates. In a second step, the
mutation extraction algorithm is applied on this
corpus and the retrieved information is transferred into
the database. In total, 258,511 mutations were found
in 78,968 abstracts. Subsequently, for all candidate
genes found in these abstracts, the corresponding
sequences are obtained and checked for compliance
with the wild type amino acid at the position of
the mentioned mutation, which led to a number of
877,183 potential protein - mutation pairs. Out of
these, 127,384 are supported by sequence (74,722
if multiple mentions of the same mutation in one
abstract are counted as one) in contrast to 131,127
(77,643) mutations which have not passed the
sequence filter. In summary, from all mutations
identified by the plain algorithm, about 49% could be
supported by gene associations based on sequence
check. These data were retrieved from 41,384 (52%)
abstracts in total.</p>
      <p>
        Evaluation We evaluated our approach on two
different tasks: pure identification of a mutation in
a text, and the identification of correct
mutationprotein pairs. An evaluation of our method on
the test data from MutationFinder
        <xref ref-type="bibr" rid="ref25 ref27 ref28 ref5 ref6">(Caporaso et al.,
2007a)</xref>
        showed comparable success rates of around
87% F-measure for pure mutation mention
extraction. On the document level, from 182 abstracts
containing mutations, 163 were identified, in 4 abstracts
mutation were wrongly predicted. On the mutation
level 741 out of 907 were identified alongside 61
false positives.
      </p>
      <p>To assess the refinement possibilities for falsely
top ranked gene names, from the 182 abstracts we
took the subset of those, the gene normaliser
identified genes from one of the 10 supported species:
human, mouse, yeast, rat, fruit fly, H. pylori, S. Pombe,
C. Elegans, A. Thaliana, and D. Rerio. This led to
a subset of 22 abstracts. In the initial run, the gene
name identifier identified in 17 of 22 abstracts (77%)
the correct gene as the top ranked candidate.
However, after the gene tagging refinement by applying
the sequence filter to all candidate genes, the genes
of 3 more papers were identified correctly replacing
the original and false top candidate. This led to the
correct protein normalisation for 20 out of 22 (91%)
publications. For the remaining 2 publication, the
correct genes could not be identified, as they were
from species, the gene identifier does not yet
support. The suggested genes from mouse were first
falsely predicted, which were then not supported by
the sequence checks. By this the proposed
identifiers were brought below the threshold, resulting in
no gene identification at all for these 2 abstracts and
turning the 2 “false positives” to “false negatives”.
On-the-fly vs. database approach We evaluated
the results of the two access approaches (database
and on-the-fly) for human Aquaporin-1, as part of
the stability analysis of protein membranes (see
Section 5). The precision of the on-the-fly approach is
expected to be lower, as the first step is more general
due to relying on full text searches instead of entity
recognition. Indeed, in comparison to the unique 20
mutations found by the organism-centred approach,
9 additional mutations were found, of which all were
false positives, actually appearing in Aquaporin-2 or
4. This supports the good precision of the named
entity approach for the gene normalisation.
5</p>
    </sec>
    <sec id="sec-12">
      <title>Application</title>
    </sec>
    <sec id="sec-13">
      <title>Predicting effects of mutations based on sequence</title>
      <p>
        Integral membrane proteins play an important role
in all organisms, especially as transporters. Due to
their striking importance, mutations in membrane
proteins are known to be the cause of many
hereditary diseases, such as cystic fibrosis, or retinitis
pigmentosa. The reason are often conformational
changes in proteins, which may lead to malfunction
of a whole protein complex. Unfortunately,
identified structures for membrane proteins are still rare.
For this reason, we used a coarse grained model
presented by
        <xref ref-type="bibr" rid="ref8">(Dressel et al., 2008)</xref>
        considering
sequence information only, to assess the influence of
mutations on protein structure.
      </p>
      <p>The approach considers the solvation energy,
which is based on the probability distribution for
each amino acid within the integral part of a
membrane protein to be facing the membrane or other
proteins. The amino acid specific property inside
or outside reflects the orientation of the amino acid
side chains with respect to the centre of mass of the
neighbouring residues. For a given mutation, the
approach compares the solvation energies for
wildtype and mutant residues. If the energies differ
significantly, a destabilising effect is predicted,
especially if the energies are changing from negative to
positive or vice versa.</p>
      <p>To quantify the ability of this model to
predict the influence of mutations on the stability of
membrane proteins, we compared already examined
and published effects of mutations with the
predictions of the sequence based model. For this
purpose, we screened the literature for single point
mutations reported for five membrane proteins from
the family of G protein-coupled receptors
(bacteriorhodopsin and halorhodopsin from Halobacterium
salinarum, bovine rhodopsin, Na+/H+ antiporter
from Escherichia coli, and human aquaporin-1). As
described in Section 4, Protein-centred approach
and Figure 1B, articles relevant for these proteins
were identified by searching PubMed via the NCBI
Entrez Programming Utilities. Abstracts for each
protein were queried by the protein and gene name
including the synonyms as derived from the
corresponding PDB/UniProt entry.</p>
      <p>The MutationTagger was applied on these five
sets of abstracts for the extraction of mutation
information. The application of sequence checks brought
the results down to a reasonable number of proposed
mutations, which were presented as HTML
documents and subsequently manually curated. We only
used the publications where a single point mutation
was discussed in the context of stability or
stability related function. Double or multiple mutations
were not considered, as the determination of a direct
relation between the reported effect and one of the
mutations is not possible. If an appropriate mutation
was found in the literature, we compared the
solvation energies of both wild-type and mutant residues
to decide, if the mutation was stabilising, slightly
stabilising, slightly destabilising, or destabilising.
Example Mutation T93P for bovine rhodopsin was
reported to lead to a conformational change of the
protein. Considering the two solvation energies of
wild type Threonine (-0.66 a.u.) and mutant Proline
(0.08 a.u.) a destabilising effect can be predicted,
although both amino acids are actually classified as
neutral. Without the change of sign from - to +, an
only slightly destabilising effect would have been
hypothesised.</p>
      <p>Relevance We were able to show the ability of our
mutation mining approach to retrieve publications
containing mutation information for given proteins
at a good precision. Due to the quick and precise
retrieval of mutation data we were able to assess the
soundness of the coarse grained model for the
prediction of stabilising regions in membrane proteins.
25 out of 35 mutational effects reported in the
literature for any of these five membrane proteins
correlate with the predictions based on the solvation
energy. These cases suggest a relation between
mutations and stability issues in membrane proteins.
Acknowledgement: We are grateful for financial
support by the EU project Sealife and the BMBF
Format Project CLSD and to Frank Dressel and Dirk
Labudde for discussions on the application.
6</p>
    </sec>
    <sec id="sec-14">
      <title>Conclusion</title>
      <p>We developed a rule- and regular expression-based
approach that allows for the retrieval of protein point
mutations from the whole PubMed database
specifically for any given protein. This flexibility makes
it a powerful tool for immediately finding relevant
data for follow-up studies, as we showed in the
application on five membrane proteins. In addition,
MutationTagger can be utilised for the species-wide
identification of mutations in proteins mentioned in
PubMed. We started to set up a mutation database
which allows for systematically querying mutation
related information, and finding relevant literature
for subsequent studies. The sequence checks applied
on identified mutations and candidate proteins have
been proven to be an efficient, yet not sufficient
filter for determing mutation-protein associations. The
filter shows good sensitivity but improvable
specificity, especially regarding the species level.
Furthermore, we were able to show, that the mutation
information from literature can even further improve
the quality of the gene tagging algorithm we used,
which already showed very good results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>B</given-names>
            <surname>Alberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Bray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K</given-names>
            <surname>Hopkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Raff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K</given-names>
            <surname>Roberts</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P</given-names>
            <surname>Walter</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Essential Cell Biology</article-title>
          . Garland Science Textbooks, London.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Ashburner</surname>
          </string-name>
          , Catherine Ball, Judith Blake, David Botstein,
          <string-name>
            <given-names>Heather</given-names>
            <surname>Butler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cherry</surname>
          </string-name>
          , Allan Davis, Kara Dolinski, Selina Dwight, Janan Eppig, Midori Harris, David Hill, Laurie Issel-Tarver,
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Kasarskis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Suzanna</given-names>
            <surname>Lewis</surname>
          </string-name>
          , John Matese, Joel Richardson, Martin Ringwald, Gerald Rubin, and
          <string-name>
            <given-names>Gavin</given-names>
            <surname>Sherlock</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Gene ontology: tool for the unification of biology. the gene ontology consortium</article-title>
          .
          <source>Nature genetics.</source>
          ,
          <volume>25</volume>
          :
          <fpage>25</fpage>
          -
          <lpage>29</lpage>
          , May.
          <volume>10</volume>
          .1038/75556.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>C</given-names>
            <surname>Broud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G</given-names>
            <surname>Collod-Broud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C</given-names>
            <surname>Boileau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Soussi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C</given-names>
            <surname>Junien</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Umd (universal mutation database): a generic software to build and analyze locus-specific databases</article-title>
          .
          <source>Hum Mutat</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <fpage>86</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>MN Cantor and YA Lussier</source>
          .
          <year>2004</year>
          .
          <article-title>Mining omim for insight into complex diseases</article-title>
          .
          <source>Medinfo</source>
          ,
          <volume>11</volume>
          (Pt 2):
          <fpage>753</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Gregory</surname>
          </string-name>
          <string-name>
            <given-names>Caporaso</given-names>
            , Jr William A.
            <surname>Baumgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>David A.</given-names>
            <surname>Randolph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bretonnel Cohen</surname>
          </string-name>
          , and Lawrence Hunter. 2007a.
          <article-title>Mutationfinder: A highperformance system for extracting point mutation mentions from text</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>23</volume>
          :
          <fpage>1862</fpage>
          -
          <lpage>1865</lpage>
          , Jul.
          <volume>10</volume>
          .1093/bioinformatics/btm235.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Gregory</surname>
          </string-name>
          <string-name>
            <given-names>Caporaso</given-names>
            , William A.
            <surname>Baumgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>David A.</given-names>
            <surname>Randolph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bretonnel Cohen</surname>
          </string-name>
          , and Lawrence Hunter. 2007b.
          <article-title>Rapid pattern development for concept recognition systems: application to point mutations</article-title>
          .
          <source>Journal of bioinformatics and computational biology</source>
          ,
          <volume>5</volume>
          :
          <fpage>1233</fpage>
          -
          <lpage>1259</lpage>
          , Dec.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Doms</surname>
          </string-name>
          and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Schroeder</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Gopubmed: exploring pubmed with the gene ontology</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <volume>33</volume>
          :
          <fpage>W783</fpage>
          -6, Jul.
          <volume>10</volume>
          .1093/nar/gki470.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>F</given-names>
            <surname>Dressel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Marsico</surname>
          </string-name>
          ,
          <string-name>
            <surname>A Tuukkanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R</given-names>
            <surname>Winnenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Labudde</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M</given-names>
            <surname>Schroeder</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Stabilizing regions in membrane proteins</article-title>
          .
          <source>In From Computational Biophysics to Systems Biology (CBSB08)</source>
          , pages
          <fpage>197</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>M</given-names>
            <surname>Erdogmus and OU Sezerman</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Application of automatic mutation-gene pair extraction to diseases</article-title>
          .
          <source>J Bioinform Comput Biol</source>
          ,
          <volume>5</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1261</fpage>
          -
          <lpage>75</lpage>
          , Dec.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>D</given-names>
            <surname>Fredman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G</given-names>
            <surname>Munns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Rios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F</given-names>
            <surname>Sjholm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Siegfried</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B</given-names>
            <surname>Lenhard</surname>
          </string-name>
          , H Lehvslaiho, and
          <string-name>
            <given-names>AJ</given-names>
            <surname>Brookes</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Hgvbase: a curated resource describing human dna variation and phenotype relationships</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <volume>32</volume>
          (Database issue):
          <fpage>D516</fpage>
          -9, Jan.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <article-title>Jo¨rg Hakenberg, Loic Royer</article-title>
          , Conrad Plake, Hendrik Strobelt, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Schroeder</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Me and my friends: gene mention normalization with background knowledge</article-title>
          .
          <source>In Proceedings of the Second BioCreative Challenge Evaluation Workshop</source>
          , pages
          <fpage>141</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>J</given-names>
            <surname>Hakenberg</surname>
          </string-name>
          ,
          <string-name>
            <surname>C Plake</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R</given-names>
            <surname>Leaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Schroeder</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G</given-names>
            <surname>Gonzales</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Inter-species normalization of gene mentions with GNAT</article-title>
          . Bioinformatics. to appear.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>A</given-names>
            <surname>Hamosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>AF</given-names>
            <surname>Scott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Amberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C</given-names>
            <surname>Bocchini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Valle</surname>
          </string-name>
          , and
          <source>VA McKusick</source>
          .
          <year>2002</year>
          .
          <article-title>Online mendelian inheritance in man (omim), a knowledgebase of human genes and genetic disorders</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <volume>30</volume>
          (
          <issue>1</issue>
          ):
          <fpage>52</fpage>
          -
          <lpage>5</lpage>
          , Jan.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>O</given-names>
            <surname>Horaitis and RG Cotton</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>The challenge of documenting mutation across the genome: the human genome variation society approach</article-title>
          .
          <source>Hum Mutat</source>
          ,
          <volume>23</volume>
          (
          <issue>5</issue>
          ):
          <fpage>447</fpage>
          -
          <lpage>52</lpage>
          , May.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>R</given-names>
            <surname>Kanagasabai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>KH</given-names>
            <surname>Choo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Ranganathan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>CJ</given-names>
            <surname>Baker</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>A workflow for mutation extraction and structure annotation</article-title>
          .
          <source>J Bioinform Comput Biol</source>
          ,
          <volume>5</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1319</fpage>
          -
          <lpage>37</lpage>
          , Dec.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>J</given-names>
            <surname>Kazius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K Wurdinger</given-names>
            ,
            <surname>Iterson</surname>
          </string-name>
          <string-name>
            <surname>M van</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Kok</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Bck</surname>
          </string-name>
          , and
          <string-name>
            <given-names>AP</given-names>
            <surname>Ijzerman</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Gpcr nava database: natural variants in human g protein-coupled receptors</article-title>
          .
          <source>Hum Mutat</source>
          , Oct.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>R</given-names>
            <surname>Klinger</surname>
          </string-name>
          ,
          <article-title>CM Friedrich, HT Mevissen</article-title>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Fluck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Hofmann-Apitius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>LI</given-names>
            <surname>Furlong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F</given-names>
            <surname>Sanz</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Identifying gene-specific variations in biomedical text</article-title>
          .
          <source>J Bioinform Comput Biol</source>
          ,
          <volume>5</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1277</fpage>
          -
          <lpage>96</lpage>
          , Dec.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>M</given-names>
            <surname>Kroos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>RJ</given-names>
            <surname>Pomponio</surname>
          </string-name>
          , Vliet L van, RE Palmer, M Phipps,
          <string-name>
            <surname>der Helm R Van</surname>
            ,
            <given-names>D</given-names>
          </string-name>
          <string-name>
            <surname>Halley</surname>
          </string-name>
          ,
          <article-title>and A Reuser and</article-title>
          .
          <year>2008</year>
          .
          <article-title>Update of the pompe disease mutation database with 107 sequence variants and a format for severity rating</article-title>
          .
          <source>Hum Mutat</source>
          , Apr.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>K</given-names>
            <surname>Lage</surname>
          </string-name>
          , EO Karlberg, ZM Strling, PI Olason, AG Pedersen,
          <string-name>
            <given-names>O</given-names>
            <surname>Rigina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>AM</given-names>
            <surname>Hinsby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z</given-names>
            <surname>Tmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F</given-names>
            <surname>Pociot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N</given-names>
            <surname>Tommerup</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y</given-names>
            <surname>Moreau</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S</given-names>
            <surname>Brunak</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>A human phenome-interactome network of protein complexes implicated in genetic disorders</article-title>
          .
          <source>Nat Biotechnol</source>
          ,
          <volume>25</volume>
          (
          <issue>3</issue>
          ):
          <fpage>309</fpage>
          -
          <lpage>16</lpage>
          , Mar.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Lawrence C.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Florence</given-names>
            <surname>Horn</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Fred E.</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Automatic extraction of protein point mutations using a graph bigram association</article-title>
          .
          <source>PLoS computational biology</source>
          ,
          <volume>3</volume>
          :e16,
          <string-name>
            <surname>Feb</surname>
          </string-name>
          .
          <volume>10</volume>
          .1371/journal.pcbi.
          <volume>0030016</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>AC</given-names>
            <surname>Martin</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Mapping pdb chains to uniprotkb entries</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>21</volume>
          (
          <issue>23</issue>
          ):
          <fpage>4297</fpage>
          -
          <lpage>301</lpage>
          , Dec.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>F</given-names>
            <surname>Milhavet</surname>
          </string-name>
          ,
          <string-name>
            <surname>L Cuisset</surname>
          </string-name>
          , HM Hoffman,
          <string-name>
            <given-names>R</given-names>
            <surname>Slim</surname>
          </string-name>
          , H ElShanti,
          <string-name>
            <given-names>I</given-names>
            <surname>Aksentijevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Lesage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H</given-names>
            <surname>Waterham</surname>
          </string-name>
          , C Wise,
          <string-name>
            <surname>de Menthiere C Sarrauste</surname>
            ,
            <given-names>and I</given-names>
          </string-name>
          <string-name>
            <surname>Touitou</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>The infevers autoinflammatory mutation online registry: update with new genes and functions</article-title>
          .
          <source>Hum Mutat</source>
          , Apr.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Ogata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Goto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Fujibuchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bono</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Kanehisa</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Kegg: Kyoto encyclopedia of genes and genomes</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <volume>27</volume>
          :
          <fpage>29</fpage>
          -
          <lpage>34</lpage>
          , Jan.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>D</given-names>
            <surname>Rebholz-Schuhmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Marcel</surname>
          </string-name>
          ,
          <string-name>
            <surname>S Albert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R</given-names>
            <surname>Tolle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G</given-names>
            <surname>Casari</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H</given-names>
            <surname>Kirsch</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Automatic extraction of mutations from medline and cross-validation with omim</article-title>
          .
          <source>Nucleic Acids Res</source>
          ,
          <volume>32</volume>
          (
          <issue>1</issue>
          ):
          <fpage>135</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>A</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <surname>A Olowoyeye</surname>
          </string-name>
          ,
          <source>PH Baenziger, J Dantzer</source>
          , MG Kann,
          <string-name>
            <surname>P Radivojac</surname>
          </string-name>
          , R Heiland,
          <source>and SD Mooney</source>
          .
          <year>2007</year>
          .
          <article-title>Mutdb: update on development of tools for the biochemical analysis of genetic variation</article-title>
          .
          <source>Nucleic Acids Res</source>
          , Sep.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>PD</given-names>
            <surname>Stenson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E</given-names>
            <surname>Ball</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K</given-names>
            <surname>Howells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Phillips</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Mort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and DN</given-names>
            <surname>Cooper</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Human gene mutation database: towards a comprehensive central mutation database</article-title>
          .
          <source>J Med Genet</source>
          ,
          <volume>45</volume>
          (
          <issue>2</issue>
          ):
          <fpage>124</fpage>
          -
          <lpage>6</lpage>
          , Feb.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>R</given-names>
            <surname>Witte and CJ Baker</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Towards a systematic evaluation of protein mutation extraction systems</article-title>
          .
          <source>J Bioinform Comput Biol</source>
          ,
          <volume>5</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1339</fpage>
          -
          <lpage>59</lpage>
          , Dec.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <article-title>CL Worth, GR Bickerton, A Schreyer, JR Forman</article-title>
          , TM Cheng,
          <string-name>
            <given-names>S</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>S Gong</surname>
          </string-name>
          , DF Burke, and
          <string-name>
            <given-names>TL</given-names>
            <surname>Blundell</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>A structural bioinformatics approach to the analysis of nonsynonymous single nucleotide polymorphisms (nssnps) and their relation to disease</article-title>
          .
          <source>J Bioinform Comput Biol</source>
          ,
          <volume>5</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1297</fpage>
          -
          <lpage>318</lpage>
          , Dec.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>YL</given-names>
            <surname>Yip</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N</given-names>
            <surname>Lachenal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V</given-names>
            <surname>Pillet</surname>
          </string-name>
          , and AL Veuthey.
          <year>2007</year>
          .
          <article-title>Retrieving mutation-specific information for human proteins in uniprot/swiss-prot knowledgebase</article-title>
          .
          <source>J Bioinform Comput Biol</source>
          ,
          <volume>5</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1215</fpage>
          -
          <lpage>31</lpage>
          , Dec.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>YL</given-names>
            <surname>Yip</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Famiglietti</surname>
          </string-name>
          ,
          <string-name>
            <surname>A Gos</surname>
          </string-name>
          , PD Duek, FP David,
          <string-name>
            <given-names>A</given-names>
            <surname>Gateau</surname>
          </string-name>
          , and
          <string-name>
            <given-names>A</given-names>
            <surname>Bairoch</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Annotating single amino acid polymorphisms in the uniprot/swiss-prot knowledgebase</article-title>
          .
          <source>Hum Mutat</source>
          , Jan.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>