<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using structural bioinformatics to investigate the impact of non synonymous SNPs and disease mutations: scope and limitations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Methods</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Switch Laboratory, VIB, Vrije Universiteit Brussel</institution>
          ,
          <addr-line>Pleinlaan 2, 1050 Brussels</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Background: Linking structural e ects of mutations to functional outcomes is a major issue in structural bioinformatics, and many tools and studies have shown that speci c structural properties such as stability and residue burial can be used to distinguish neutral variations and disease associated mutations. Results: We have investigated 39 structural properties on a set of SNPs and disease mutations from the Uniprot Knowledge Base that could be mapped on high quality crystal structures and show that none of these properties can be used as a sole classi cation criterion to separate the two data sets. Furthermore, we have reviewed the annotation process from mutation to result and identi ed the liabilities in each step. Conclusions: Although excellent annotation results of various research groups underline the great potential of using structural bioinformatics to investigate the mechanisms underlying disease, the interpretation of such annotations cannot always be extrapolated to proteome wide variation studies. Di culties for large-scale studies can be found both on the technical level, i.e. the scarcity of data and the incompleteness of the structural tool suites, and on the conceptual level, i.e. the correct interpretation of the results in a cellular context.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Corresponding author</title>
      <sec id="sec-1-1">
        <title>Background</title>
        <p>
          The molecular phenotype of a coding non
synonymous SNP or disease associated mutation describes
the functional and structural properties of a protein
that are a ected by a single amino acid
substitution [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. In this study we want to address whether
the concept of the in silico determined molecular
phenotype can be employed for large-scale classi
cation of SNPs and disease mutations. The attempt to
classify a large set of mutations based on an
incomplete molecular phenotype may seem naive at rst
glance, had it not been suggested that individual
properties such as protein stability, the accessibility
of the amino acid substitution site, and the location
of variants in surface pockets are predictive
determinants of the phenotypic e ect of a variation [1{4].
A comparative study of protein stability predictors
by Blundell and co-workers demonstrated that
although protein stability changes caused by mutation
can be relatively accurately estimated in silico, these
predictions by themselves do not yield accuracy on
large-scale classi cation between benign and
disruptive mutations [5{7].
        </p>
        <p>Furthermore, computational analyses rely
heavily on the quality of the data under scrutiny and
the computational methods used to evaluate these
data. Before investigating 39 structural properties of
proteins and amino acid substitutions for their
predictive power regarding SNP classi cation, we have
investigated what major liabilities are encountered
when implementing an structural approach to SNP
annotation and classi cation. The results are
compared with those achieved by the best performers
among the state-of-the-art tools.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Results and Discussion</title>
        <p>
          In this study we have identi ed the common issues
that are encountered when performing large-scale
analyses of structural properties of human coding
variation. The rst issue concerns the availability of
structural data for nsSNPs and disease mutations,
while the second involves the availability of
computational tools to predict structural properties. The
last issue concerns the quality of classi cation: are
the training and evaluation data sets used in the
analyses su cient to extrapolate results for larger
studies, and do the properties used have su cient
predictive power to separate the two data sets?
Structural coverage of human genetic variation
Despite structural genomics projects, the gap
between sequence and structural information is still
wide, and the coverage of variation data with
structural data is estimated to be as low as 14% [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. We
have investigated the boundaries of structural
coverage by varying the quality requirements on the
structural model (Supplementary Figure S1A), the
sequence identity between query sequence and
modelled structure (Figure S1B), the percentage of the
wild type sequence covered by the structural model
(Figure S1C), and the length of the alignment
between query and target (Figure S1D). Without
applying any restrictions, about 12% of all nsSNPs
present in the Ensembl Variation Database (release
44) can be mapped on a structural model, in
accordance with the estimate cited previously. However,
this percentage is valid only when no restrictions
regarding sequence identity, sequence coverage or
structure quality are applied. Our standard
restrictions on building high-con dence structural models
using the FoldX force eld are X-ray structures with
a resolution lower than 2.5 A and sequence identity
higher than 80%. Applying these restrictions to the
Ensembl data results in a data set of 5416 nsSNPs
(circa 4% of the data, Figure S1B).
        </p>
        <p>
          Predictability of structural properties
The second issue for a large-scale structural
bioinformatics approach is the structural properties that are
predictable with state of the art tools: how well can
we describe the structural behaviour of a protein and
its mutants? Previous structural studies have
identi ed protein stability, aggregation and misfolding
as determinants of correct functioning on the single
protein level [
          <xref ref-type="bibr" rid="ref11 ref12 ref7">7,11,12</xref>
          ]. Mutations a ecting the
functional sites of a protein, such as DNA, ligand and
protein interaction sites, are not considered within
this scope, but the investigation of these sites will
most certainly be of great importance to assess the
impact of amino acid substitutions.
        </p>
        <p>
          Tools have been developed that describe the
structure and dynamics of a protein: stability,
aggregation, amyloidosis, and folding. We have used
computational methods that are capable of assessing
the e ects of a mutation on protein stability (FoldX),
aggregation (Tango) and amyloidosis (Waltz).
Although algorithms exist that can predict folding of
small single domain proteins (e.g. Rosetta [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ],
FoldX [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], SimFold [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]), to date no computational
method exists that can predict folding events on
large multi-domain proteins, or that is applicable in
genome wide studies.
        </p>
        <p>
          Although we have not investigated
proteinprotein interactions in this study, we have included
an analysis of the binding of proteins to molecular
chaperones, as it is directly related to correct folding
of the protein. The high abundance of chaperones in
the cell emphasises their crucial role in the cell [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ],
but this is not re ected in the availability of
computational tools for chaperone binding. We have used
the only available tool, the Hsp70 binding
predictor Limbo [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], to assess chaperone binding variation
caused by amino acid alteration.
        </p>
        <p>
          The predictive power of structural properties
Following the recommendations of Care et al [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ],
we have used the SwissProt annotated disease and
polymorphism data (SwissProt Variation Index
release 52) as the evaluation data for our analyses.
Mapping of these variants on high quality structural
models (X-ray structures with resolution 2:5A,
sequence identity with the model above 80%) yielded
a data set of 240 positive (disease-associated)
mutations and 400 negative variations (neutral nsSNPs)
in 98 proteins. To ensure that the analyses are
comparable, we applied the sequence based predictors to
the same small data set as the predictors that use
3D structures or structural models.
        </p>
        <p>
          Before we evaluated the discriminative power of
the individual structural parameters, we wanted to
assess whether our data showed distinguishable
patterns for three important parameters. The rst two
criteria, stability di erence and the degree of burial
of the mutation site, have previously been
identied as providing information about the severity of
a mutation [
          <xref ref-type="bibr" rid="ref19 ref4">4, 19</xref>
          ]. The third criterion is di erence
in aggregation propensity, which has been cited as
likely to be an important factor in disease
susceptibility [
          <xref ref-type="bibr" rid="ref12 ref20">12, 20</xref>
          ] but thus far has not been applied in a
proteome wide mutation analysis.
        </p>
        <p>
          Figure 1 shows the distributions for the
stability di erences (A) and di erences in aggregation
propensity (B) between wild type and variant
proteins, and the burial of the mutation site (C). The
rst observation of both the stability and the
aggregation analysis is that the observed changes are
not discrete but follow a smooth distribution from
negative to positive change. Second, there are
noticeable di erences between SNPs and disease
mutations, but they cannot be distinguished by a simple
cut-o value on the output, as there is large
overlap between the distributions. This is con rmed by
the P-values obtained from paired student t-tests,
which are 0.96 for the stability distributions, 0.99
for the aggregation distributions, and 0.99 for the
burial distributions, respectively. For the stability
distributions, we see that disease mutations are
generally more destabilising than SNPs, but their
distributions overlap largely. A similar analysis has been
performed on SwissProt variants using the Site
Directed Mutator stability predictor [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], and the
distributions of stability di erences of disease mutations
and neutral variations are similar to our ndings.
        </p>
        <p>In a rst series of properties to test as classi ers,
we have investigated 15 properties of the amino acid
substitution site that contribute to the assessment of
the e ect of the mutation using the FoldX algorithm
(Table 3). Cut o values were generated that
varied between the minimal and maximal values
measure for the speci c property, and the true and false
positive rate, and the Matthews correlation coe
cient (MCC) were calculated for each cut-o value.
Table 3 lists the data for both the best MCC and
the MCC90, i.e. the coe cient that is measured at
high speci city (true negative rate = 90%). The
corresponding ROC curves for these analyses can be
found in Supplementary Figure S1.</p>
        <p>The same strategy was then applied to predicted
values of structural di erences between mutant and
wild type proteins (24 properties). Statistics were
calculated for stability and entropy parameters, as
well as for di erences concerning protein
aggregation, amyloidosis and chaperone binding (Table 4,
Supplementary Figure S2).</p>
        <p>
          The results obtained from these detailed
analyses are unanimous: none of the parameters evaluated
can be used to separate the data. All MCC values
are close to zero, and thus the predictions are no
better than a random predictor would perform on the
data. The high accuracy of FoldX for stability
estimation has been proven in various studies [
          <xref ref-type="bibr" rid="ref10 ref6 ref9">6,9,10</xref>
          ], so
we have high con dence in our stability estimations.
In accordance with the analyses of [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], we nd that
high stability di erences alone are no su cient
criterion to distinguish deleterious mutations and neutral
variation. These results show that the dominant
effect of for instance stability that was proposed in
earlier large-scale studies [
          <xref ref-type="bibr" rid="ref22 ref4">4, 22</xref>
          ] can not be always
generalised for other data.
        </p>
        <p>The fact that none of the properties representing
conformational di erences between wild type and
variant protein contain enough information to
distinguish neutral and deleterious variation implies that
large-scale classi cation based on singular structural
properties is not feasible and requires a better
understanding of how the complex interplay between
biophysical and biochemical properties of a protein
conspire to di erent tolerance for mutations in
different proteins.</p>
        <p>Recent studies that combine structural and
evolutionary information using machine learning
techniques are able to classify relatively large data sets
obtained for the SwissProt database successfully
(summarised in Table S2). Machine learning
approaches suggest that data integration is indeed the
way forward, but the creation of this black box style
of classi er does not o er insight into the biological
processes. In the same way that using evolutionary
information to classify SNPs obscures the how and
why a speci c mutation is deleterious, using black
box machine learning methods will not teach us what
the underlying reason of disease is. Although
knowing that an amino acid is critical for correct function
is of course useful, in a structural bioinformatics
approach the focus is more on the molecular
mechanism underlying disease.</p>
        <p>
          A simple combination of the SNPe ect structural
bioinformatics toolsuite on our evaluation data set
showed that in our case, at least a linear
combination of these methods is not su cient to classify the
data (TPR = 0.73, TNR= 0.27, MCC=0). A large
part of the polymorphism data is predicted to have
deleterious e ect. To assess the \predictiveness" of
our data set, we applied the well-established
evolutionary method SIFT [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] to our data and found that
SIFT was also not able to classify e ectively. In fact
the results were even worse than our naive classi er
(TPR=0.69, TNR=0.21, MCC=-0.12).
        </p>
        <p>
          As an illustration of the in uence of the data
set used for evaluation on the performance of a
predictor, we list the results for the variation in
performance of SNP classi cation of SIFT, that uses
evolutionary information to label SNPs
(Supplementary Table S3). The Matthews correlation coe cient
varies between -0.12 on our data set over 0.25 on
human mutagenesis data, up to 0.59 on the HIV-1
protease mutagenesis set in the original SIFT paper [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].
This is yet another informative example on how
crucial the choice of training and test data are to build
and evaluate predictors: generalisation of results is
only possible when the training data are expressive
enough to represent the entire feature space.
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>Conclusions</title>
        <p>
          The concept of using the molecular phenotypic
effect of a nsSNP to assess its e ect on the structure
and function of the protein it alters was rst
introduced by Bork and co-workers [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. The question
has been raised to how much of this molecular
phenotype is necessary to evaluate the contribution of
a SNP to a disease phenotype: are there singular
dominant properties that determine the impairment
of structure and function, or do we need to consider
the full ensemble of molecular properties to interpret
the impact of the SNP? Other research groups have
proposed that single properties such as stability [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
and solvent accessibility [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] can be used to classify
SNPs. We have examined all the individual
structural bioinformatics tools that were proposed in the
SNPe ect toolsuite [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] for their ability to act as a
binary classi er for deleterious and neutral SNPs.
Neither of the individual properties that were
examined could serve this purpose. Because several
approaches were able to classify similar data sets as
the one we have used, we applied the most used
evolutionary method, SIFT [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], to our data set. As it
was not able to classify our data set accurately, we
argued that generalisation of the results presented by
the state of the art classi ers might be an important
issue. We illustrated this problem with the
variability of performance of SIFT on 8 di erent data sets
used in various analyses.
        </p>
        <p>From these analyses we concluded that strict
classi cation of SNPs is not feasible at the time, both
because there are still many technical di culties to
overcome, and because the biological interpretation
of the molecular phenotype in relation to a disease
phenotype is a complex matter. Even at the single
molecule level, we cannot assess how tolerant a
speci c protein is to structural variation. The inherent
rigidity of a protein might in uence the change in
stability that is allowed before severe conformational
changes are introduced. Furthermore, on the
cellular level biological interpretation is even harder: we
can not predict the role of the protein quality control
system plays in this tolerance level, not all
interactions are described at the molecular level, and much
more. Even if we can predict the molecular e ect
accurately, this might not necessarily result in a
disease phenotype because of functional redundancy of
the protein.</p>
        <p>However, not being able to classify human
variation into disease mutations and neutral or bene cial
variation does not mean that this approach or the
methods developed are useless. By using high
quality bioinformatics tools, we can select from a large
pool of variations the candidates that are interesting
for detailed investigation. This in itself is a valuable
contribution, because the amount of variation data
available is too massive to be investigated
experimentally. In silico analyses can and will be used
successfully as an addition to in vitro and in vivo
studies.</p>
        <sec id="sec-1-3-1">
          <title>Assembly of data sets</title>
          <p>
            Statistics on the structural coverage and validation
status of human non synonymous coding SNPs were
performed on data from the Ensembl human
variation database release 44, containing 12.2 million
SNPs, of which 133698 cause an amino acid
variation in a known transcript. The mapping of SNPs
on protein structures was evaluated using the
\ensppdbmapping" DAS service provided by the SPICE
server [
            <xref ref-type="bibr" rid="ref27">27</xref>
            ]. Positive and negative data sets for
the evaluation of SNP classi cation were designed
with data from the SwissProt variation index [
            <xref ref-type="bibr" rid="ref28">28</xref>
            ] in
the UniProt knowledge base (version 52.0, March
2007, [
            <xref ref-type="bibr" rid="ref29">29</xref>
            ]) that were mapped onto known PDB
structures and high quality homologs thereof. The
quality criteria described in the results section
(models with resolution of 3 Aor higher, sequence
identity of 80% or more) lead to structural models of
400 SNPs (negative) and 240 disease associated
mutations (positive).
          </p>
        </sec>
        <sec id="sec-1-3-2">
          <title>Structural bioinformatics tools</title>
          <p>We have used the FoldX force eld [33] for all
mutant properties regarding structural location, protein
stability and its various components, the Tango [34]
and Waltz [35, submitted] algorithms to assess the
propensity for aggregation of wild type and variant
proteins, and the Limbo algorithm [17, submitted] to
evaluate the chaperone-binding properties of amino
acid sequences. A novel tool developed by Lenaerts
et al (unpublished) was used to estimate the
entropy of a speci c amino acid site in a high-resolution
structure. Detailed descriptions of these ve tools
can be found in the Supplementary Material.</p>
        </sec>
      </sec>
      <sec id="sec-1-4">
        <title>Authors contributions</title>
        <p>Conceived and designed the experiments: JR JS FR.
Performed the experiments: JR. Analysed the data:
JR JS FR. Wrote the paper: JR.</p>
      </sec>
      <sec id="sec-1-5">
        <title>Acknowledgements</title>
        <p>Joke Reumers was supported by a grant from the Federal
Research O ce (FWO, IUAP P6/43), Belgium, and the
Institute for the encouragement of Scienti c Research
and Innovation of Brussels (ISRIB), Belgium.
29. UniProt Consortium:</p>
        <p>Resource (UniProt).
35(Database issue):D193{7.</p>
        <p>The
Nucleic</p>
        <p>Universal Protein</p>
        <p>Acids Res 2007,
30. Zweig MH, Campbell G: Receiver-operating
characteristic (ROC) plots: a fundamental evaluation
tool in clinical medicine. Clin Chem 1993, 39(4):561{
577.
31. Matthews BW: Comparison of the predicted
and observed secondary structure of T4 phage
lysozyme. Biochim Biophys Acta 1975, 405(2):442{451.
Best MCC</p>
        <p>Threshold</p>
        <p>MCC90
1.61
-1.05
-1.76
-0.10
0.32
1.96
-0.98
-0.6
1.5
0.22
0.43
0.73
0.93
0.19
0.22
&lt;0
-0.01
0.05
0.10
0.15
0.16
0.06
0.15
-0.1
0.05
0</p>
      </sec>
      <sec id="sec-1-6">
        <title>Additional Files</title>
        <p>Figure 1 { gure1.pdf
Additional le 2 | supplementary.pdf
Several of the less critical gures and tables are added as supplementary material, together with detailed
descriptions of the structural bioinformatics tools used.
Using structural bioinformatics to investigate the impact of
non synonymous SNPs and disease mutations: scope and
limitations
Supplementary Material
Joke Reumers1, Joost Schymkowitz1 and Frederic Rousseau 1
1Switch Laboratory, VIB, Vrije Universiteit Brussel, Pleinlaan 2, 1050 Brussels, Belgium
Email: Joke Reumers - joke.reumers@vub.ac.be; Joost Schymkowitz - joost.schymkowitz@vub.ac.be; Frederic Rousseau
frederic.rousseau@vub.ac.be;</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Corresponding author</title>
      <p>Methods
Structural bioinformatics tools
FoldX
The FoldX force eld was developed for the fast and
accurate estimation of the free change upon
mutation on the stability of a protein or a protein
complex [23{26]. It uses an all-atom representation of
these macromolecules, and has been validated on a
test database of more than 1000 mutants from more
than 20 di erent proteins. It currently yields a
correlation of 0.78 with a standard deviation of 0.41
kcal/mol.</p>
      <p>Modelling and evaluation of mutations in FoldX
is performed with the BuildModel command. It is
used rst to model a homologous sequence on a
structural model and to optimise the side chains to</p>
      <p>t the new sequence, and then to evaluate the e ect
of a single amino acid variation. The Gibbs free
energy of a protein is calculated with the Stability
command. The various structural parameters used in
the classi cation tests (backbone clash, backbone H
bond formation , sidechain H bond formation,
electrostatics , solvation of hydrophobic residues,
solvation of polar residues, torsion clash, Van der Waals
contribution,Van der Waals clash)</p>
      <p>Entropy calculations based on side chain sampling
In addition to the entropy calculations intrinsic to
the FoldX force eld, we use a novel method based
on extensive sampling of side chain conformations
as developed by Lenaerts et al. (unpublished). The
sampling method produces for each side chain the
probability (P (X)) of nding the residue's side chain
in a particular conformational state. From these
probabilities entropy can easily be derived:</p>
      <p>H(X) =</p>
      <p>X iP (xi)log2P (xi)</p>
      <p>
        The method uses a rotamer database based on
conditional statistics of dihedral angles derived from
the WHAT IF data set [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. All amino acids from
this data and their corresponding dihedral angles
(10 bin) were used to derive the following
probabilities: P ( i); P ( ij i 1) and P ( ij i 1; i 2), except
for 1(P ( 1) and P ( 1j ; )). A set of n random
rotamers can be derived from the probability
distribution thus calculated. This will allow sampling of
rotamers with greater resolution than classical
rotamer libraries.
      </p>
      <p>The sampling itself is performed by Monte Carlo
based sampling method with Metropolis criterion (at
298K). The Metropolis criterion states that a certain
conformational change is accepted with a
probability p that depends on the free energy change G
associated with the conformational change as given
by the following formula:</p>
      <p>
        Tango
The -aggregation prediction algorithm Tango [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]
uses a statistical mechanics approach to represent
a competition between major conformational states:
the random coil and the native conformations, as
well as -turn, -helix and -aggregate. Two
windows of variable length slide over the sequence, and
each such window can populate these conformational
states according to a Boltzmann distribution. The
frequency of population of each structural state for
a given segment will be relative to its energy, which
is derived from statistical and empirical parameters.
      </p>
      <p>To predict the -aggregating segments of a peptide,
Tango calculates the partition function of the phase
space involving these conformational states. In our
analysis we have used Tango to calculate the di
erence in aggregation tendency that results from an
single amino acid variation.</p>
      <p>
        Waltz
Current methods for the prediction of the sequence
determinants of amyloidosis su er from two major
problems: overpredicting amorphous cross
aggregates and missing amylogenic sequences that are
enriched in the polar Q and N residues, such as the
prion protein. The Waltz algorithm [29,
submitted] tackles these problems by taking into account
amyloid hexapeptides from 48 new amyloid
forming sequences, derived from 31 proteins. About half
the proteins in this extended data set were not
previously known to contain amyloidogenic sequences
such as presenilin-2, titin and myosin. Waltz
combines terms from amino acid sequence scoring in the
learning set, physical property analysis and
homology modelling. The method shows 84% sensitivity
at 92% speci city on the AmylHex data set [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ],
and correctly identi es mutations in human proteins
known to be associated with amyloid deposition.
      </p>
      <p>
        Limbo
Limbo is a Hsp70 binding site predictor that was
built using a dual method combining sequence and
structural information [31, submitted].
Experimental DnaK binding data of 53 non-redundant
peptide sequences was used to generate a
sequencebased position-speci c scoring matrix (PSSM) based
on logarithm of the odds scores. Following an
in silico alanine scan of the substrate peptide in
the crystal structure of a DnaK-substrate complex
(PDBID 1DKX Zhu1996) using FoldX, a
structurebased PSSM that re ects the individual contribution
of certain substrate residue types for DnaK
binding was generated. The Limbo DnaK binding site
predictor was obtained by combining the
structurebased PSSM with a normalisation factor of 0.2 with
the sequence-based PSSM. Limbo is able to correctly
predict 89% of the true positives in a tested peptide
set (high sensitivity), with a concurrent amount of
only 5.9% false positives for a speci c score threshold
(high speci city). The robustness of the predictor
was evaluated with a cross-validation test, resulting
in a true positive rate of 72% true positives and a
false positive rate of 5.9%. The predictor was able
to identify an entire known DnaK binding site in the
heat-shock promoter 32 [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. We have used Limbo
to rank mutated proteins according to their DnaK
binding a nity.
Supplementary Table S1 - Types of data sets used to train and test SNP classi ers.
      </p>
      <p>
        Origin data set
Neutral variations
Mutagenesis studies
Orthologs
SwissProt SNP
OMIM
dbSNP
Disease mutations
Mutagenesis studies
COSMIC database
HGMD
OMIM
SwissProt Disease
Data [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
Data [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
      </p>
      <p>Size
of data set</p>
      <p>
        Number
of studies
[1{9]
[
        <xref ref-type="bibr" rid="ref10 ref3 ref9">3, 9, 10</xref>
        ]
[3, 8, 11{14]
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
[
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ]
[1{9]
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
[
        <xref ref-type="bibr" rid="ref13 ref15 ref18 ref3 ref8">3, 8, 13, 15, 18</xref>
        ]
[3, 8, 10{14, 19, 20]
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
[
        <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
        ]
Supplementary Table S2 - Performance of state-of-the-art predictors on representative data sets.
The performance of a few selected tools on SwissProt disease associated mutations and SNP data are shown.
      </p>
      <p>
        Study
Bao et al [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
Capriotti et al [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
Karchin et al [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
Ng &amp; Heniko [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]
Wang &amp; Moult [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
Worth et al [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
Yue &amp; Moult [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
      </p>
      <p>Method</p>
      <p>FPR</p>
      <p>FNR</p>
      <p>TPR</p>
      <p>TNR</p>
      <p>MCC
Random Forest
HybridMeth
SVM
SIFT
Stability
Combined
SVM
Supplementary Table S3 - Variation of the performance of SIFT on di erent data sets.</p>
      <p>Dataset
Test set
Human
lac I repressor
HIV 1-protease
T4 lysozyme
SwissProt disease
SwissProt + dbSNP
SwissProt</p>
      <p>A</p>
      <p>2 104
C
0
0</p>
      <p>2 104
1.5 104
s
P
N
S
ro 1 104
f
e
b
m
u
N</p>
      <p>5000
B
0</p>
      <p>0
100 100
80 80
TPR4600 TPR4600
20 20
0 0 20 40 FPR 60 80 100 0 0 20 40 FPR 60 80 100
Figure S2. ROC curves for classi cation of disease mutations and neutral variation by using structural properties of the
amino acid substitution site.
100
80
R60
TP40
20
0 0</p>
      <p>Solvation hydrophobic
20</p>
      <p>40 FPR 60
Main chain burial
80
100
20</p>
      <p>40 FPR 60
Side chain burial
80
100
20</p>
      <p>40 FPR 60
Solvation polar
80
100
20</p>
      <p>40 FPR 60
Van der Waals clash
80
100
100 100
80 80
TPR4600 TPR4600
20 20
0 0 20 40 FPR 60 80 100 0 0 20 40 FPR 60 80 100
Figure S2 (continued). ROC curves for classi cation of disease mutations and neutral variation by using structural
properties of the amino acid substitution site.
100
80
60
40
20
0
0
20
80
100
0
20
40
60
80</p>
      <p>100</p>
      <p>FPR
40</p>
      <p>60</p>
      <p>FPR
properties of the amino acid substitution site.
Overall stability difference</p>
      <p>Overall stability difference</p>
      <p>(surface residues)</p>
      <p>40 FPR 60
Backbone H bond
80
100
20</p>
      <p>40 FPR 60
Sidechain H bond
80
100
100 100
80 80
TPR4600 TPR4600
20 20
0 0 20 40 FPR 60 80 100 0 0 20 40 FPR 60 80 100
Figure S3. ROC curves for classi cation of disease mutations and neutral variation by using structural di erences between
the wild type and variant protein.
Electrostatics</p>
      <p>Entropy main chain</p>
      <p>40 FPR 60
Entropy side chain
80
100
20 40 FPR 60 80
Solvation hydrophobic</p>
      <p>100
80
100
20
80</p>
      <p>100
40 FPR 60</p>
      <p>Torsion
20
100 100
80 80
TPR4600 TPR4600
20 20
0 0 20 40 FPR 60 80 100 0 0 20 40 FPR 60 80 100
Figure S3 (continued). ROC curves for classi cation of disease mutations and neutral variation by using structural
di erences between the wild type and variant protein.</p>
      <p>Van der Waals clash
0
20
40
60
80
100
0
20
40
60
80
100</p>
      <p>FPR
100
80
60
40
20
0
80
60
40
20
0
80
60
40
20
0
100
R
P
T
0
20
40
60
80</p>
      <p>100</p>
      <p>FPR
60
40
20
0
80
60
40
20
0
100
100
R
P
T
80
60
40
20
0</p>
      <p>FPR
di erences between the wild type and variant protein.
40 FPR 60</p>
      <p>Waltz
80
100
20 40 FPR 60 80
Waltz (positive scores)</p>
      <p>100
100
20
80</p>
      <p>100
40 FPR 60</p>
      <p>Limbo
20 40 FPR 60 80</p>
      <p>Waltz (negative scores)
100 100
80 80
TPR4600 TPR4600
20 20
0 0 20 40 FPR 60 80 100 0 0 20 40 FPR 60 80 100
Figure S3 (continued). ROC curves for classi cation of disease mutations and neutral variation by using structural
di erences between the wild type and variant protein.
100
80
R60
TP40
20
0 0</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Chasman</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adams</surname>
            <given-names>RM</given-names>
          </string-name>
          :
          <article-title>Predicting the functional consequences of non-synonymous single nucleotide polymorphisms: Structure-based assessment of amino acid variation</article-title>
          .
          <source>J Mol Biol</source>
          <year>2001</year>
          ,
          <volume>307</volume>
          (
          <issue>2</issue>
          ):
          <volume>683</volume>
          {
          <fpage>706</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cli ord</surname>
            <given-names>RJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edmonson</surname>
            <given-names>MN</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buetow</surname>
            <given-names>KH</given-names>
          </string-name>
          :
          <article-title>Large-scale analysis of non-synonymous coding region single nucleotide polymorphisms</article-title>
          .
          <source>Bioinformatics</source>
          <year>2004</year>
          ,
          <volume>20</volume>
          (
          <issue>7</issue>
          ):
          <volume>1006</volume>
          {
          <fpage>1014</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ferrer-Costa</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orozco</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de la Cruz</surname>
            <given-names>X</given-names>
          </string-name>
          :
          <article-title>Sequencebased prediction of pathological mutations</article-title>
          .
          <source>Proteins</source>
          <year>2004</year>
          ,
          <volume>57</volume>
          (
          <issue>4</issue>
          ):
          <volume>811</volume>
          {
          <fpage>819</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jiang</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>T</given-names>
          </string-name>
          :
          <article-title>Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2006</year>
          , 7:
          <fpage>417</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Krishnan</surname>
            <given-names>VG</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Westhead</surname>
            <given-names>DR</given-names>
          </string-name>
          :
          <article-title>A comparative study of machine-learning methods to predict the e ects of single nucleotide polymorphisms on protein function</article-title>
          .
          <source>Bioinformatics</source>
          <year>2003</year>
          ,
          <volume>19</volume>
          (
          <issue>17</issue>
          ):
          <volume>2199</volume>
          {
          <fpage>2209</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Needham</surname>
            <given-names>CJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bradford</surname>
            <given-names>JR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bulpitt</surname>
            <given-names>AJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Care</surname>
            <given-names>MA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Westhead</surname>
            <given-names>DR</given-names>
          </string-name>
          :
          <article-title>Predicting the e ect of missense mutations on protein function: analysis with Bayesian networks</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2006</year>
          , 7:
          <fpage>405</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ng</surname>
            <given-names>PC</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heniko</surname>
            <given-names>S</given-names>
          </string-name>
          :
          <article-title>Predicting deleterious amino acid substitutions</article-title>
          .
          <source>Genome Res</source>
          <year>2001</year>
          ,
          <volume>11</volume>
          (
          <issue>5</issue>
          ):
          <volume>863</volume>
          {
          <fpage>874</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Saunders</surname>
            <given-names>CT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baker</surname>
            <given-names>D</given-names>
          </string-name>
          :
          <article-title>Evaluation of structural and evolutionary contributions to deleterious mutation prediction</article-title>
          .
          <source>J Mol Biol</source>
          <year>2002</year>
          ,
          <volume>322</volume>
          (
          <issue>4</issue>
          ):
          <volume>891</volume>
          {
          <fpage>901</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Yue</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moult</surname>
            <given-names>J</given-names>
          </string-name>
          :
          <article-title>Loss of protein structure stability as a major causative factor in monogenic disease</article-title>
          .
          <source>J Mol Biol</source>
          <year>2005</year>
          ,
          <volume>353</volume>
          (
          <issue>2</issue>
          ):
          <volume>459</volume>
          {
          <fpage>473</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ferrer-Costa</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orozco</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de la Cruz</surname>
            <given-names>X</given-names>
          </string-name>
          :
          <article-title>Characterization of disease-associated single amino acid polymorphisms in terms of sequence and structure properties</article-title>
          .
          <source>J Mol Biol</source>
          <year>2002</year>
          ,
          <volume>315</volume>
          (
          <issue>4</issue>
          ):
          <volume>771</volume>
          {
          <fpage>786</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bao</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cui</surname>
            <given-names>Y</given-names>
          </string-name>
          :
          <article-title>Prediction of the phenotypic effects of non-synonymous single nucleotide polymorphisms using structural and evolutionary information</article-title>
          .
          <source>Bioinformatics</source>
          <year>2005</year>
          ,
          <volume>21</volume>
          (
          <issue>10</issue>
          ):
          <volume>2185</volume>
          {
          <fpage>2190</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Bao</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cui</surname>
            <given-names>Y</given-names>
          </string-name>
          :
          <article-title>Functional impacts of nonsynonymous single nucleotide polymorphisms: Selective constraint and structural environments</article-title>
          .
          <source>FEBS Lett</source>
          <year>2006</year>
          ,
          <volume>580</volume>
          (
          <issue>5</issue>
          ):
          <volume>1231</volume>
          {
          <fpage>4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Capriotti</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calabrese</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Casadio</surname>
            <given-names>R</given-names>
          </string-name>
          :
          <article-title>Predicting the insurgence of human genetic diseases associated to single point protein mutations with support vector machines and evolutionary information</article-title>
          .
          <source>Bioinformatics</source>
          <year>2006</year>
          ,
          <volume>22</volume>
          (
          <issue>22</issue>
          ):
          <volume>2729</volume>
          {
          <fpage>2734</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Karchin</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diekhans</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            <given-names>DJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pieper</surname>
            <given-names>U</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eswar</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haussler</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sali</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>LS-SNP: largescale annotation of coding non-synonymous SNPs based on multiple information sources</article-title>
          .
          <source>Bioinformatics</source>
          <year>2005</year>
          ,
          <volume>21</volume>
          (
          <issue>12</issue>
          ):
          <volume>2814</volume>
          {
          <fpage>2820</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Stitziel</surname>
            <given-names>NO</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tseng</surname>
            <given-names>YY</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pervouchine</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goddeau</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasif</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            <given-names>J</given-names>
          </string-name>
          :
          <article-title>Structural location of disease-associated single-nucleotide polymorphisms</article-title>
          .
          <source>J Mol Biol</source>
          <year>2003</year>
          ,
          <volume>327</volume>
          (
          <issue>5</issue>
          ):
          <volume>1021</volume>
          {
          <fpage>1030</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Worth</surname>
            <given-names>CL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bickerton</surname>
            <given-names>GRJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreyer</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forman</surname>
            <given-names>JR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cheng</surname>
            <given-names>TMK</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gong</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burke</surname>
            <given-names>DF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blundell</surname>
            <given-names>TL</given-names>
          </string-name>
          :
          <article-title>A structural bioinformatics approach to the analysis of nonsynonymous single nucleotide polymorphisms (nsSNPs) and their relation to disease</article-title>
          .
          <source>J Bioinform Comput Biol</source>
          <year>2007</year>
          ,
          <volume>5</volume>
          (
          <issue>6</issue>
          ):
          <volume>1297</volume>
          {
          <fpage>1318</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Burke</surname>
            <given-names>DF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Worth</surname>
            <given-names>CL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Priego</surname>
            <given-names>EM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cheng</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smink</surname>
            <given-names>LJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Todd</surname>
            <given-names>JA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blundell</surname>
            <given-names>TL</given-names>
          </string-name>
          :
          <article-title>Genome bioinformatic analysis of nonsynonymous SNPs</article-title>
          .
          <source>BMC Bioinformatics</source>
          <year>2007</year>
          , 8:
          <fpage>301</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Worth</surname>
            <given-names>CL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burke</surname>
            <given-names>DF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blundell</surname>
            <given-names>TL</given-names>
          </string-name>
          :
          <article-title>Estimating the effects of single nucleotide polymorphisms on protein structure: how good are we at identifying likely disease associated mutations</article-title>
          ?
          <source>In Proceedings of Molecular Interactions - Bringing Chemistry to Life</source>
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Ng</surname>
            <given-names>PC</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heniko</surname>
            <given-names>S</given-names>
          </string-name>
          :
          <article-title>Accounting for human polymorphisms predicted to a ect protein function</article-title>
          .
          <source>Genome Res</source>
          <year>2002</year>
          ,
          <volume>12</volume>
          (
          <issue>3</issue>
          ):
          <volume>436</volume>
          {
          <fpage>446</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Wang</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moult</surname>
            <given-names>J</given-names>
          </string-name>
          :
          <article-title>SNPs, protein structure, and disease</article-title>
          .
          <source>Hum Mutat</source>
          <year>2001</year>
          ,
          <volume>17</volume>
          (
          <issue>4</issue>
          ):
          <volume>263</volume>
          {
          <fpage>270</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Halushka</surname>
            <given-names>MK</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fan</surname>
            <given-names>JB</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bentley</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsie</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weder</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cooper</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipshutz</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chakravarti</surname>
            <given-names>A</given-names>
          </string-name>
          :
          <article-title>Patterns of single-nucleotide polymorphisms in candidate genes for blood-pressure homeostasis</article-title>
          .
          <source>Nat Genet</source>
          <year>1999</year>
          ,
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <volume>239</volume>
          {
          <fpage>247</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Cargill</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Altshuler</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ireland</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sklar</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ardlie</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patil</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaw</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lane</surname>
            <given-names>CR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lim</surname>
            <given-names>EP</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalyanaraman</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nemesh</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ziaugra</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedland</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rolfe</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warrington</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lipshutz</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daley</surname>
            <given-names>GQ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lander</surname>
            <given-names>ES</given-names>
          </string-name>
          :
          <article-title>Characterization of single-nucleotide polymorphisms in coding regions of human genes</article-title>
          (vol
          <volume>22</volume>
          , pg 231,
          <year>1999</year>
          ).
          <source>Nat Genet</source>
          <year>1999</year>
          ,
          <volume>23</volume>
          (
          <issue>3</issue>
          ):
          <volume>373</volume>
          {
          <fpage>373</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Serrano</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guerois</surname>
            <given-names>R</given-names>
          </string-name>
          :
          <article-title>Fold-X: An algorithm to predict and engineer folding pathways</article-title>
          .
          <source>Abstr Pap Am Chem Soc</source>
          <year>2001</year>
          ,
          <volume>221</volume>
          :U395{
          <fpage>U395</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Guerois</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nielsen</surname>
            <given-names>JE</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano</surname>
            <given-names>L</given-names>
          </string-name>
          :
          <article-title>Predicting changes in the stability of proteins and protein complexes: A study of more than 1000 mutations</article-title>
          .
          <source>J Mol Biol</source>
          <year>2002</year>
          ,
          <volume>320</volume>
          (
          <issue>2</issue>
          ):
          <volume>369</volume>
          {
          <fpage>387</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Schymkowitz</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borg</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stricher</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nys</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rousseau</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano</surname>
            <given-names>L</given-names>
          </string-name>
          :
          <article-title>The FoldX web server: an online force eld</article-title>
          .
          <source>Nucleic Acid Res</source>
          <year>2005</year>
          ,
          <volume>33</volume>
          :W382{
          <fpage>W388</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Schymkowitz</surname>
            <given-names>JWH</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rousseau</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martins</surname>
            <given-names>IC</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferkingho - Borg</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stricher</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano</surname>
            <given-names>L</given-names>
          </string-name>
          :
          <article-title>Prediction of water and metal binding sites and their a nities by using the Fold-X force eld</article-title>
          .
          <source>Proc Natl Acad Sci USA</source>
          <year>2005</year>
          ,
          <volume>102</volume>
          (
          <issue>29</issue>
          ):
          <volume>10147</volume>
          {
          <fpage>10152</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Vriend</surname>
            <given-names>G</given-names>
          </string-name>
          :
          <article-title>What If - a molecular modeling and drug design program</article-title>
          .
          <source>J Mol Graph</source>
          <year>1990</year>
          ,
          <volume>8</volume>
          :
          <fpage>52</fpage>
          {.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Fernandez-Escamilla</surname>
            <given-names>AM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rousseau</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schymkowitz</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano</surname>
            <given-names>L</given-names>
          </string-name>
          :
          <article-title>Prediction of sequence-dependent and mutational e ects on the aggregation of peptides and proteins</article-title>
          .
          <source>Nat Biotechnol</source>
          <year>2004</year>
          ,
          <volume>22</volume>
          (
          <issue>10</issue>
          ):
          <volume>1302</volume>
          {
          <fpage>1306</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Maurer-Stroh</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuemmerer</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez de la Paz</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martins</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reumers</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rousseau</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schymkowitz</surname>
            <given-names>J</given-names>
          </string-name>
          :
          <article-title>Accurate prediction of sequence determinants of amyloid formation using the Waltz algorithm</article-title>
          .
          <source>Submitted</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Thompson</surname>
            <given-names>MJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sievers</surname>
            <given-names>SA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karanicolas</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ivanova</surname>
            <given-names>MI</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baker</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisenberg</surname>
            <given-names>D</given-names>
          </string-name>
          :
          <article-title>The 3D pro le method for identifying bril-forming segments of proteins</article-title>
          .
          <source>Proc Natl Acad Sci U S A</source>
          <year>2006</year>
          ,
          <volume>103</volume>
          (
          <issue>11</issue>
          ):
          <volume>4074</volume>
          {
          <fpage>4078</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Van Durme</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maurer-Stroh</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilkinson</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rousseau</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schymkowitz</surname>
            <given-names>J</given-names>
          </string-name>
          :
          <article-title>Accurate prediction of the sequence determinants of DnaK-peptide binding via a method that integrates homology modelling and experimental data</article-title>
          .
          <source>Submitted</source>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>McCarty</surname>
            <given-names>JS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rudiger</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schonfeld</surname>
            <given-names>HJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>SchneiderMergener</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakahigashi</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yura</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bukau</surname>
            <given-names>B</given-names>
          </string-name>
          :
          <article-title>Regulatory region C of the E. coli heat shock transcription factor, sigma32, constitutes a DnaK binding site and is conserved among eubacteria</article-title>
          .
          <source>J Mol Biol</source>
          <year>1996</year>
          ,
          <volume>256</volume>
          (
          <issue>5</issue>
          ):
          <volume>829</volume>
          {
          <fpage>37</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>