<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Network-Based Disease Candidate Gene Prioritization: Towards Global Di usion in Heterogeneous Association Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Joana P. Goncalves</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara C. Madeira</string-name>
          <email>smadeirag@kdbio.inesc-id.pt</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yves Moreau</string-name>
          <email>yves.moreau@esat.kuleuven.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>BIOI, ESAT-SCD, Department of Electrical Engineering</institution>
          ,
          <addr-line>KULeuven, Kasteelpark Arenberg 10, 3001 Leuven-Heverlee</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IST, Technical University of Lisbon</institution>
          ,
          <addr-line>Av. Rovisco Pais 1, 1049-001 Lisboa</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Knowledge Discovery and Bioinformatics group (KDBIO), INESC-ID</institution>
          ,
          <addr-line>Rua Alves Redol 9, 1000-029 Lisboa</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Disease candidate gene prioritization addresses the association of novel genes with disease susceptibility or progression. Networkbased approaches explore the connectivity properties of biological networks to compute an association score between candidate and diseaserelated genes. Although several methods have been proposed to date, a number of concerns arise: (i) most networks used rely exclusively on curated physical interactions, resulting in poor coverage of the Human genome and leading to sparsity issues; (ii) most methods fail to incorporate interaction con dence weights; (iii) in some cases, relevance scores are computed as local measures based on the direct interactions with the disease-related genes, ignoring potentially relevant indirect interactions. In this study, we seek a robust network-based strategy by evaluating the performance of selected prioritization strategies using genes known to be involved in 29 di erent diseases.</p>
      </abstract>
      <kwd-group>
        <kwd>protein-protein interaction</kwd>
        <kwd>network</kwd>
        <kwd>random walk</kwd>
        <kwd>disease candidate genes</kwd>
        <kwd>prioritization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Biomarkers play a crucial role in modern medical practice as a means of
improving accuracy in diagnosis, prognosis and treatment. In particular, research
has been actively devising associations of novel genes with disease susceptibility
or progression, relying on high-throughput technologies and the proliferation of
accessible resources of biological data to enable large-scale genome-wide studies.</p>
      <p>Most computational methods proposed for disease gene prioritization aim to
identify putative candidates based on their similarity with genes known to be
involved in the occurrence of a particular phenotype, according to: intrinsic
properties, functional annotations, coherent transcriptional responses via expression
data analysis, orthologous relations with genes from model organisms or even
co-occurrence in the literature [22]. Alternative strategies adopt a systemic
approach and explore the topology of biological networks, including protein-protein
interactions, regulatory data or metabolic pathways. These approaches rely on
the assumption that genes co-occurring in a particular network substructure or
interacting tend to participate together in related biological processes to identify
novel genes based on their linkage with the known disease genes [22].</p>
      <p>
        Integrative network-based analysis has been addressed [
        <xref ref-type="bibr" rid="ref11 ref8">8,11,15,16,20,23,26</xref>
        ],
combining knowledge from distinct resources in association networks to unravel
novel disease genes. However, most of these approaches rely solely on physical
interactions [
        <xref ref-type="bibr" rid="ref8">8,23</xref>
        ], potentially inferred via orthologous relations with model
organisms [26], often resulting in insu cient coverage of the Human genome. Others
include additional interactions predicted from coexpression, pathway, functional
or literature data, but still devise sparse networks [
        <xref ref-type="bibr" rid="ref11">11, 15</xref>
        ]. Although the risk for
false positive interactions may rise, the integration of knowledge from
heterogeneous sources generates denser networks which tend to be less biased toward a
particular evidence, more robust to noise and thus able to perform better in the
prioritization task [16].
      </p>
      <p>Network-based prioritization methods further di er in how they de ne the
ranking of the candidates from the known disease-related genes. Local measures
are usually computed based on the direct links or shortest paths between the
candidates and the disease-related genes [15, 16], while global strategies di use
or smooth a disease-related signal through the network. In this work, we
evaluate whether the latter should be preferred over the former, as the inclusion of
indirect associations is able to compensate for missing linkage, ultimately
mitigating sparsity and \small world" e ect issues [20], and global similarities have
recently been shown to outperform local measures [15].</p>
      <p>
        Random walks or di usion kernels arise as natural candidates for the
diffusion approach and their application to prioritization has been proven e
ective [
        <xref ref-type="bibr" rid="ref6 ref8">6, 8, 15, 23</xref>
        ]. Not only they compute fast using iterative methods, even for
large networks [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], they are also able to straightforwardly establish a ranking
of the candidates based on the global connectivity of the network.
Nevertheless, some of the proposed methods [
        <xref ref-type="bibr" rid="ref8">8, 15</xref>
        ] ignore or fail to incorporate weights
expressing the con dence on the evidence of every particular association [16].
Furthermore, their scores are based on the steady-state probability obtained
after a large number of iterations or upon convergence. In this study, we assess the
claim that limited di usion is usually su cient for ranking purposes [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ] and
on our intuition which leads us to expect the prior knowledge to be somehow
lost or of very little importance to the ranking after di using to a large extent.
      </p>
      <p>Throughout this paper, we address the aforementioned topics by
analyzing the performance of di erent prioritization strategies in three case studies:
(i) Integrative heterogeneous protein association network vs integrative
proteinprotein physical interaction network (PPPIN); (ii) Global ranking measure vs
local ranking measure; (iii) Con dence weights, degree of di usion and
parameter variation.</p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>
        A protein-protein association network can be described as a weighted undirected
graph, a special case of a weighted directed graph, de ned as G = (V; E), where
V is the set of vertices and E is the set of edges. Each vertex in V and edge
in E correspond to a gene and an association between two genes, respectively.
Let A and D denote the adjacency and diagonal matrices of G, respectively.
Auv is the weight w(u; v) of the edge (u; v) between source u and target v.
Also, Duu = P(u;v)2E Auv; 8u 2 V , that is, the sum of the weights of the edges
for which u is the source. Prioritizing disease candidates thus formulates as
obtaining a ranking on V given a set S 2 V of seed genes. For the local scoring
scheme Endeavour's measure was used [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. As global network-based strategies,
the PageRank with priors and Heat Di usion random walks were applied: an
initial signal expressing the relevance of the genes in the context of the disease
in the form of a preference vector, p(0), is di used over the network by performing
a limited number of iterations, N .
2.1
      </p>
      <p>
        Endeavour's Measure: Intersection of Interactors
Endeavour computes a local network-based measure, whereby the score of each
gene is computed as the overlap between the sets of genes interacting with the
seed genes and those interacting with the candidate gene itself [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]:
Sv =
      </p>
      <p>X
(u;v)2E
intSeeds(u)
intSeeds(u) =
1 , if 9z 2 S : (u; z) 2 E
0 , otherwise
2.2</p>
      <p>
        Heat Di usion and PageRank with Priors
Heat Di usion is a discrete approximation of the heat kernel [28] rst introduced
in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], in which the rate of di usion is controlled by a non-negative parameter,
the heat di usion coe cient t. The iterative equation is given by
p(i+1) =
v
1
t
N
p(i) + t
v N
      </p>
      <p>X
(u;v)2E
p(i)
u</p>
      <p>Auv :
Duu
PageRank with priors is an extension of the original PageRank algorithm to
consider the original probability distribution of the scores [25]. A parameter ,
called \back probability" expresses the probability of jumping to the initial node
at each iteration. The iterative equation is
p(i+1) =
v
p(0) + (1
v
)</p>
      <p>X</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        Evaluation sudies were performed using Human data from the STRING database
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and a PPPIN from Entrez Gene [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] as representatives of protein-protein
heterogeneous association and physical interaction networks, respectively. 620
genes known to be related with 29 diseases were used as prior knowledge to
prioritize candidates in a leave-one-out cross-validation scheme.
3.1
      </p>
      <p>
        Data and Preprocessing
Networks The STRING database [
        <xref ref-type="bibr" rid="ref12">12, 18</xref>
        ] integrates physical interactions and
predicted associations based on knowledge obtained from heterogeneous sources
of transcriptional, functional, metabolic, literature and orthology data. For a fair
comparison with Endeavour, we downloaded and parsed version 7.1 of STRING
[18], including evidences from MINT [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], HPRD [19], BIND [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], DIP [27],
BioGRID [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], KEGG [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and Reactome [24] databases. Associations from STRING
v8.2 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] were also retrieved to assess to which extent the additional knowledge
integrated from IntAct [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], PID [21] and GO [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] protein complexes would
improve the prioritization performance relative to the previous release. A PPPIN
was downloaded from the NCBI Entrez Gene FTP repository [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. 130797 Human
interactions were selected from 448534 entries, for which both interactant genes
were tagged with tax ID 9606. From these, 4611, 51275 and 74911 were originally
from BIND [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], BioGRID [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and HPRD [19], respectively. Genes' identi ers
followed Entrez Gene nomenclature. Preprocessing of these networks involved
ltering redundant edges and devising an explicit representation of a directed
graph. In the case of the STRING releases, original weights were used to express
the con dence of every association, while in the PPPIN all edges were attributed
weight 1. STRING v7.1 contained 16050 genes and 698534 unique associations.
STRING v8.2 covered 17448 Human genes with 1256016 non-redundant
associations. Finally, the PPPIN had 47873 physical interactions between 10175 genes.
Seed sets 620 disease genes were selected from the OMIM [17] database
spanning 29 disease-speci c sets, with an average of 21 genes per set. As genes were
identi ed according to Ensembl nomenclature, the seeds could be directly used
with STRING. For the PPPIN, however, we performed a conversion between
Ensembl and Entrez Gene identi ers. A mapping was parsed from a le downloaded
from the NCBI Entrez Gene FTP repository [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and used to generate the
corresponding seed sets using Entrez Gene names. Additionally, we ltered the genes
absent from at least one of the networks or for which the conversion between
Ensembl and Entrez did not succeed. In total, 94 seed genes were lost (14 with
no conversion, 80 absent from the PPPIN). A single occurrence of a gene with
several Entrez aliases happened. In this case, only the alias present in the PPPIN
was kept. For validation purposes, seed sets containing randomly selected genes
were generated. The number of seeds in each set was randomly chosen in the
range [5; 100] and the genes were randomly selected from the Human STRING
v8.2 network. 546 genes were retrieved.
3.2
      </p>
      <p>Evaluation Measures and Experimental Setting
Evaluation measures Ideally, in a leave-one-out cross-validation scheme, we
would expect the prioritization strategy to rank the left-out gene known to be
related with the disease at the top. Under this assumption, we assess the
performance of the scoring methods overall and per disease based on four evaluation
measures: the number of left-out genes ranked in the top 10 and 20 positions,
the Area under the ROC curve (AUC) score, and the mean average precision.</p>
      <p>For a given combination of di usion parameter and number of iterations
N , n rankings are generated (one per left-out gene). The AUC score is given by
SAUC ;N =
n</p>
      <p>Pn r(N)
k=1 mk(kN) ;
n
where r(N) is the ranking position of the kth left-out gene in the kth ranked list
k
and m(kN) is the number of ranked genes in the kth list.</p>
      <p>Mean average precision (MAP) is an evaluation measure that combines
precision and recall. Essentially, MAP averages the precisions computed by truncating
the list after each of the relevant entities is found. Only one relevant entity must
be found, the left-out gene. Thus, precision at rank r is either 0, before it has
been found, or 1r . Moreover, in our setting the ranked lists contain equal number
of genes, allowing us to simplify our MAP score for n lists with the same size to:
SMAP ;N =</p>
      <p>Pn 1
k=1 r(N)</p>
      <p>k
n
Experimental setting In each validation run, one di erent gene was deleted
from the set of seed genes and added to 99 randomly selected candidate genes.
A ranking method was then applied to compute a score for every gene in the
network. Finally, the ranking of the 100 candidate genes was de ned according
to the retrieved scores. In the case of the Heat Di usion, the scores of the seed
genes were initialized to 1. For PageRank, an initial seed score of 1=jSj was used.
Performance was assessed by computing AUC and MAP scores, and counting the
number of left-out genes ranked in the top 10 and top 20 positions, both overall
and per disease. We sought the best performance of each method using several
combinations of parameters. Heat di usion coe cients t and back probabilities
of 0.1, 0.3, 0.5, 0.7 and 0.9 with 2, 5, 10, 15 and 20 iterations using STRING and
2, 5, 10, 20, 100 iterations using the PPPIN were tried. In the case studies, results
are shown only for the parameter settings which achieved the best performance
in each case. We further ranked the randomly generated seed sets using the
leaveone-out cross-validation in STRING v8.2 to assess whether the Heat Di usion
method was able to take advantage of the information contained in the seed sets
to improve the identi cation of the left-out seeds. Overall, AUC and MAP scores
of 0.501 and 0.05 were achieved and only 57 and 92 genes were ranked in the top
10 and 20 positions. Similar results were obtained per seed set (data not shown),
in accordance with what would have been expected for random seed sets.
3.3
Heat Di usion and PageRank with priors achieved similar results in both
networks (Table 1). For this reason, we abstain ourselves of comparing the results
of both random walks, considering the results equivalent when applied to the
same network. Throughout this section, we will always refer to one of them as
a representative of a global measure. A brief description of the prioritization
performances obtained for each case study follows.</p>
      <sec id="sec-3-1">
        <title>Method Network Parameters AUC MAP TOP 10 TOP 20 #BRM #BRN HeatDi usion STRING8 t = 0:3; N = 10 0:962 0:711 484 502 26% 68% PageRank STRING8 = 0:7; N = 2 0:961 0:693 485 502 20% 69%</title>
      </sec>
      <sec id="sec-3-2">
        <title>HeatDi usion PPPIN</title>
        <p>
          PageRank PPPIN
Global measure vs Local measure A network-based global ranking was
obtained using the Heat Di usion method with t = 0:3; N = 10, while Endeavour
[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] was used to score the genes using its local measure. Both rankings were
based on STRING v7.1, the version included in Endeavour. Overall, the random
walk global measure outperformed the local interaction overlap in all evaluation
measures (see Table 2), that is, the higher number of left-out genes was ranked
on the top positions, also achieving better ranks in general, using the latter.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Method Network AUC MAP TOP 10 TOP 20</title>
        <p>HeatDi usion (t = 0:3; N = 10) STRING v7.1 0:942 0:643 536 569
Endeavour STRING v7.1 0:806 0:326 393 464</p>
        <p>Regarding the AUC scores per disease (see Table 3), the Heat Di usion
method outperformed Endeavour in all diseases except Ehlers-Danlos syndrome
(0.944 opposed to 0.948, respectively). This was also the only disease for which
the number of genes ranked in the top 20 positions was higher using the local
measure (Endeavour was able to rank one more gene in the top 20). However,
the MAP score was better for the Heat Di usion method and, in fact, 9 of the 10
seed genes ranked in the top 10 positions by both methods scored higher using
the global measure.</p>
        <p>For the remaining diseases, Heat Di usion was always able to rank the same
or a higher number of genes in both the top 10 and the top 20 positions.
Regarding the MAP scores, Heat Di usion outperformed Endeavour in every disease
and was able to rank all genes of both amyotrophic lateral sclerosis and Usher
syndrome in the rst position.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Disease</title>
      </sec>
      <sec id="sec-3-5">
        <title>Alzheimer's disease</title>
        <p>amyotrophic lateral sclerosis
anemia
breast cancer
cardiomyopathy
cataract</p>
      </sec>
      <sec id="sec-3-6">
        <title>Charcot-Marie-Tooth disease</title>
        <p>colorectal cancer
deafness
diabetes
dystonia
Ehlers-Danlos syndrome
emolytic anemia
epilepsy
ichthyosis
leukemia
lymphoma
mental retardation
muscular dystrophy
myopathy
neuropathy
obesity</p>
      </sec>
      <sec id="sec-3-7">
        <title>Parkinson's disease</title>
        <p>retinitis pigmentosa
spastic paraplegia
spinocerebellar ataxia</p>
      </sec>
      <sec id="sec-3-8">
        <title>Usher syndrome</title>
        <p>xeroderma pigmentosum</p>
      </sec>
      <sec id="sec-3-9">
        <title>Zellweger syndrome</title>
        <p>
          Protein-Protein Associations vs Protein-Protein Physical Interactions
Heat Di usion achieved better performance using STRING v8.2, with AUC score
0.962, opposed to 0.862 using the PPPIN (see Table 1). Furthermore, STRING
enabled to rank more than 90% of the genes in the top 10 positions, while using
the PPPIN less than 60% were in top 10. In a one-to-one comparison, Heat
Di usion ranked 68% of the genes better using STRING, while only 11% of the
ranks were better using the PPPIN. Table 4 compares the results obtained for
the Heat Di usion method using STRING v8.2 with PageRank with priors in a
PPPIN, one of the best performing strategies in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], per disease.
        </p>
      </sec>
      <sec id="sec-3-10">
        <title>Disease</title>
      </sec>
      <sec id="sec-3-11">
        <title>Alzheimer's disease</title>
        <p>amyotrophic lateral sclerosis
anemia
breast cancer
cardiomyopathy
cataract</p>
      </sec>
      <sec id="sec-3-12">
        <title>Charcot-Marie-Tooth disease</title>
        <p>colorectal cancer
deafness
diabetes
dystonia
Ehlers-Danlos syndrome
emolytic anemia
epilepsy
ichthyosis
leukemia
lymphoma
mental retardation
muscular dystrophy
myopathy
neuropathy
obesity</p>
      </sec>
      <sec id="sec-3-13">
        <title>Parkinson's disease</title>
        <p>retinitis pigmentosa
spastic paraplegia
spinocerebellar ataxia</p>
      </sec>
      <sec id="sec-3-14">
        <title>Usher syndrome</title>
        <p>xeroderma pigmentosum</p>
      </sec>
      <sec id="sec-3-15">
        <title>Zellweger syndrome</title>
        <p>Regarding the disease-speci c scores (see Table 4), the lowest AUC (and
MAP) values for the combination Heat Di usion and STRING v8.2 were of
0.926 (0.727) for mental retardation, and 0.930 (0.476) for lymphoma, which are
still good results. For ve diseases, namely amyotrophic lateral sclerosis,
EhlersDanlos syndrome, spastic paraplegia, Usher syndrome and Zellweger syndrome,
the heterogeneous association network approach was actually able to rank all the
seed genes in the rst position of the ranking. On the other hand, the PageRank
di usion in the PPPIN achieved AUC scores above 0.9 only for two diseases:
colorectal cancer with 0.912 and xeroderma pigmentosum with 0.98. The lowest
AUC and MAP scores were obtained for amyotrophic lateral sclerosis (0.53 and
0.028) and spastic paraplegia (0.49 and 0.083). The PPPIN strategy could not
rank any of the seed genes for amyotrophic lateral sclerosis in the top 10 positions
and only one was identi ed in the rst 20. Also, only one gene out of the 5 seeds
for spastic paraplegia was ranked in the top 10/20. In this case, the performance
for both diseases is comparable to the one obtained using the random seed sets
(data now shown).</p>
        <p>Con dence weights, number of iterations and di usion rate We assessed
the contribution of STRING's weights expressing the degree of con dence in the
associations between genes to the performance of the prioritization method by
di using the initial preference vector using the ltered disease-speci c seed sets
on the network after setting all associations' weights to 1. Although the resulting
AUC and MAP scores (0.957 and 0.662) were not substantially di erent from
the ones obtained using the con dence weights (0.962 ans 0.711), they actually
re ected in less 9 genes ranked in the top 10 (data not shown). Overall, the
number of genes in the top 20 was the same, with slight variations per disease.
From the ve diseases achieving maximum performance in the di erentially
association weighted setting, only for Ehlers-Danlos syndrome, spastic paraplegia
and Zellweger syndrome these results could be maintained.</p>
        <p>In both random walk approaches, the best results were achieved using a
limited number of iterations. STRING v8.2 provided consistent and stable
performance when varying the number of di usion steps. On the PPPIN, the best
ranking was always obtained using two iterations. It would then stabilize for
larger numbers of steps, although measuring considerably lower in the
evaluation, since it was never able to rank more than 289 or 346 genes - out of 526
in the top 10 and top 20, respectively.</p>
        <p>Regarding the parameter controlling the rate of di usion, the Heat Di usion
method delivered quite similar performance for the set of heat coe cients tried:
in STRING v8.2, resulting in AUC scores ranging from 0.960 to 0.962 for each
di usion coe cient, considering equal number of iterations; in the PPPIN, AUC
scores ranging between 0.859 and 0.862 with 2 iterations, N = 2, and between
0.766 and 0.771 using 5, 10, 20 and 100 iterations.These results indicate its
robustness to variations in this parameter. For PageRank with priors, the impact
of the back probability value was not neglegible. For the lowest back
probabilities (0.01 and 0.05) the scores were unstable leading to considerable performance
variations, even using STRING v8.2. For = f0:1; 0:3; 0:5; 0:7; 0:9g, the
PageRank AUC scores in STRING v8.2 varied between 0.936 and 0.961 considering the
results obtained using the same number of iterations. In the PPPIN, PageRank
obtained AUC scores between 0.859 and 0.861 using 2 iterations and ranging
between 0.758 and 0.775 using 5, 10, 20 and 100 iterations.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>Prioritization results con rmed our hypothesis that networks integrating gene
associations retrieved or predicted using data from heterogeneous sources should
be in general more informative and potentially able to perform better in the
identi cation of genes associated with a given disease when compared to networks
containing only physical interactions. Advantages of the former are supported by
three key observations: (1) associations derived from the combination of several
types of evidence should be more reliable and accurate; (2) heterogeneous data
integration enables a better coverage of the genome and larger network density,
confering robustness to noise; (3) con dence weights can be devised in order to
di erentiate associations and mitigate the impact of false positive associations,
particularly when based on a limited number of sources.</p>
      <p>Nevertheless, our analysis shows that heterogeneous association networks do
not present su cient guarantee for maximum performance by themselves. In fact,
the network-based score measuring the degree of relatedness of each candidate
gene with a given disease based on a set of known disease-related genes proved
to play a major role. Essentially, based on the results we could conclude that in
comparison to neighborhood-limited scores a network-based measure able to
capture global connectivity properties by considering indirect associations between
genes is not only (1) more robust, as it compensates for the sparsity related to
direct associations and tackles the \small world" e ect issue; but also (2) more
informative, deriving a score based on a systemic view of the interactome. This
claim has also been previously hinted at in [15, 16].</p>
      <p>
        Propagation schemes tested in the computation of global network-based scores
di used an initial preference vector expressing the distribution of the known
disease-related genes through the network using random walks. These methods
compute fast using iterative procedures, even for large networks. Furthermore,
we could verify that in the context of prioritization in association or physical
interaction networks the maximum performance can be achieved using only a
limited number of iterations. Heat Di usion and PageRank with priors delivered
high quality results and achieved similar performance under appropriate
parameter settings, supporting the claim of equivalence [
        <xref ref-type="bibr" rid="ref8">8, 25</xref>
        ] for other approaches of
the same kind, namely HITS with priors and K-Step Markov. The importance of
con dence weights was inconclusive, as the di erence in performance exhibited
by our experiments was residual. We believe, however, that appropriate
association con dence weights may improve accuracy of network-based prioritization
results.
      </p>
      <p>Acknowledgments This work was partially supported by FCT (INESC-ID
multiannual funding) through the PIDDAC Program funds. JPG is the recipient
of a doctoral grant supported by FCT (SFRH/BD/36586/2007).
IntAct{open source resource for molecular interaction data. Nucleic Acids Research
35(suppl 1), D561{565 (2007)
15. Kohler, S., Bauer, S., Horn, D., Robinson, P.: Walking the interactome for
prioritization of candidate disease genes. The American Journal of Human Genetics
82(4), 949958 (2008)
16. Linghu, B., Snitkin, E.S., Hu, Z., Xia, Y., Delisi, C.: Genome-wide prioritization of
disease genes and identi cation of disease-disease associations from an integrated
human functional linkage network. Genome biology 10(9), R91 (2009)
17. McKusick, V.A.: Mendelian Inheritance in Man and Its Online Version, OMIM.</p>
      <sec id="sec-4-1">
        <title>The American Journal of Human Genetics 80(4), 588{604 (2007)</title>
        <p>18. von Mering, C., Jensen, L.J., Kuhn, M., Cha ron, S., Doerks, T., Kruger, B., Snel,</p>
      </sec>
      <sec id="sec-4-2">
        <title>B., Bork, P.: STRING 7{recent developments in the integration and prediction of</title>
        <p>protein interactions. Nucleic Acids Research 35(Database issue), D358{62 (2007)
19. Mishra, G.R., Suresh, M., Kumaran, K., Kannabiran, N., Suresh, S., Bala, P.,</p>
      </sec>
      <sec id="sec-4-3">
        <title>Shivakumar, K., Anuradha, N., Reddy, R., Raghavan, T.M., Menon, S., Hanu</title>
        <p>manthu, G., Gupta, M., Upendran, S., Gupta, S., Mahesh, M., Jacob, B., Mathew,</p>
      </sec>
      <sec id="sec-4-4">
        <title>P., Chatterjee, P., Arun, K.S., Sharma, S., Chandrika, K.N., Deshpande, N., Pal</title>
        <p>vankar, K., Raghavnath, R., Krishnakanth, R., Karathia, H., Rekha, B., Nayak, R.,</p>
      </sec>
      <sec id="sec-4-5">
        <title>Vishnupriya, G., Kumar, H.G.M., Nagini, M., Kumar, G.S.S., Jose, R., Deepthi, P.,</title>
      </sec>
      <sec id="sec-4-6">
        <title>Mohan, S.S., Gandhi, T.K.B., Harsha, H.C., Deshpande, K.S., Sarker, M., Prasad,</title>
      </sec>
      <sec id="sec-4-7">
        <title>T.S.K., Pandey, A.: Human protein reference database{2006 update. Nucleic Acids</title>
        <p>Research 34(suppl 1), D411{414 (2006)
20. Nitsch, D., Tranchevent, L.C., Thienpont, B., Thorrez, L., Van Esch, H., Devriendt,</p>
      </sec>
      <sec id="sec-4-8">
        <title>K., Moreau, Y.: Network analysis of di erential expression for the identi cation of</title>
        <p>disease-causing genes. PloS ONE 4(5), e5526 (2009)
21. Schaefer, C.F., Anthony, K., Krupa, S., Bucho , J., Day, M., Hannay, T., Buetow,</p>
      </sec>
      <sec id="sec-4-9">
        <title>K.H.: PID: the Pathway Interaction Database. Nucleic Acids Research 37(suppl 1),</title>
        <p>D674{679 (2009)
22. Ti n, N., Andrade-Navarro, M.A., Perez-Iratxeta, C.: Linking genes to diseases:
it's all in the data. Genome Medicine 1(8), 77 (Jan 2009)
23. Vanunu, O., Magger, O., Ruppin, E., Shlomi, T., Sharan, R.: Associating Genes and</p>
      </sec>
      <sec id="sec-4-10">
        <title>Protein Complexes with Disease via Network Propagation. PLoS Computational</title>
        <p>Biology 6(1) (2010)
24. Vastrik, I., D'Eustachio, P., Schmidt, E., Joshi-Tope, G., Gopinath, G., Croft, D.,
de Bono, B., Gillespie, M., Jassal, B., Lewis, S., Matthews, L., Wu, G., Birney, E.,</p>
      </sec>
      <sec id="sec-4-11">
        <title>Stein, L.: Reactome: a knowledge base of biologic pathways and processes. Genome</title>
        <p>Biology 8(3), R39 (2007)
25. White, S., Smyth, P.: Algorithms for estimating relative importance in networks.</p>
      </sec>
      <sec id="sec-4-12">
        <title>In: KDD '03: Proceedings of the Ninth ACM SIGKDD International Conference</title>
        <p>on Knowledge Discovery and Data Mining. pp. 266{275. ACM, New York, NY,
USA (2003)
26. Wu, X., Jiang, R., Zhang, M.Q., Li, S.: Network-based global inference of human
disease genes. Molecular Systems Biology 4(189), 189 (2008)
27. Xenarios, I., Rice, D.W., Salwinski, L., Baron, M.K., Marcotte, E.M., Eisenberg,</p>
      </sec>
      <sec id="sec-4-13">
        <title>D.: DIP: the Database of Interacting Proteins. Nucleic Acids Research 28(1), 289{</title>
        <p>291 (2000)
28. Yang, H., King, I., Lyu, M.: Di usionrank: a possible penicillin for web
spamming. In: Proceedings of the 30th Annual International ACM SIGIR Conference
on Research and Development in Information Retrieval. p. 438. ACM (2007)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>NCBI</given-names>
            <surname>Entrez Gene FTP Repository</surname>
          </string-name>
          (
          <year>Jan 2010</year>
          ), ftp://ftp.ncbi.nih.gov/gene/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Aerts</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Van Loo,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>De Smet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Lambrechts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Maity</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Tranchevent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.C.</given-names>
            ,
            <surname>De Moor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Coessens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Marynen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Carmeliet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Moreau</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          :
          <article-title>Gene prioritization through genomic data fusion</article-title>
          .
          <source>Nature Biotechnology</source>
          <volume>24</volume>
          (
          <issue>5</issue>
          ),
          <volume>537</volume>
          {
          <fpage>44</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ball</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blake</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Botstein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Butler</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cherry</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dolinski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dwight</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eppig</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <article-title>Others: Gene Ontology: tool for the uni cation of biology</article-title>
          .
          <source>Nature Genetics</source>
          <volume>25</volume>
          (
          <issue>1</issue>
          ),
          <volume>2529</volume>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bader</surname>
            ,
            <given-names>G.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Betel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hogue</surname>
            ,
            <given-names>C.W.V.</given-names>
          </string-name>
          :
          <article-title>BIND: the Biomolecular Interaction Network Database</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>31</volume>
          (
          <issue>1</issue>
          ),
          <volume>248</volume>
          {
          <fpage>250</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Breitkreutz</surname>
            ,
            <given-names>B.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reguly</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boucher</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Breitkreutz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Livstone</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oughtred</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lackner</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bahler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dolinski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tyers</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The BioGRID Interaction Database: 2008 update</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>36</volume>
          (
          <issue>suppl 1</issue>
          ), D637{
          <volume>640</volume>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Can</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camoglu</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          :
          <article-title>Analysis of protein-protein interaction networks using random walks</article-title>
          .
          <source>In: Proceedings of the 5th International Workshop on Bioinformatics - BIOKDD '05</source>
          . p.
          <fpage>61</fpage>
          . ACM Press, New York, New York, USA (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Chatr-Aryamontri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanzoni</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cesareni</surname>
          </string-name>
          , G.:
          <article-title>Searching the protein interaction space through the MINT Database</article-title>
          .
          <source>Methods in Molecular Biology</source>
          <volume>484</volume>
          ,
          <issue>305</issue>
          {
          <fpage>317</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aronow</surname>
            ,
            <given-names>B.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jegga</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>Disease candidate gene identi cation and prioritization using protein interaction networks</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>10</volume>
          ,
          <issue>73</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Chung</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yau</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Coverings, heat kernels and spanning trees</article-title>
          .
          <source>Electronic Journal of Combinatorics</source>
          <volume>6</volume>
          ,
          <issue>R12</issue>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Francisco</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goncalves</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madeira</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Using personalized ranking to unravel relevant regulations in the Saccharomyces cerevisiae regulatory network</article-title>
          .
          <source>In: Jornadas de Bioinformatica</source>
          <year>2009</year>
          . Lisbon, Portugal (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Franke</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bakel</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fokkens</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jong</surname>
            , D., E.d, Egmont-petersen,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wijmenga</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Reconstruction of a functional human gene network, with an application for prioritizing positional candidate genes</article-title>
          .
          <source>The American Journal of Human Genetics</source>
          <volume>78</volume>
          ,
          <issue>1011</issue>
          {
          <fpage>1025</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Jensen</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuhn</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stark</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cha ron</surname>
          </string-name>
          , S.,
          <string-name>
            <surname>Creevey</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muller</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doerks</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Julien</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonovic</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bork</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>von Mering</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>STRING 8{a global view on proteins and their functional interactions in 630 organisms</article-title>
          .
          <source>Nucleic acids research</source>
          37(Database issue),
          <source>D412{6</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kanehisa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goto</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kawashima</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Okuno</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hattori</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The KEGG resource for deciphering the genome</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>32</volume>
          (
          <issue>suppl 1</issue>
          ), D277{
          <volume>280</volume>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kerrien</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alam-Faruque</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aranda</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bancarz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bridge</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derow</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dimmer</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feuermann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrichsen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huntley</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kohler</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khadake</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leroy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liban</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lieftink</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montecchi-Palazzi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orchard</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Risse</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robbe</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roechert</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thorneycroft</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Apweiler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Hermjakob</surname>
          </string-name>
          , H.:
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>