<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A new database for drug-discovery address key-issues in mining of knowledge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ole Kristian Ekseth</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Svein-Olav Hvasshovd</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science (IDI) NTNU</institution>
          ,
          <addr-line>Trondheim</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The life of individuals are strongly in uenced by their health. An example concerns salinity resistant plants, an invention which may alleviate issues of climate change and rising sea levels. A di erent issue conserns drug discovery for humans, such as accurate and inexpensive cures available for the poor, personalized drugs, etc. In drug discovery the applied strategy is to combine domain experts with data made accessible through o -the-shelf software, and from the latter expect to identify new drugs. While computational drug-discovery is known to be working when number of candidate-factors are su ciently small, the established methods and software are unfeasible for mining in big-data knowledge bases. In this paper we address the above issue. We present an holistic approach for searches in big-data with complex relations. We demonstrate how our novel strategies for integration of large heterogeneous datasets results in knowledge discovery. In our work we address issues of: semantics, entity similarity, clustering, data-engine, hypothesis testing, and user-interfaces. To verify our approach we implement data from 37 external data-resources, resulting in a database with more than 30 million bio-medical relationships. When we compare our ndings with existing literature we observe how our holistic approach for big-data mining discover 1000+ novel candidates for drug interaction. To address keyissues in knowledge discovery we have constructed 10+ new softwareapproaches for data-mining, tools which enable the development of a new method for mining of big-data. To enable reuse of our approaches, they are available from: http://www.knittingTools.org/, http://www. knittingTools.org/gui_lib_mine.cgi, https://bitbucket.org/oekseth/ mine-data-analysis/downloads/, and https://bitbucket.org/oekseth/ hplysis-cluster-analysis-software.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In life-science a recurring task is to understand how and why entities relate: to
construct a hypothesis which translates discrete observations into a conceptual
gure capturing core-traits of an evaluated subject, as exempli ed in Fig. 3 for
the research of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. An example concerns the e ects of Cytoplasmic
Phospholipase A2 (cP LA2) enzyme which is associated to a number of diseases, such as
      </p>
      <p>
        Copyright held by the author(s). NOBIDS 2017
Alzheimer [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Rhematism [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In knowledge-discovery researchers use
manual approaches to identify candidate interactions, as exempli ed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] where the
authors use literature to manually construct a \heterogeneous network with 351
node" [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        In contrast to established approaches for data-mining, an understanding of
drug-interactions require the analysis of possible interactions, as exempli ed in
Fig. 1. While the \PubMed" database [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] contains \more than 27 million
citations for biomedical literature" [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the \Uni ed Protein Resources (UniProt)"
describes more than 47 million protein sequences [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. \The Economist" asserts
that 50 per-cent of published research-literature are erroneous [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], hence the
established use of manually selected research ndings to identify new drug
candidates is challenging.
      </p>
      <p>
        The high cost of drug development discourage the development of drugs for
the poor [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The cost of developing a single drug vary from $802 million to
$2.2 billion [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The drug-company of \AstraZeneca" spend on average $11+
billion ([
        <xref ref-type="bibr" rid="ref10">10,11</xref>
        ]) on each accepted drug. The main-cost of drug development is
the number of failed drugs [12], e.g., as observed by [13]: \only, one in 5,000
medicines makes it to the marked" [14]. Of importance is to address the above
issues in drug discovery, i.e., as \today's pharmaceutical industry cannot sustain
su cient innovation" [15] with today's cost of drug-development. Hence, the
importance of accurate tools for knowledge discovery.
      </p>
      <p>In this paper we relate the above perspectives, demonstrating how a new
holistic approach for mining of big data enable user-interactive drug discovery.
In the method and associated software we unify the approaches of user-centric
and software-centric approaches for data-mining, as depicted in Fig. 5. What
we assert is that an holistic approach which increase accuracy and performance
of data from disparate sources, software for mining, and tacit understanding, is
su cient to address major issues in drug discovery, a view supported by [15].
\R&amp;D e ciency represents the ability of an R&amp;D system to translate inputs (for
example, ideas, investments, e ort) into de ned outputs (for example, internal
milestones that represent resolved uncertainty for a given project or product
launches), generally over a de ned period of time" [15]. The ensemble of methods
and software, summarized in Fig. 5, address challenges which have prevented
established semantic data-bases from knowledge discovery, e.g., as observed with
respect to the issues encountered by [17,?,19].</p>
      <p>In the work we have identi ed and addressed the issues of:
1. Disparate data: automatic approaches to unify distinctively di erent data,
where results are exempli ed in Fig. 1.
2. Execution-time: high-performance software for accurate and large-scale
datamining, as exempli ed in Fig. 4;
3. User searches: interactive real-time data-mining which stimulate use of tacit
knowledge, as exempli ed in Fig. 3 and Fig. 6.</p>
      <p>The remainder of the paper is organised as follows. In section 2 we brie y
survey related approaches, before we in section 3 describe the approach. In the
result-section 4 we identify evaluate/discuss how the holistic approach address
10,000
s
r
i
a
)p8,000
X
=
e
c
n
ta6,000
s
i
d
,
e
t
a
ic4,000
d
e
r
p
(
f
to2,000
n
u
o
C
0
1</p>
      <p>The semantic inference distance for predicates
activates state change
regulates state change
active in pathway through casuality
state transition
cites research
state transition for protein–protein
2
3
4
5
6
7
8</p>
      <p>Semantic inference distance.
current issues in big-data mining. This paper ends with a brief summary of
observations in section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>A challenge in data-mining conserns the slow performance of software, as
observed for [20] in Fig. 4. A possible explanation of the latter is an unawareness
of high-performance software implementation strategies [21]. To exemplify, the
major e orts in \systems biology is on developing fundamental computational
and informatics tools" [22], an assertion motivated by how \a concerted e ort
to bring all the useful tools for pathway analysis in a common platform is still
missing" [23]. When combining the observations of ([22,?]) and Fig. 4 we
realize how poor-performing software represents a hurdle in knowledge discovery.
To summarize, we observe that established approaches for data-mining su ers
from: 1
1. Disparate data: insu cient data-coverage and prediction, e.g., in [23,24,25,26,17];
2. Execution-time: high query response-time, e.g., in [20,?];</p>
      <sec id="sec-2-1">
        <title>User asks a question/query. web-interface</title>
      </sec>
      <sec id="sec-2-2">
        <title>View result. &lt;Logical Module&gt;</title>
      </sec>
      <sec id="sec-2-3">
        <title>Parse: apply ‘syntactic sugar’, and ‘validation of syntax’. web-server</title>
      </sec>
      <sec id="sec-2-4">
        <title>Parse: concatenate ‘provenance’ and ‘relationship’.</title>
      </sec>
      <sec id="sec-2-5">
        <title>Request query to be computed. data-server::client</title>
      </sec>
      <sec id="sec-2-6">
        <title>Interpret/parse query. data-server::computer</title>
      </sec>
      <sec id="sec-2-7">
        <title>Receive result, and handle</title>
        <p>warnings (ie, if any).</p>
        <p>Construct result-representation.
y
it
x
e
l
p
m
o
C
l
a
c
i
g
o
L
:
d
n
e
g
e
L</p>
      </sec>
      <sec id="sec-2-8">
        <title>Compute query. &lt;Logical Operation&gt;</title>
        <p>
          Caption: Above cartoon/figure depicts/illustrate the ‘flow of data’ initiated when a user requests the computation of a query: how a large number of ‘code/logical modules’ are involved in an
operation which from the outside (ie, by a non-expert) may be seen as a trivial task. Of interest is to observe the modules involved in the handling ‘of the user-request’: for the “web-interface”
CSS FandigHT.M2L:areC‘reospmonspiblue’ tfoarttheiolaynouat,lwhcileoJQmuepry,lJeaxvaiStcryipt avndePresrl-uCGsI huansdleesrt-hei nustere-inrpautc;ttheiousner-inipnut is krenceoivewd/hlaenddlegd eby thseePaerrl-CcGhI-ebass.ed T“wehb-seerver”
(currently running on an “Apache” ‘framework) which construct a valid KNIT-TSV query; the KNIT-TSV query is received by the “data-server::client” written in “C” (where the
prograambmoinvg-leanguagge uofr“Ce” isrcheolsaenttoessimpulifystheerin-tsereacatiorncwhithethse “Atpoachet” hweeb-secrvoerm),wphicuhitniatiatteisoancoamlmuniincaftioenrweithnKcniettingp-ToroolscSKeBs-sse.rveTr(ahnditshererafete-r sending the
query in question to the SKB-server); Knitting-Tools “data-server::computer” (which is written in the programming-languages of “C” and “C++”) accesses the ‘already loaded
SKB-daetac-otbjsect’t,bhefeore evoaluwatingothfe qiunerfyoanrd mretaurntinigotnhe,resaulnttodthec“doatma-seprvleer::xcliietnty”, whoicfh asgaoinfrtewturnathreer,esufltotortheuPserel-CrG-Iq-buaseedr“iweebs-sesrvuerp”.Tpheo“wretbe-sderver” (a)
relate the provenance-relationships to each of the relationships (thereby simplifying the viewing of the result), (b) construct “JavaScript” and HTML data-structures (for simplifying and
‘beauatiftyingw’thwewus.erkexnpeiritentceianndgusTero-inotelracstio.n)o,brefgor.e (c) Winithiatiinlge/conmstruactijnogdrifferreentsdeataa-rrecprhesenetatioonsr/latysoutsa(erge,tabilens,vheeast-tmeapds, raidnial viiemws, petcr).oTvheirnesgult is then
seen by the user (though the web-interface), where the JQeury and JavaScript support allow user-interaction (ie, without the data to be reloaded for each/every request).
performance of software categorized in the upper part of the gure, their return on
investment is limited, i.e., as captured from the gure. In contrast the majority of our
e orts are invested on improving the software modules depicted in the bottommost
part in the above gure, as exempli ed in Fig. 5.
3. User searches: user-interfaces which limits domain-experts from accurate
data-searches and result-interpretation, e.g., in [
          <xref ref-type="bibr" rid="ref11 ref12 ref13">18,28,29,30</xref>
          ].
        </p>
        <p>In below we brie y examine the above issues, focusing on issues concerned
with disparate data and execution-time.
2.1</p>
        <sec id="sec-2-8-1">
          <title>Execution-Time: Tools for similarity, feature-selection and clustering</title>
          <p>
            There are more than 106 research-works concerned with data-mining . An
exam1
ple concerns the k-means cluster-algorithm, where new permutations are
published every year, e.g., with respect to [
            <xref ref-type="bibr" rid="ref14">31</xref>
            ]. The work of [
            <xref ref-type="bibr" rid="ref15">32</xref>
            ] observes how
existing software for data-analysis under-utilizes computer-hardware. A popular
software-tool for cluster-analysis is the \cluster C" software [20]. From Fig. 4
we observe how the approach manages to outperform the software of [20] by
a factor of 100,000x+. While [
            <xref ref-type="bibr" rid="ref16">33</xref>
            ] provide a GPU-optimized implementation of
\DB-SCAN" [
            <xref ref-type="bibr" rid="ref17">34</xref>
            ], the GPU implementations limited support for user-de ned
parameters result in inaccurate cluster-predictions for numerous data-sets [
            <xref ref-type="bibr" rid="ref18">35</xref>
            ], e.g.,
1
          </p>
          <p>
            Observation from searching on \Google Scholar" for terms such as PCA, k-means,
Sum of Squared Error (SSE), spearman, Euclid, correlation, similarity, etc.
with respect to issues in missing data and similarity-metrics. In order to
evaluate accuracy of cluster-algorithms, application of feature-selection, and
manydimensional hypothesis-testing, metrics for cluster-consistency are used [
            <xref ref-type="bibr" rid="ref19">36</xref>
            ].
Examples of cluster-consistency metrics are \Silhouette" [
            <xref ref-type="bibr" rid="ref20">37</xref>
            ], \Sum of Squared
Error (SSE)" [
            <xref ref-type="bibr" rid="ref21">38</xref>
            ], \Rand's Index" [
            <xref ref-type="bibr" rid="ref22">39</xref>
            ] and \Rands Adjusted Index" [
            <xref ref-type="bibr" rid="ref23">40</xref>
            ], etc.
Therefore, accurate software for data mining need to be optimized both with
respect to number and execution-time of integrated metrics.
2.2
          </p>
        </sec>
        <sec id="sec-2-8-2">
          <title>Disparate data and Execution-Time: Engines for data-access</title>
          <p>
            A major challenge in big-data analytics concerns the slow performance of
databaseengines [
            <xref ref-type="bibr" rid="ref24 ref25">41,23,22,42</xref>
            ]. To exemplify, the authors of [23] asserts that there is no
sound computational framework for database-management. The work of [
            <xref ref-type="bibr" rid="ref26">43</xref>
            ]
observes how \big data analytics requires technologies to e ciently process large
quantities of data" [
            <xref ref-type="bibr" rid="ref26">43</xref>
            ]. To address the performance lag in database-engines
current approaches seeks to pre-compute queries [
            <xref ref-type="bibr" rid="ref27 ref28">17,44,45</xref>
            ], reduce RDF-dataset
through slicing [
            <xref ref-type="bibr" rid="ref29">46</xref>
            ], etc. However, a prevailing issue concerns the high time-cost
of queries: the search-engine of [
            <xref ref-type="bibr" rid="ref30">47</xref>
            ] use more than 13 minutes to evaluate a
simple query. What may be argued is that the choice of accurate data-engines
may address the performance issue. There is a large number of di erent
dataengines for high-performance querying of semantic data [48,49,50,?,?]. One of
these is the \Sesame" data-engine [
            <xref ref-type="bibr" rid="ref36">53</xref>
            ], a data-engine which is unable to provide
real-time query-answer-time to simple queries [
            <xref ref-type="bibr" rid="ref37">54</xref>
            ]. Our earlier work [
            <xref ref-type="bibr" rid="ref38">55</xref>
            ]
identi es how the established B-tree ([
            <xref ref-type="bibr" rid="ref24 ref39">56,41</xref>
            ]) data-structure results in a 10,000x+
performance-delay when compared to accessing data stored in an in-memory
2dsparse data-structure, as discussed in [
            <xref ref-type="bibr" rid="ref38">55</xref>
            ]. In our [
            <xref ref-type="bibr" rid="ref38">55</xref>
            ] we demonstrates how a 2d
sparse data-structure may be used as an alternative to established data-engines,
a work which observe how application of a 2d sparse data-structure outperforms
MySQL by 10,000,000x+ for important bio-medical queries.
3
          </p>
          <p>
            Method: A holistic method for knowledge discovery
In the integration of real-time user access to 30 million bio-medical relationships
we have faced the challenges described in research, as exempli ed in section 2.
From the works of others we realize that it is not feasible to follow the
established strategies. To exemplify, major e orts by [17,?,19] are placed on
translating data-formats into RDF. However, their approaches have not resulted in
knowledge discovery. In the uni cation of data-resources we have addressed
issues in assimilation of the graph-structured \BioPax" [16] formats and
evidenceannotations in \OBO" [
            <xref ref-type="bibr" rid="ref43">60</xref>
            ], i.e., where latter by de nition is not supported by
the \SPAR-QL" query-language. An example of erroneous name-mappings is
seen for an entity with name \HDR" asserted by [61,?] to be an exact synonym
of the \gata3" gene. In contrast the established view is that \HDR" describes a
mechanism in cells [63]. The latter example is one of many fallacies observed in
Fig. 3: Our support for knowledge-inferences in ltered data. The above gure
exemplify how our approach enable knowledge-discovery, as described in [
            <xref ref-type="bibr" rid="ref40">57</xref>
            ]. Each of the
sub- gures represent distinct data-sets capturing di erent hypothesis in [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. The di
erences and similarities between the sub- gures provide clues of how guinea pigs develop.
Importantly, the above separation between entities re ect the ndings in [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ], hence our
interactive data-mining approach provide support for accurate and fast data- ltering
of user-de ned data-sets.
multi-origin databases, hence integration of data need to take care when using
assertions from unreliable sources.
          </p>
          <p>
            A di erent aspect concerns the execution-time of user-queries, exempli ed
in Fig. 2. To address the high time-cost of translating external data-bases into
RDF, and searching RDF data-stores, we have designed a data-engine which
accepts semantic relationships. When measuring the response-time of queries we
observe how our new data-engine address issues in execution-time, as described
in our [
            <xref ref-type="bibr" rid="ref38">55</xref>
            ]. The/Our semantic data-engine address issues such as:
1. Disparate data: integration of evidence annotations, hence less relationships
to investigate during evidence-centered user-queries;
2. Execution-time: memory-cache aware data-searches, e ective use of SSE [64],
memory-tiling [65], etc;
3. User searches: pre-computation of statistics, and ranks, for database-vertices
enable accurate suggestion when users type name of entities, as exempli ed
in Fig. 6.
          </p>
          <p>The above described strategy exemplify approaches which reduce
searchtime without introducing erroneous heuristics, e.g., in contrast to [18]. Fig. 5
presents a summary of the approaches undertaken to optimize the performance
aspects of bio-medical knowledge discovery, hence a holistic method for
datamining. To exemplify, we from Fig. 5 observe how the holistic approach address
issues in disparate data through a combination of manually curated rules (to
address quality issues in data-resources, e.g., the \HDR" use-case), application
of clustering to unify entities both with respect to their database-resource (e.g.,
\uniprot"), etc.</p>
          <p>1,000</p>
          <p>800
]
s
[
e 600
m
i
T
r
e
s
U 400
200
0</p>
          <p>Time-cost of pariwise simliarity-metrics in established software-libraries</p>
          <p>Kendall: “cluster-C”
Kendall: “hpLysis”
100
200
300
400
500
600</p>
          <p>Time to compute pairwise simliarity for a squared matrix
The holistic approach for drug-discovery, introduced in this paper, is constructed
to relate domain-experts to accurate interactions minted from big data. From
below sub-sections we assert that the approach manages to correctly address the
issues described in section 2.</p>
        </sec>
        <sec id="sec-2-8-3">
          <title>Disparate data: Data-access in the bio-medical domain</title>
          <p>
            There are more than 220 di erent knowledge sources in the bio-medical domain
[66]. Correctness and usability are characteristics which describe the most
popular tools for knowledge integration. An example tool is cPath, written by [66],
which is built around a MySQL databas1e [67]. From the performance
measurements of data-structures in sub-section 4.4 we observe how semantic searches
through MySQL results in a 10,000,000x+ performance-delay. The best tools
provide access to data which has been manually curated by eld experts, such
as the Reactome tool [26] or the BioGrid tool [25]. The back-bone of the tools
is often the Gene Ontology (GO) [68], which is used to de ne the lexicographic
order of the genes and proteins. Translating compartmentalized knowledge into
an ontology for reasoning, such as the RDF format, is seen in [
            <xref ref-type="bibr" rid="ref28">24,45</xref>
            ]. The
problem with both approaches is the performance and quality issues in knowledge
discovery.
4.2
          </p>
        </sec>
        <sec id="sec-2-8-4">
          <title>Execution-Time: Relationship between implementation, execution-time, and their in uence</title>
          <p>
            In data-mining the execution-time of software may render high-quality analytical
approaches useless, e.g., as inferred for large data-sets in Fig. 4. The application
of established implementation-strategies results in under-performing code due
to the challenges of compilers to identify strategies for performance tuning (as
it otherwise would not have been a time-di erence between di erent software
implementations). In below we list a subset of observations from the holistic
optimization of approaches for data-mining:
1. Search-time: the test-cases listed in sub-section 4.4 relate the time-cost of
semantic searches in our novel database-engine to the established use of
Btrees ([
            <xref ref-type="bibr" rid="ref24 ref39">56,41</xref>
            ]), observing a time-di erence of 10,000x+;
2. Data-mining: Fig. 4 compare the time-cost of strategies for computing \Kendall's
Tau". When our approach is compared to the popular \cluster-C" software
[20], a time-di erence of 100,000x+ is observed;
3. Software complexity: Fig. 5 examplify how accurate and fast user-searches
involve steps in data-curation, analysis of semantic similarity, and
construction of user-interfaces.
          </p>
          <p>The above observations indicates that the application of low-level optimization
strategies is a central part in e orts for mining of big-data, re ecting observations
in section 2, hence the importance of an holistic optimization strategy.
4.3</p>
        </sec>
        <sec id="sec-2-8-5">
          <title>User searches: application centered perspective</title>
          <p>A sound interaction between software-tools and human domain-experts is seen
as an essential part of knowledge discovery. In below we exemplify a subset of
the strategies we have applied in the holistic approach (Fig. 5):
1. Semantic user-interactions: Fig. 6 exemplify how users are provided with
support for both semantic queries (topmost sub- gures), interactive exploration
(sub- gures in the middle) and signature queries (bottom-right sub- gure);
2. Testing hypothesis: in Fig. 3 we observe how a combination of the \MINE"
metric [69] and our web-based framework for data-mining facilitate
knowledge discovery on ltered query-subsets.
Pool of KBs
and DBs
Translating
id’s into</p>
          <p>Normalized
column-based N(7)
Standardization (sanitation)</p>
          <p>Input to</p>
          <p>Column-based</p>
          <p>N(7)
Translating
axioms into</p>
          <p>KT-axiom-spec
KEnfSoTewtrrcaultenicvdsteuglareDeteasintato uSnyifnicoantyiomn</p>
          <p>Knit ing-Tools -CORE(KT)</p>
        </sec>
      </sec>
      <sec id="sec-2-9">
        <title>RSeeamsoannitnicg Re-Nenegtwinoerekring</title>
        <p>Legends:</p>
        <p>Disparate data
Execution-time
User-searches
- &gt; Update
KB-mappings.
&lt;- get
normalized
KB-identity
Mapping
between equal
KB-identifiers
Clustering into
disjoint
KB-entries</p>
      </sec>
      <sec id="sec-2-10">
        <title>ComPpaletxerSnesarch Data-engine</title>
        <p>Knit ing-Tools -SKB
quReeryala-tcimceess CoCmopraerilsaotinonand</p>
        <p>Formats: OBO,
Bel , CSV/TSV</p>
        <p>Local and
client-server
access
GUI-query-API</p>
        <p>Visualisation
API</p>
        <p>Similarity</p>
        <p>Clustering</p>
        <p>Feature Selection and
Group Assignment</p>
        <p>Test hypothesis
www</p>
        <p>Terminal</p>
        <p>Selection Criterias</p>
        <p>Fig. 5: An holistic approach for data-mining. The above gure depicts a collection of
labelled boxes, such as Pools of KBs and DBs and Visualization. The legend-text
(topConstSruKcBtion of right) describe the classi cation of the di erent background rectangles, as discussed
in section 3. An example of a classi cation concerns the process of handling disparate
data, where tasks for data-parsing, entity optimization, and format-uni cation, are
combined into an automated approach. The size of the background-boxes re ect their
computational complexity, as discussed in Fig. 2. The uniqueness of our approach
concerns how we relate the existing strategies into one uni ed model, thereby avoiding
overheads associated with generalised approaches (such as RDF centered
integrationstrategies). To exemplify, when approaches such as [19] apply Standardization they use
standardized rules for all of the integrated data-resources, hence entities from di erent
data-resources are syntactically correct while semantically inaccurate.</p>
        <sec id="sec-2-10-1">
          <title>Reproducibility: interfaces approaches for data-mining</title>
          <p>The results, summarized
of our software, as listed
in
in
this paper, may
below:
to
validate
and
elaborate
our
be
re-produced
through
application
4.4
1.
2.
3.</p>
          <p>Semantics and data:
MINE data-mining:
MINE high-perform
downloads/;
Software for</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Use-cases: discovery</title>
      <p>The holistic approach represented in this paper manages to address issues in
bigdata analysis. The term big data depends on the complexity of data, algorithms,
and use-cases to be evaluated, e.g., where [70] asserts that a study of 857 proteins
implies a large-scale analysis. This section therefore seeks to address strategies
for:
http://www.knittingTools.org/;
http://www.knittingTools.org/gui_lib_mine.cgi;
ance software: https://bitbucket.org/oekseth/mine- data- analysis/
data-analysis: https://bitbucket.org/oekseth/hplysis- cluster- analysis- software.
how the holistic
approach improves
drug</p>
      <p>Formats: OBO,
Bel , CSV/TSV</p>
      <p>Local and
client-server
access
- A weakness of our approach, is the complexity of it, a complexity whic
maintain several (??) isolated parts.
- A strength of our layout, is the (relative) disjointness of the component
makes it easier to verify cor ectness (and to some automatical y extend fi
in naming of entities).</p>
      <p>Concatenation of
reasoning-res
ults</p>
      <p>Clustering</p>
      <p>Proximity
Centrality
Interactive
www.kn
Filtering and
Analysis
Local and
client-server
access
graph-based
visualization
(??) ANN
combined with
PCA Analysis
GUI-query-API
signature
searches
Paral el-plots</p>
      <p>Heatmaps
Dendograms</p>
      <p>Legend-plot
Circle-plots</p>
      <p>Dynamic fetching</p>
      <p>of
node-information
Visualization through www.knittingTools.org
Fig. 6: A user-interface for semantic queries. The above sub- gures depict the
userinterface for semantic searches, a user-interface designed for domain-experts without
interest in programming. The two uppermost sub- gures depicts the query-form for
submitting conditional questions. The drop-down menu observed in the top-left gure
exempli es the support for auto-complete. In the middle-right sub- gure the
queryresults is seen. To handle the frequent issue of 1000+ identi ed relationships, the table
include lter-options. When a user lters a subset of the query-result, it is visualized
in the middle-left sub- gure.
1. Disparate data: algorithms and data which may be accurately queried;
2. Execution-time: why the enabled performance-increase is important in drug
discovery;
3. User searches: how domain-experts may identify accurate prediction from
complex data.
5.1</p>
      <sec id="sec-3-1">
        <title>Disparate data and user-searches</title>
        <p>The \Knitting-Tools" web-server includes a number of pre-computed use-cases
(http://knittingtools.org/examples.cgi). The results have been manually
investigated and veri ed. Accuracy of predictions depends on data (Fig. 1) and
accessibility (Fig. 6). Below use-cases provides a brief introduction to how
complex searches may be applied on big-data.</p>
        <p>Use-case(1): What is known for \notch2"? (http://knittingtools.org/
query.cgi?queryID=get_allRelations_forA_vertex). The use-case illustrates
an exploratory search to identify all relationships, synonyms and provenance
(e.g., the set of databases) describing a vertex of interest, e.g., the \notch2".
Ampli es the use of basic search functionality to fetch relationships in
KnittingTools, both for visual evaluation (Fig. 6), and as a prior data-gathering step
before application of software for pattern identi cation (sub-section 5.2 and
subsection 5.3).</p>
        <p>Use-case(2): What is known for \cdk4", and how was this known? (http://
knittingtools.org/query.cgi?queryID=get_allRelations_forA_vertex_evidence_
and_synonyms). Extends use-case(1) with logic to fetch the synonymous vertices
(for each vertex in the set of identi ed relations) and the provenance for each
relationships. Provides insight into why the identi ed relationships were predicted,
i.e., their provenance.</p>
        <p>question(3): Identify the regulations associated to the important event of
apoptosis (i.e., 'controlled cell death'). (http://knittingtools.org/query.cgi?
queryID=intro_basic_bio_2). The query identi es relationships associated to
pathways and regulations for chemical entities, proteins, genes, and pathways.
In the result provenance is associated to each relationship, a provenance which
becomes visible when clicking the green-plus button in the result-table (Fig. 6).
5.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Pattern identi cation and usability</title>
        <p>
          The MINE software combine an highly accurate algorithm for pattern
identi cation [69] with a web-interface for interactive data-exploration (Fig. 3).
To evaluate the applicability of the MINE web-based software (http://www.
knittingTools.org/gui_lib_mine.cgi) the data-sets of [71], [72], and [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] are
evaluated. While the data-set by [
          <xref ref-type="bibr" rid="ref1">1,71</xref>
          ] provide explanation factors for growth of
guinea pigs, [72] analysis the variation in the guinea pigs goat-spots. The
conclusions presented by the authors are supported through application of the MINE
web-interface.
5.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Execution-Time and knowledge discovery</title>
        <p>For large data-sets the prediction accuracy relates directly to the execution-time
of data-mining software, i.e., a users are otherwise needed to explore smaller
data-samples and test fewer number of hypothesis. The below paragraphs
exempli es how our proposed approach increase accuracy of large-scale data-analysis.</p>
        <p>Application(1). Large-scale ontology engineering (https://bitbucket.org/
oekseth/hplysis-ontology/). We have developed a new hpLysis-onto software
for high-performance engineering of bio-medical ontology. Ontologies are used in
a large number of application, e.g., to identify similarities of gene products from
experimental outcomes [73] The hpLysis-onto software address performance
issues in computation of transitive closures and transitive reductions, an issue
hampering analysis of large and complex data-sets. For the task to compute
transitive closures for all vertices, the software of [74] consumes more than one
day on the 24 MB \Gene Ontology" [68]. In contrast, the hpLysis-onto software
manages to answer the latter query in less than one second, hence a signi cant
improvement in performance.</p>
        <p>
          Application(2). Large scale Semantic similarity (https://bitbucket.
org/oekseth/hplysis-cluster-analysis-software). The hpLysis software is
updated with a new high-performance library for computation of 20+
semantic similarity metrics. \Generally speaking, semantic similarity measures involve
the GO tree topology, information content of GO terms, or a combination of
both" [75]. The software proposed by [75] takes several hours to complete. The
hpLysis-semantic software improves the performance of established software
approaches by 1000x+, i.e., without reducing the prediction accuracy. The latter is
enabled through increased utilization of computer memory hardware. Semantic
similarity-metrics are used to identify important traits in data-sets [76], e.g., to
(1) relate hypothetical assumptions to gene-expression-levels [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] and (2) with
respect to \Word Sense Disambiguation" (WSD) for automated analysis of
textcorpuses [77].
        </p>
        <p>
          Application(3). Many-dimensional data-analysis ([
          <xref ref-type="bibr" rid="ref19">36</xref>
          ]). The hpLysis
software provides an API for high-performance computation of 20+ cluster-algorithms,
320+ pairwise similarity-metrics, 10+ metrics for string-similarity, and 20+
metrics for cluster-validity. The work enables a performance-improvement of
600x+ for pairwise similarity-metrics such as \Canberra" and \Cosine", while
100,000x+ performance-improvement when compared to \Kendall's Tau" (Fig.
4). In large-scale data-analysis the execution-time severely hamper the types of
relationship which may be explored, e.g., when analysing gene-expressions
datasets for possible interactions, when using pathway-relationships (sub-section 5.1),
application of ontology annotations for similarity-assessment, mining of
bibliometric data-bases (e.g., in [78]), etc. Therefore, hpLysis improves both accuracy
of data-generation and analysis of user-de ned data (Fig. 6).
6
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conlusion and Future Work</title>
      <p>We have presented both a method, a database, and 10+ software, for
datamining. This paper argue that the holistic approach, which captures an
ensemble of approaches and high performance software, manages to overcome the
current hurdles in big-data drug-discovery. Fig. 6 exemplify how domain-experts
may interact with our real-time support for querying 30+ million bio-medical
relationships. In order to evaluate the quality of the approach we investigate
the number of accurate and unique relationships identi ed in our approach,
exempli ed in Fig. 1. The 1000+ novel candidate interactions which are identi ed
highlight the ability of our approach to automatically identify relationship which
are not known in literature. Through an optimized data-engine the relationships
are accessible for users in real-time, exempli ed in Fig. 3.</p>
      <p>In this paper we have described an approach to unify our semantic
interface (www.knittingTools.org) with our high-performance software application
(e.g., https://bitbucket.org/oekseth/hplysis-cluster-analysis-software).
Through concrete use-cases we have exempli ed how the approach address
issues in disparate data, execution-time, and user searches, ie, parameters which
are critical in discovery of knowledge. From the examples we observe how our
method and 10+ novel software approaches address issues in big-data
drugdiscovery. Therefore, we assert that our novel holistic approach may in uence
strategies for mining of big-data.
We plan to address the weakness of the user-interfaces and the unknown
quality of our knowledge inferences. In order to improve our user-interfaces we are
now initiating e orts in usability testing for di erent target groups. Similarily,
we have initiated e orts to evaluate the drug-impact of our putative knowledge
discoveries. Both of the issues require year-long lab-experiments, hence the
importance of quality and performance enabled through our novel method and
software.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>The authors would like to thank MD K.I. Ekseth at UIO, Dr. O.V. Solberg
at SINTEF, Dr. S.A. Aase at GE Healthcare, MD B.H. Helleberg at NTNU{
medical, Dr. Y. Dahl, Dr. T. Aalberg, Dr. J.C. Meyer, and K.T. Dragland at
NTNU, and the High Performance Computing Group at NTNU for their support.
11. Herper, M.: The Truly Staggering Cost Of Inventing New
Drugs. https://www.forbes.com/sites/matthewherper/2012/02/10/
the-truly-staggering-cost-of-inventing-new-drugs/#1a0129244a94 (2012)
12. DiMasi, J.A., Grabowski, H.G., Hansen, R.W.: The cost of drug development. New</p>
      <p>England Journal of Medicine 372(20), 1972{1972 (2015)
13. Association, C.B.R., et al.: Fact sheet: New drug development process. FDA Special</p>
      <p>Consumer Report
14. Thomas, K.: The price of health: the cost of developing new medicines.
https://www.theguardian.com/healthcare-network/2016/mar/30/
new-drugs-development-costs-pharma (2016)
15. Paul, S.M., Mytelka, D.S., Dunwiddie, C.T., Persinger, C.C., Munos, B.H.,
Lindborg, S.R., Schacht, A.L.: How to improve r&amp;d productivity: the pharmaceutical
industry's grand challenge. Nature reviews. Drug discovery 9(3), 203 (2010)
16. Demir, E., Cary, M.P., Paley, S., Fukuda, K., Lemer, C., Vastrik, I., Wu, G.,
D'Eustachio, P., Schaefer, C., Luciano, J., et al.: The biopax community standard
for pathway data sharing. Nature biotechnology 28(9), 935{942 (2010)
17. Blonde, W.: Metarel, an ontology facilitating advanced querying of biomedical
knowledge. PhD thesis, Department of Mathematical Modelling, Statistics and
Bioinformatics, Ghent University, Ghent, Belgium (2012)
18. Antezana, E., Blond, W., Egan~a, M., Rutherford, A., Stevens, R., De Baets, B.,
Mironov, V., Kuiper, M.: Biogateway: a semantic systems biology tool for the life
sciences. BMC bioinformatics 10(10), 11 (2009)
19. Belleau, F., Nolin, M.-A., Tourigny, N., Rigault, P., Morissette, J.: Bio2rdf:
towards a mashup to build bioinformatics knowledge systems. Journal of biomedical
informatics 41(5), 706{716 (2008)
20. de Hoon, M.J., Imoto, S., Nolan, J., Miyano, S.: Open source clustering software.</p>
      <p>Bioinformatics 20(9), 1453{1454 (2004)
21. Hucka, M., Finney, A., Sauro, H.M., Bolouri, H., Doyle, J.C., Kitano, H.: The erato
systems biology workbench: enabling interaction and exchange between software
tools for computational biology (2002)
22. Butcher, E.C., Berg, E.L., Kunkel, E.J.: Systems biology in drug discovery. Nature
biotechnology 22(10), 1253 (2004)
23. Chowdhury, S., Sarkar, R.R.: Comparison of human cell signaling pathway
databasesevolution, drawbacks and challenges. Database 2015 (2015)
24. Beisswanger, E., Lee, V., Kim, J.-J., Rebholz-Schuhmann, D., Splendiani, A.,
Dameron, O., Schulz, S., Hahn, U., et al.: Gene regulation ontology (gro):
design principles and use cases. Studies in health technology and informatics 136, 9
(2008)
25. Stark, C., Breitkreutz, B.-J., Chatr-Aryamontri, A., Boucher, L., Oughtred, R.,
Livstone, M.S., Nixon, J., Van Auken, K., Wang, X., Shi, X., et al.: The biogrid
interaction database: 2011 update. Nucleic acids research 39(suppl 1), 698{704
(2011)
26. Croft, D., OKelly, G., Wu, G., Haw, R., Gillespie, M., Matthews, L., Caudy, M.,
Garapati, P., Gopinath, G., Jassal, B., et al.: Reactome: a database of reactions,
pathways and biological processes. Nucleic acids research 39(suppl 1), 691{697
(2011)
27. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O.,
Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al.: Scikit-learn: Machine
learning in python. Journal of Machine Learning Research 12(Oct), 2825{2830
(2011)
63. Davis, L., Maizels, N.: Homology-directed repair of dna nicks via pathways distinct
from canonical double-strand break repair. Proceedings of the National Academy
of Sciences 111(10), 924{932 (2014)
64. Intel: SSE computer-hardware-low-level parallelism. https://software.intel.</p>
      <p>com/sites/landingpage/IntrinsicsGuide/. Online; accessed 06. June 2017
65. Drepper, U.: What every programmer should know about memory. Red Hat, Inc
11, 2007 (2007)
66. Cerami, E., Bader, G., Gross, B., Sander, C.: cpath: open source software for
collecting, storing, and querying biological pathways. BMC bioinformatics 7(1),
497 (2006)
67. MySQL: MySQL database engine. https://www.mysql.com/ (2017)
68. Ashburner, M., Ball, C.A., Blake, J.A., Botstein, D., Butler, H., Cherry, J.M.,
Davis, A.P., Dolinski, K., Dwight, S.S., Eppig, J.T., et al.: Gene ontology: tool for
the uni cation of biology. Nature genetics 25(1), 25{29 (2000)
69. Reshef, D.N., Reshef, Y.A., Finucane, H.K., Grossman, S.R., McVean, G.,
Turnbaugh, P.J., Lander, E.S., Mitzenmacher, M., Sabeti, P.C.: Detecting novel
associations in large data sets. science 334(6062), 1518{1524 (2011)
70. Butland, G., Peregr n-Alvarez, J.M., Li, J., Yang, W., Yang, X., Canadien, V.,
Starostine, A., Richards, D., Beattie, B., Krogan, N., et al.: Interaction network
containing conserved and essential protein complexes in escherichia coli. Nature
433(7025), 531{537 (2005)
71. Peaker, M., Taylor, E.: Sex ratio and litter size in the guinea-pig. Journal of
reproduction and fertility 108(1), 63{67 (1996)
72. Wright, S., Chase, H.B.: On the genetics of the spotted pattern of the guinea pig.</p>
      <p>Genetics 21(6), 758 (1936)
73. Pesquita, C., Faria, D., Falcao, A.O., Lord, P., Couto, F.M.: Semantic similarity
in biomedical ontologies. PLoS computational biology 5(7), 1000443 (2009)
74. Antezana, E., Egana, M., Baets, B., Kuiper, M., Mironov, V.: Onto-perl: An api for
supporting the development and analysis of bio-ontologies. Bioinformatics (2008)
75. Ehsani, R., Drabl s, F.: Topoicsim: a new semantic similarity measure based on
gene ontology. BMC bioinformatics 17(1), 296 (2016)
76. Rada, R., Mili, H., Bicknell, E., Blettner, M.: Development and application of
a metric on semantic nets. IEEE transactions on systems, man, and cybernetics
19(1), 17{30 (1989)
77. McInnes, B.T., Pedersen, T.: Evaluating measures of semantic similarity and
relatedness to disambiguate terms in biomedical text. Journal of biomedical informatics
46(6), 1116{1124 (2013)
78. Aalberg, T., Zumer, M.: The value of marc data, or, challenges of frbrisation.</p>
      <p>Journal of Documentation 69(6), 851{872 (2013)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>McPhee</surname>
            ,
            <given-names>H.C.</given-names>
          </string-name>
          , et al.:
          <article-title>Genetic growth di erentiation in guinea pigs</article-title>
          .
          <source>Technical report</source>
          , United States Department of Agriculture, Economic Research Service (
          <year>1931</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Stephenson</surname>
            ,
            <given-names>D.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lemere</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Selkoe</surname>
            ,
            <given-names>D.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clemens</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Cytosolic phospholipase a 2 (cpla 2) immunoreactivity is elevated in alzheimer's disease brain</article-title>
          .
          <source>Neurobiology of disease 3(1)</source>
          ,
          <volume>51</volume>
          {
          <fpage>63</fpage>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Feuerherm</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johansen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Rheumatoid arthritis treatment</article-title>
          .
          <source>US Patent App. 13/783</source>
          ,088 (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          :
          <article-title>Mining online heterogeneous healthcare networks for drug repositioning</article-title>
          .
          <source>In: Healthcare Informatics (ICHI)</source>
          ,
          <source>2016 IEEE International Conference On</source>
          , pp.
          <volume>106</volume>
          {
          <issue>112</issue>
          (
          <year>2016</year>
          ). IEEE
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. for Biotechnology Information,
          <string-name>
            <surname>N.C.</surname>
          </string-name>
          :
          <article-title>PubMed data-base for biomedical litterature</article-title>
          . https://www.ncbi.nlm.nih.gov/pubmed/ (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Consortium</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          , et al.:
          <article-title>Uniprot: the universal protein knowledgebase</article-title>
          .
          <source>Nucleic acids research</source>
          <volume>45</volume>
          (
          <issue>D1</issue>
          ),
          <volume>158</volume>
          {
          <fpage>169</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Economist</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>How science goes wrong: Trouble at the lab</article-title>
          .
          <source>The Economist</source>
          <volume>409</volume>
          (
          <issue>8858</issue>
          ),
          <volume>21</volume>
          {
          <fpage>24</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cuatrecasas</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Drug discovery in jeopardy</article-title>
          .
          <source>Journal of Clinical Investigation</source>
          <volume>116</volume>
          (
          <issue>11</issue>
          ),
          <volume>2837</volume>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>DiMasi</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grabowski</surname>
            ,
            <given-names>H.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>R.W.:</given-names>
          </string-name>
          <article-title>Innovation in the pharmaceutical industry: new estimates of r&amp;d costs</article-title>
          .
          <source>Journal of health economics 47</source>
          ,
          <volume>20</volume>
          {
          <fpage>33</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Al-Huniti</surname>
          </string-name>
          , N.:
          <article-title>Quantitative Decision-Makingin Drug Development</article-title>
          . http://www. phuse.eu/download.aspx?type=cms&amp;docID=
          <volume>5334</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          28.
          <string-name>
            <surname>Clustergrammer</surname>
          </string-name>
          , J.:
          <article-title>Clustergrammer heatmap visualization</article-title>
          . http://amp.pharm. mssm.edu/clustergrammer/
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          29.
          <string-name>
            <surname>Metsalu</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vilo</surname>
          </string-name>
          , J.:
          <article-title>Clustvis: a web tool for visualizing clustering of multivariate data using principal component analysis and heatmap</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>43</volume>
          (
          <issue>W1</issue>
          ),
          <volume>566</volume>
          {
          <fpage>570</fpage>
          (
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .1093/nar/gkv468
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          30.
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>E.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dannenfelser</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>N.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maayan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Network2canvas: network visualization on a canvas with enrichment analysis</article-title>
          .
          <source>Bioinformatics</source>
          <volume>29</volume>
          (
          <issue>15</issue>
          ),
          <year>1872</year>
          {
          <year>1878</year>
          (
          <year>2013</year>
          ). doi:
          <volume>10</volume>
          .1093/bioinformatics/btt319
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          31.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>Z.:</given-names>
          </string-name>
          <article-title>A new algorithm to optimize maximal information coe cient</article-title>
          .
          <source>PloS one 11(6)</source>
          ,
          <volume>0157567</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          32.
          <string-name>
            <surname>Mekkat</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Natarajan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
          </string-name>
          , W.-C.,
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Performance characterization of data mining benchmarks</article-title>
          .
          <source>In: Proceedings of the 2010 Workshop on Interaction Between Compilers and Computer Architecture</source>
          , p.
          <volume>11</volume>
          (
          <year>2010</year>
          ). ACM
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          33.
          <string-name>
            <surname>Andrade</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramos</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madeira</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sachetto</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rocha</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Gdbscan: A gpu accelerated algorithm for density-based clustering</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>18</volume>
          ,
          <issue>369</issue>
          {
          <fpage>378</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          34.
          <string-name>
            <surname>Ester</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kriegel</surname>
            ,
            <given-names>H.-P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sander</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , et al.:
          <article-title>A density-based algorithm for discovering clusters in large spatial databases with noise</article-title>
          .
          <source>In: Kdd</source>
          , vol.
          <volume>96</volume>
          , pp.
          <volume>226</volume>
          {
          <issue>231</issue>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          35.
          <string-name>
            <surname>Ekseth</surname>
            ,
            <given-names>O.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hvasshovd</surname>
          </string-name>
          , S.-O.:
          <article-title>How an optimized DB-SCAN implementation reduce execution-time and memory-requirements for large data-sets. Accepted for publication (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          36.
          <string-name>
            <surname>Ole Kristian</surname>
          </string-name>
          <article-title>Ekseth: hpLysis: a high-performance software-library for big-data machine-learning</article-title>
          . https://bitbucket.org/oekseth/ hplysis-cluster
          <article-title>-analysis-software/</article-title>
          .
          <source>Online; accessed 06</source>
          . June 2017
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          37.
          <string-name>
            <surname>Rousseeuw</surname>
            ,
            <given-names>P.J.:</given-names>
          </string-name>
          <article-title>Silhouettes: a graphical aid to the interpretation and validation of cluster analysis</article-title>
          .
          <source>Journal of computational and applied mathematics 20</source>
          ,
          <volume>53</volume>
          {
          <fpage>65</fpage>
          (
          <year>1987</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          38.
          <string-name>
            <surname>Lloyd</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Least squares quantization in pcm</article-title>
          .
          <source>IEEE transactions on information theory 28</source>
          (
          <issue>2</issue>
          ),
          <volume>129</volume>
          {
          <fpage>137</fpage>
          (
          <year>1982</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          39.
          <string-name>
            <surname>Rand</surname>
            ,
            <given-names>W.M.:</given-names>
          </string-name>
          <article-title>Objective criteria for the evaluation of clustering methods</article-title>
          .
          <source>Journal of the American Statistical association</source>
          <volume>66</volume>
          (
          <issue>336</issue>
          ),
          <volume>846</volume>
          {
          <fpage>850</fpage>
          (
          <year>1971</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          40.
          <string-name>
            <surname>Yeung</surname>
            ,
            <given-names>K.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruzzo</surname>
            ,
            <given-names>W.L.</given-names>
          </string-name>
          :
          <article-title>Details of the adjusted rand index and clustering algorithms, supplement to the paper an empirical study on principal component analysis for clustering gene expression data</article-title>
          .
          <source>Bioinformatics</source>
          <volume>17</volume>
          (
          <issue>9</issue>
          ),
          <volume>763</volume>
          {
          <fpage>774</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          41.
          <string-name>
            <surname>Jagadish</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olken</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Database management for life sciences research</article-title>
          .
          <source>ACM SIGMOD Record</source>
          <volume>33</volume>
          (
          <issue>2</issue>
          ),
          <volume>15</volume>
          {
          <fpage>20</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          42.
          <string-name>
            <surname>Eltabakh</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ouzzani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aref</surname>
            ,
            <given-names>W.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elmagarmid</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laura-Silva</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arshad</surname>
            ,
            <given-names>M.U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salt</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baxter</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Managing biological data using bdbms</article-title>
          .
          <source>In: Data Engineering</source>
          ,
          <year>2008</year>
          .
          <article-title>ICDE 2008</article-title>
          . IEEE 24th International Conference On, pp.
          <volume>1600</volume>
          {
          <issue>1603</issue>
          (
          <year>2008</year>
          ). IEEE
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          43.
          <string-name>
            <surname>Lau</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang-Turner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karacapilidis</surname>
          </string-name>
          , N.:
          <article-title>Requirements for big data analytics supporting decision making: A sensemaking perspective</article-title>
          .
          <source>In: Mastering DataIntensive Collaboration and Decision Making</source>
          , pp.
          <volume>49</volume>
          {
          <fpage>70</fpage>
          . Springer, ??? (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          44.
          <string-name>
            <surname>Blond</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antezana</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mironov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baets</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          :
          <article-title>Using the relation ontology Metarel for modelling Linked Data as multi-digraphs (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          45.
          <string-name>
            <surname>Blond</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mironov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antezana</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venkatesan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baets</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Reasoning with bio-ontologies: using relational closure rules to enable practical querying</article-title>
          .
          <source>Oxford Bioinformatics</source>
          <volume>27</volume>
          , 1562{
          <fpage>1568</fpage>
          (
          <year>2011</year>
          ). doi:
          <volume>10</volume>
          .1093/bioinformatics/btr164
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          46.
          <string-name>
            <surname>Marx</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shekarpour</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soru</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brasoveanu</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saleem</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baron</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weichselbraun</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngomo</surname>
            ,
            <given-names>A.-C.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Torpedo: Improving the state-of-the-art rdf dataset slicing</article-title>
          .
          <source>In: Semantic Computing (ICSC)</source>
          ,
          <source>2017 IEEE 11th International Conference On</source>
          , pp.
          <volume>149</volume>
          {
          <issue>156</issue>
          (
          <year>2017</year>
          ). IEEE
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          47.
          <string-name>
            <surname>Papanikolaou</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavlopoulos</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pa lis</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Theodosiou</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schneider</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Satagopam</surname>
            ,
            <given-names>V.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ouzounis</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eliopoulos</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Promponas</surname>
            ,
            <given-names>V.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iliopoulos</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Biotextquest+: a knowledge integration platform for literature mining and concept discovery</article-title>
          .
          <source>Bioinformatics</source>
          <volume>30</volume>
          (
          <issue>22</issue>
          ),
          <volume>3249</volume>
          {
          <fpage>3256</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          48.
          <string-name>
            <surname>Kolpakov</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poroikov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharipov</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kondrakhin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zakharov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagunin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milanesi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Cyclonetan integrated database on cell cycle regulation and carcinogenesis</article-title>
          .
          <source>Nucleic acids research 35(suppl 1)</source>
          ,
          <volume>550</volume>
          {
          <fpage>556</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          49.
          <string-name>
            <surname>Demir</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Babur</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodchenkov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aksoy</surname>
            ,
            <given-names>B.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fukuda</surname>
            ,
            <given-names>K.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gross</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Sumer,
          <string-name>
            <given-names>O.S.</given-names>
            ,
            <surname>Bader</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.D.</given-names>
            ,
            <surname>Sander</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Using biological pathway data with paxtools</article-title>
          .
          <source>PLoS computational biology 9</source>
          (
          <issue>9</issue>
          ),
          <volume>1003194</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          50.
          <string-name>
            <surname>Masseroli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinoli</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venco</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaitoua</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jalili</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palluzzi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Genometric query language: a novel approach to large-scale genomic data management</article-title>
          .
          <source>Bioinformatics</source>
          <volume>31</volume>
          (
          <issue>12</issue>
          ),
          <year>1881</year>
          {
          <year>1888</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          51.
          <string-name>
            <surname>Mironov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seethappan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blond</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antezana</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Splendiani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Gauging triple stores with actual biological data</article-title>
          .
          <source>BMC bioinformatics 13(1)</source>
          ,
          <volume>3</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          52.
          <string-name>
            <surname>Wylot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cudre-Mauroux</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Diplocloud:
          <article-title>E cient and scalable management of rdf data in the cloud</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>28</volume>
          (
          <issue>3</issue>
          ),
          <volume>659</volume>
          {
          <fpage>674</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          53.
          <string-name>
            <surname>Huysmans</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Richelle</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wodak</surname>
            ,
            <given-names>S.J.:</given-names>
          </string-name>
          <article-title>Sesam: a relational database for structure and sequence of macromolecules</article-title>
          .
          <source>Proteins: Structure, Function, and Bioinformatics</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <volume>59</volume>
          {
          <fpage>76</fpage>
          (
          <year>1991</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          54.
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , He in, J.:
          <article-title>Lubm: A benchmark for owl knowledge base systems</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>3</volume>
          (
          <issue>2</issue>
          ),
          <volume>158</volume>
          {
          <fpage>182</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          55.
          <string-name>
            <surname>Ekseth</surname>
            ,
            <given-names>O.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hvasshovd</surname>
          </string-name>
          , S.-O.:
          <article-title>hpLysis database-engine: A new data-scheme for fast semantic queries in biomedical databases. Under review: Provides details of the in-memory data-engine: contact oekseth@gmail.com for the paper</article-title>
          . (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          56.
          <string-name>
            <surname>Bayer</surname>
          </string-name>
          , R.:
          <article-title>Symmetric binary b-trees: Data structure and maintenance algorithms</article-title>
          .
          <source>Acta Informatica</source>
          <volume>1</volume>
          ,
          <issue>290</issue>
          {
          <fpage>306</fpage>
          (
          <year>1972</year>
          ).
          <volume>10</volume>
          .1007/BF00289509
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          57.
          <string-name>
            <surname>Ekseth</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hvasshovd</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>hpLysis MINE: A high-performance approach for computation of the accurate MINE simliarty-metric</article-title>
          . http://www.knittingtools.org/ gui_lib_mine.cgi.
          <source>Online; accessed 06</source>
          . June 2017
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          58.
          <string-name>
            <surname>Knight</surname>
            ,
            <given-names>W.R.:</given-names>
          </string-name>
          <article-title>A computer method for calculating kendall's tau with ungrouped data</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          <volume>61</volume>
          (
          <issue>314</issue>
          ),
          <volume>436</volume>
          {
          <fpage>439</fpage>
          (
          <year>1966</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          59.
          <string-name>
            <surname>Kendall</surname>
            ,
            <given-names>M.G.</given-names>
          </string-name>
          :
          <article-title>Rank correlation methods</article-title>
          . (
          <year>1948</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          60.
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ashburner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosse</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bug</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceusters</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eilbeck</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ireland</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mungall</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          , et al.:
          <article-title>The obo foundry: coordinated evolution of ontologies to support biomedical data integration</article-title>
          .
          <source>Nature</source>
          biotechnology
          <volume>25</volume>
          (
          <issue>11</issue>
          ),
          <volume>1251</volume>
          {
          <fpage>1255</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          61.
          <string-name>
            <surname>Ho</surname>
            <given-names>mann</given-names>
          </string-name>
          , R.:
          <article-title>A wiki for the life sciences where authorship matters</article-title>
          .
          <source>Nature genetics 40(9)</source>
          ,
          <volume>1047</volume>
          {
          <fpage>1051</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          62.
          <string-name>
            <surname>Chawla</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tripathi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thommesen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , L greid, A.,
          <string-name>
            <surname>Kuiper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Tfcheckpoint: a curated compendium of speci c dna-binding rna polymerase ii transcription factors</article-title>
          .
          <source>Bioinformatics</source>
          <volume>29</volume>
          (
          <issue>19</issue>
          ),
          <volume>2519</volume>
          {
          <fpage>2520</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>