<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Case-based Reasoning for the Analysis of Methylation Data in Oncology</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christopher L. Bartlett</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Isabelle Bichindaritz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Intelligent Bio Systems Laboratory, Biomedical and Health Informatics State University of New York at Oswego</institution>
          ,
          <addr-line>7060 NY-104, Oswego, NY 13126</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Researchers seek to identify biological markers which accurately di erentiate cancer subtypes and their severity from normal controls. One such biomarker, DNA methylation, has recently become more prevalent in genetic research studies in oncology. This paper proposes to apply these ndings in a study of the diagnostic accuracy of DNA methylation signatures for classifying metastasis samples. Very high classi cation performance measures were obtained from di erentially methylated positions and regions, as well as from selected gene signatures. Perfect accuracy was achieved with the top 5 feature-selected genes using three similar cases and the K-nearest neighbor classi er. This work contributes to the path toward the identi cation of biological signatures for oncology samples using case-based reasoning.</p>
      </abstract>
      <kwd-group>
        <kwd>machine learning</kwd>
        <kwd>case-based reasoning</kwd>
        <kwd>bioinformatics</kwd>
        <kwd>breast cancer</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The term epigenetics was rst introduced into modern biology by Conrad
Waddington as a means of de ning interactions between genes and their products that
result in phenotypic variations. Waddington's landscape presents a cell
becoming more di erentiated as time goes on. One of the events that can cause this
di erentiation is methylation. Methylation is a covalent attachment of a methyl
group to cytosine. Figure 1 shows the addition of this methyl group to cytosine.
Cytosine (C) is one of the four bases that construct DNA and one of only two
bases that can be methylated. While adenine can be methylated as well, cytosine
is typically the only base that's methylated in mammals. Once this methyl group
is added, it forms 5-methylcytosine where the 5 references the position on the
6-atom ring where the methyl group is added. Under the majority of
circumstances, a methyl group is added to a cytosine followed by a guanine (G) which
is known as CpG. While the methyl group is added onto the DNA, it doesn't
alter the underlying sequence but it still has profound e ects on the expression
of genes and the functionality of cellular and bodily functions. Methylation at
these CpG sites has been known to be a fairly stable epigenetic biomarker that
usually results in silencing the gene. Further, the amount of methylation can be</p>
      <p>increased (known as hypermethylation) or decreased (known as
hypomethylation) and improper maintenance of epigenetic information can lead to a variety
of human diseases.</p>
      <p>DNA methylation, has recently become more prevalent in genetic research
studies in oncology. This paper proposes to apply these ndings in a study of
the diagnostic accuracy of DNA methylation signatures for classifying metastatic
samples in breast cancer. This paper outlines the methods used to be able to
apply case-based reasoning (CBR) and instance-based learning to methylation
data, most often analyzed through statistical methods. Methylation data require
a preprocessing pipeline leading to improved analysis, as this article shows. First,
potential confounding factors such as batch e ect and potential covariates are
eliminated. Following, varied methods for the selection of subsets of methylation
probes from the 485,577 highly dimensional dataset are applied. Feature selection
methods further re ne and select appropriate probes, eventually grouping them
in genomic regions. These stages amount to case elaboration and constitute the
bulk of the work for classi cation or prediction. This paper shows that the case
elaboration mechanisms greatly improves the classi cation capability of
casebased reasoning. Following sophisticated case elaboration processes, very high
classi cation performance measures were obtained from di erentially methylated
positions and regions, as well as from selected gene signatures.</p>
      <p>Speci cally, we o er the following signi cant contributions:
1. One of the rst applications of CBR using methylation data. While
studies using gene expression data in a CBR context have been performed
previously, very few (if any), applications using methylation data have been
produced.</p>
      <sec id="sec-1-1">
        <title>2. Multi-level case elaboration and re nement which examine biolog</title>
        <p>ical and statistical di erences. Signi cantly di erent methylation levels
in the DNA, both at the microarray probe level and with a higher-order
cluster of probes that serve similar functions were utilized and compared.
Lastly, these probes are mapped to genes and ranked through a feature
selection stage that attempted to locate the smallest possible signature of
di erential methylation.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The utility of DNA methylation for the purposes of classi cation has been
recently studied to di erentiate blood samples in mental disorder subtypes [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and
cancer tumor tissue from normal tissue. This section will discuss a few such
examples before concluding with the inspiration for the project outlined in this
paper. The rst such example is a prognostic classi er developed by Dos Reis
et al., [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for well-di erentiated thyroid carcinoma (WDTC) based on 21 DNA
methylation probes that predicted a poor outcome in patients with 63%
sensitivity and 92% speci city for their internal data and 64% sensitivity and 88%
speci city for data from The Cancer Genome Atlas. Similarly, Mundbjerg et
al., [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] constructed an aggressiveness classi er from 25 methylation probes that
could determine aggressive versus non-aggressive subtypes of prostate cancer.
Testing on 496 prostate samples from tumors and adjacent-normal (AN) tissue,
they found 97.4% speci c and a 96.2% sensitivity.
      </p>
      <p>
        Hao et al., [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] determined that DNA methylation could predict cancer versus
normal tissue with accuracies above 95% in a three-cohort study of four common
cancers. Testing in breast, colon, liver and lung cancer, di erentially methylated
CpG sites were used to classify tumor versus normal tissue. Hao et al., [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] used
whole-genome methylation data from The Cancer Genome Atlas to construct a
training cohort of 1,619 tumor samples and 173 matched adjacent normal tissue
samples, and a validation cohort of 791 tumor samples and 93 matched
adjacent normal tissue samples. The correct diagnosis rate for their training set was
98.4%, which was then replicated in the validation cohort for a statistically
similar rate of 97.1%. A third, independent cohort of Chinese cancer samples (394
tumor samples and 324 matched adjacent normal tissue samples) resulted in a
correct diagnosis rate of 95.0%. Methylation patterns were also able to correctly
identify 29 of 30 colorectal cancer metastases in the liver, 32 of 34 colorectal
cancer metastases in the lung and 19 of 20 breast cancer metastases [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This
particular study promoted a positive outlook on the utility of DNA methylation
for the classi cation and characterization of cancer.
      </p>
      <p>
        Within the domain of CBR, there exist several applications using
microarray data. Anaissi, Goyal, Catchpoole, Braytee, and Kennedy [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], for example,
attempted to navigate the complexity of the highly-dimensional and imbalanced
datasets often found in microarray analysis by focusing on case retrieval. Their
framework uses a k-nearest neighbor (kNN) classi er with a weighted
featurebased similarity measure to retrieve similar patients from a case base of acute
lymphblastic leukemia. Gene expression data is employed to determine this
similarity, and the treatment and outcome is used to propose solutions. Feature
selection, dimensionality reduction, and feature weighting is used to handle the
high-dimensionality of the data and removal of irrelevant features. They utilize
oversampling to deal with the imbalanced classes. More speci cally, they use the
synthetic minority oversampling technique (SMOTE) methodology which arti
cially creates minority samples based on interpolation between members of the
original minority class. After these pre-processing stages, a new sample is given
to the kNN classi er to retrieve similar cases.
      </p>
      <p>
        Ramos-Gonzalez et al., [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] used a two-level feature selection process for gene
expression data in squamous cell carcinoma and adenocarcinoma. Their
methodology has a preliminary feature selection which uses a non-parametric
MannWhitney test to locate genes whose expression levels variation are statistically
di erentiated between subtypes. Following is a feature selection stage with
Gradient Boosted Regression Trees that further re nes the feature list into a greatly
reduced subset that still maintains a high classi cation accuracy. A
distancebased approach is used to retrieve similar cases, while additional diagnostic
information may be requested that assists in correcting the prediction.
      </p>
      <p>
        More recently, Lamy, Sekar, Guezennec, Bouaud and Seroussi [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] proposed
a CBR method that visualizes results. The CBR system was rather
straightforward, retrieving cases through a distance measure, though their specialization
was in the explainability. Qualitative attributes between cases were shown
using rainbow boxes, where labeled and colored rectangles extend through columns
that represent the cases, clearly showing what was similar or dissimilar between
cases. Quantitative attributes are provided in scatter plots that center on the
query case and accurately displays the similar cases.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Material</title>
      <p>
        Methylation data for breast cancer (BRCA, 1) was downloaded from The Cancer
Genome Atlas (TCGA, 2) using the R package TCGAbiolinksGUI [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Molecular data was ltered for only the Illumina Human Methylation 450 platform
and prepared as an RStudio object. This data pertained to 892 samples and the
485,577 probes that exist on the Illumina Human Methylation 450 beadchip.
The methylation values were then extracted. values are an estimation of the
methylation levels between 0 and 1 with 0 being completely non-methylated and
1 being completely methylated. Similarly, the BRCA clinical data was
downloaded and subset for variables of relevance. These variables were the sample
de nition (describes whether the sample is a primary solid tumor, normal
tissue, or the metastatic site), tumor stage, year of birth, tissue of origin, gender
and race. This study focuses on classifying sample as either normal or a sample
from a metastatic, stage 4 tumor.
4
4.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Methods</title>
      <sec id="sec-4-1">
        <title>Data Preprocessing</title>
        <p>Metastatic tissue samples (those pertaining to the metastasized site, not the
primary cancer site) were discarded, as well as samples from males. Year of birth
was subtracted from the current year as a measure of the subject's age, regardless
of whether the subject was alive or deceased. These subjects were then assigned
1 https://portal.gdc.cancer.gov/projects/TCGA-BRCA
2 https://www.cancer.gov/tcga
an age group with those less than 50 being in group 1, between 50 and 60 being
in group 2, 60 to 70 in group 3, 70 to 80 in group 4, 80 to 90 in group 5, and those
over 90 in group 6. The 10 stage 4 primary solid tumor samples were used to
de ne the Metastatic group (M), while 95 solid tissue normal samples de ned the
Normal group (N). Removal of probes associated with covariate variables were
then performed using the R package SVA and ComBat. The resulting dataset
after pre-processing was 120,681 sites for 105 samples.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Case-based Classi cation</title>
        <p>
          Classi cation was performed in several stages that further elaborated and
rened the cases and carried out using the Waikato Environment for Knowledge
Analysis (WEKA) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and the K-nearest neighbor algorithm. In each classi
cation step, training was rst performed through iterative removal of one sample
for testing while the other samples were retained for training. For the testing
sample, the nearest one, two or three cases were retrieved by calculating the
Euclidean distance based on the similarity of features. The classi cation label
for these cases were retrieved, with the label that was in the majority being
reused for the testing sample. Despite the class imbalance, we elected not to
use oversampling or undersampling. Oversampling the minority class can swiftly
lead to over tting, while undersampling the majority class can potentially lead
to leaving out an important instance with crucial di erences that could aid in
the identi cation of the minority class. Instead, we utilized performance
measures that adjusted for the class imbalance by calculating a balanced accuracy
(BACC, computed using the average of per-class accuracy) and the weighted
average area under the ROC curve (AUC).
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Case Elaboration and Re nement</title>
        <p>First results on all features after pre-processing proved to be unsatisfactory and
attempts were made to re ne the cases by focusing on the di erentially
methylated probes between the two groups (normal and metastatic), then on the
differentially methylated regions. Finally feature selection methods were attempted
to de ne a methylation signature on the metastatic samples.</p>
        <p>Di erentially Methylated Positions Di erentially methylated positions (DMP)
were identi ed using the Chip Analysis Methylation Pipeline (ChAMP) for R.
This package uses limma to identify statistically di erent probes between the
groups, using a Benjamin-Hochberg adjusted p-value of 0.05 for signi cance.
107,497 probes were found to be di erentially methylated, and these sites were
tested to observe di erences in classi cation performance.</p>
      </sec>
      <sec id="sec-4-4">
        <title>Di erentially Methylated Regions The DMRcate method within ChAMP</title>
        <p>was used to extract the di erentially methylated regions (DMR). Regions are
clusters of probes that serve a similar function in gene transcriptional
regulation. Cross-hybridizing probes and sex-chromosome probes were removed prior
to operation to further account for potential confounding factors such as gender.
A false-discovery rate of 0.05 and a minimum probe number of 15 were provided
as primary thresholding parameters with an adjusted p-value of 0.01 as the
signi cance threshold. Probes within the located regions were then used to build
the dataset for this stage. 788 probes were located within these regions.
4.4</p>
      </sec>
      <sec id="sec-4-5">
        <title>Feature Selection</title>
        <p>Feature selection was carried out on the dataset after initial pre-processing
measures were performed, as well as on the data after di erentially methylated
position analyses. Prior to feature selection, each probe was mapped to its associated
gene. Four algorithms in WEKA consisting of the Information Gain Attribute
Evaluation, Correlation Attribute Evaluation, SMO Classi er Attribute
Evaluation and Naive Bayes Classi er Attribute Evaluation were performed. An
ensemble was then created using all of the results by tallying the rankings for each
gene in the results of each algorithm. In each list, the best gene would be ranked
rst and the second best would be ranked second and so forth. The rst stage
was to take the top 5 percent of genes. The top 5 percent after pre-processing
equated to 6,036 genes, while the top 5 percent after DMP equated to 5,377
genes. Balanced accuracy and the AUC were again used as performance
measures.</p>
        <p>Finally, features were ranked and a search by trial-and-error was performed
to determine the smallest possible methylation signature.
5
5.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <sec id="sec-5-1">
        <title>Classi cation after pre-processing</title>
        <p>The resulting dataset after pre-processing was classi ed using
leave-one-outcross-validation (LOOC) using one, two or three cases to serve as a baseline for
the comparison of case elaboration and re nement strategies. Table 1 displays the
classi cation results as well as the number of metastatic samples (M) identi ed
out of 10 total M samples. The results show that only 75% of the samples were
correctly classi ed. This is a di cult problem due to the very large number of
features (120,681).
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Di erentially Methylated Positions</title>
        <p>The rst case elaboration strategy consisted in selecting di erentially
methylated probes between normal and metastatic cases. 107,497 probes were found to
be di erentially methylated, and these sites were tested to observe di erences in
classi cation performance. The resulting balanced accuracies, AUC, and M
samples identi ed is in Table 2. This table shows that results improved only slightly.
The problem remains hard to the still large number of features (107,497).
The second case elaboration and re nement strategy consisted in selecting di
erentially methylated regions. 788 probes were located within these regions, which
greatly reduced the number of features. The balanced accuracies, AUC and M
samples identi ed after classi cation at this stage is in Table 3. It is interesting to
notice that the classi cation results signi cantly improve as measures by AUC.
In addition the classi cation e ciency is greatly improved due to the signi cant
reduction in number of features.
Finally feature selection was applied to determine a methylation signature of the
metastatic cases. The top 5 percent after pre-processing equated to 6,036 genes,
while the top 5 percent after DMP equated to 5,377 genes. Balanced accuracy
and the AUC were again used as performance measures. The results for the top
5 percent after pre-processing is available in Table 4 and after DMP in Table 5.
This method generates signi cantly improved balanced accuracy and AUC over
the previous methods.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Incremental Testing of the Highest Ranked Features</title>
        <p>
          To determine a methylation signature, the top 1 feature-selected gene, top 2
feature-selected genes and so forth were selected, until reaching the top 15
feature-selected genes. The balanced accuracies and AUC for the top 1, top
5, top 10 and top 15 genes after pre-processing are available in Table 6. The
balanced accuracies and AUC for the top 1, top 5, top 10 and top 15 genes
after DMP are available in Table 7. Comparisons between case-based classi
cation and alternate methods such as Naive Bayes and Random Forest, which
showed highest classi cation performance, were performed. These tables show
that the case-based classi ers performed at least as well as the best classi ers
in this domain. Therefore, the case elaboration and re nement strategies proved
very e ective at reducing the search space and once this task accomplished the
case-based approach is just as e ective, if not more, with the advantage of
being more explainable through the possibility of showing the cases used for the
classi cation process.
These experiments show the usefulness of feature selection to both improve the
e ciency and e ectiveness of classi cation on highly dimensional data.
Whatever the feature selection method selected, classifying on 1 to 15 features yielded
improved results in most cases. In comparison with Anaissi et al., [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], we retrieve
similar cases and perform classi cation using multiple levels which further
rene the case information. We also opted out of using a synthetic oversampling
technique which we believed may have reduced variance and impacted feature
selection.
        </p>
        <p>
          Bioinformatics is particularly interested in nding gene signatures for
diseases, therefore appreciates feature selection over other methods [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. It is
therefore not surprising that this paper con rms the importance of this method in
bioformatics and its usefulness to deal with high dimensional data.
7
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper, we have proposed to apply case-based classi cation to the task of
classifying samples between normal and primary tumor with metastasis. Di
erent strategies for case elaboration and re nement were attempted to reduce the
high dimensionality of the methylation data. Our results show that case-based
classi cation performs at least as well as the best classi ers in this domain, after
selecting a pertinent methylation signature. This methylation signature will be
invaluable for interpreting the deeper pathophysiological processes involved in
the disease process. Some limitations of this work is that we have analyzed only
one type of cancer - breast - which yielded a small dataset with only 105 cases,
including 10 primary tumors from metastatic cancer. More work on independent
data remains to perform to con rm 1) the reproducibilty of the results on these
independent datasets, and 2) the validity of the selected genetic signature.
8</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We thank the State University of New York EIPF grant #172 for their support
of this work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Anaissi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Catchpoole</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braytee</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kennedy</surname>
            ,
            <given-names>P.J.:</given-names>
          </string-name>
          <article-title>Casebased retrieval framework for gene expression data</article-title>
          .
          <source>Cancer Informatics</source>
          <volume>14</volume>
          (
          <year>2015</year>
          ). https://doi.org/10.4137/cin.s22371
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bartlett</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glatt</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bichindaritz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>: Machine Learning and Feature Selection for the Classi cation of Mental Disorders from Methylation Data</article-title>
          .
          <source>Arti cial Intelligence in Medicine Lecture Notes in Computer Science</source>
          p.
          <volume>311321</volume>
          (
          <year>2019</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -21642-940
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Colaprico</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>T.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olsen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garofano</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cava</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garolini</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabedot</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malta</surname>
            ,
            <given-names>T.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pagnotta</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castiglioni</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceccarelli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bontempi</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noushmehr</surname>
          </string-name>
          , H.:
          <article-title>Tcgabiolinks: An r/bioconductor package for integrative analysis of tcga data</article-title>
          .
          <source>Nucleic Acids Research</source>
          (
          <year>2015</year>
          ). https://doi.org/10.1093/nar/gkv1507, http://doi.org/10.1093/nar/gkv1507
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutemann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>The WEKA data mining software: an update</article-title>
          .
          <source>SIGKDD Explorations</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <volume>10</volume>
          {
          <fpage>18</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hao</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krawczyk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flagg</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hou</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang, H.,
          <string-name>
            <surname>Yi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.:
          <article-title>Dna methylation markers for diagnosis and prognosis of common cancers</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>114</volume>
          (
          <issue>28</issue>
          ),
          <volume>74147419</volume>
          (
          <year>2017</year>
          ). https://doi.org/10.1073/pnas.1703577114
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Jurisica</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glasgow</surname>
          </string-name>
          , J.:
          <article-title>Applications of case-based reasoning in molecular biology</article-title>
          .
          <source>Ai Magazine</source>
          <volume>25</volume>
          (
          <issue>1</issue>
          ),
          <volume>85</volume>
          {
          <fpage>85</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lamy</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sekar</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guezennec</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouaud</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sroussi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Explainable arti cial intelligence for breast cancer: A visual case-based reasoning approach</article-title>
          .
          <source>Arti cial Intelligence in Medicine</source>
          <volume>94</volume>
          ,
          <issue>4253</issue>
          (
          <year>2019</year>
          ). https://doi.org/10.1016/j.artmed.
          <year>2019</year>
          .
          <volume>01</volume>
          .001
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mundbjerg</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chopra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Alemoza ar, M.,
          <string-name>
            <surname>Duymich</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lakshminarasimhan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nichols</surname>
            ,
            <given-names>P.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aron</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siegmund</surname>
            ,
            <given-names>K.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ukimura</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aron</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Stern, ., Gill,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Carpten</surname>
          </string-name>
          , J.D., rntoft, T.F., S rensen, K.D.,
          <string-name>
            <surname>Weisenberger</surname>
            ,
            <given-names>D.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duddalwar</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gill</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
          </string-name>
          , G.:
          <article-title>Identifying aggressive prostate cancer foci using a DNA methylation classi er</article-title>
          .
          <source>Genome Biology</source>
          <volume>18</volume>
          (
          <issue>1</issue>
          ),
          <volume>3</volume>
          (
          <year>2017</year>
          ). https://doi.org/10.1186/s13059-016-1129-3, http://genomebiology.biomedcentral.com/articles/10.1186/s13059-016-1129-3
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ramos-Gonzlez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lpez-Snchez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castellanos-Garzn</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paz</surname>
            ,
            <given-names>J.F.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corchado</surname>
            ,
            <given-names>J.M.:</given-names>
          </string-name>
          <article-title>A CBR framework with gradient boosting based feature selection for lung cancer subtype classi cation</article-title>
          .
          <source>Computers in Biology and Medicine</source>
          <volume>86</volume>
          ,
          <issue>98106</issue>
          (
          <year>2017</year>
          ). https://doi.org/10.1016/j.compbiomed.
          <year>2017</year>
          .
          <volume>05</volume>
          .010
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. dos Reis,
          <string-name>
            <given-names>M.B.</given-names>
            ,
            <surname>Barros-Filho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.C.</given-names>
            ,
            <surname>Marchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.A.</given-names>
            ,
            <surname>Beltrami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.M.</given-names>
            ,
            <surname>Kuasne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.A.L.</given-names>
            ,
            <surname>Ambatipudi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Herceg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Kowalski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.P.</given-names>
            ,
            <surname>Rogatto</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.R.</surname>
          </string-name>
          :
          <article-title>Prognostic classi er based on genome-wide DNA methylation pro ling in welldi erentiated thyroid tumors</article-title>
          .
          <source>The Journal of Clinical Endocrinology &amp; Metabolism</source>
          <volume>102</volume>
          (
          <issue>November</issue>
          ),
          <volume>4089</volume>
          {
          <fpage>4099</fpage>
          (
          <year>2017</year>
          ). https://doi.org/10.1210/jc.2017-
          <fpage>00881</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>